Continuing from part three: the backend had turned the corner, until setup pressure increased again close to tapeout.
Close to tapeout, timing pressure changes engineering behaviour. Teams start accepting optimisations they’d normally leave alone earlier in the flow, and that’s exactly what happened here.
The late-stage decision
The client enabled more aggressive useful skew optimisation, trying to squeeze out additional setup closure before tapeout. At first, the results looked encouraging: setup violations improved, reports looked cleaner, and for a short while it genuinely felt like the backend was converging.
Then the instability spread again. Within days:
- Hold fixes started oscillating between iterations
- ECO predictability collapsed
- Previously stable regions became sensitive again
- Routing behaviour turned inconsistent run to run
At that point, the team wasn’t converging any more. They were reabsorbing clock instability every cycle.
Backing out
We made the call to back out the aggressive skew changes and return the clock behaviour to the last stable convergence point. Rather than chasing further setup gains through increasingly unstable CTS behaviour, we shifted to controlled, incremental timing ECOs on the remaining critical paths. The priority became simple: stop destabilising the backend first, then close timing without reopening convergence risk.
Why this matters more broadly
This is the part many teams underestimate. A routing issue tends to stay local to one region. A CTS instability spreads across the entire backend, especially late in the flow when everything is already tightly coupled. That’s why we never treat CTS as a standalone implementation stage. We treat it as the convergence backbone of the whole RTL-to-GDS flow.
Close to tapeout, the aim isn’t a mathematically optimal clock tree. It’s a clock tree stable enough that the rest of the chip can converge around it before schedule pressure and runway collide.
Over the years, the lesson has held: late-stage backend recovery is rarely about chasing better numbers. It’s about protecting convergence stability for long enough to reach tapeout safely.
That’s the full recovery, start to finish. If you’re navigating something similar on your own programme, get in touch.
