Continuing from part one: the backend was destabilising long before the reports clearly showed it.
The next morning, my team focused on one question: why was CTS improving the timing numbers while making the backend physically less stable at the same time? That investigation uncovered the first major issue inside the clock strategy.
The first issue we found
After nearly two days debugging the block, we realised CTS was over-optimising skew so aggressively that the backend had started fighting itself.
At first glance, the clock reports looked impressive. Skew was extremely tight, and setup closure looked better after every rerun. But one of my engineers noticed something subtler: buffer growth was accelerating disproportionately in the same regions every cycle, particularly around macro boundaries that were already congested.
CTS was chasing picosecond-level skew gains with almost no meaningful benefit left in silicon.
The hidden physical cost
Underneath, the cost was considerable:
- Excessive buffering
- Localised routing congestion
- Difficult hold behaviour
- Unstable ECO regions
- Signal routing detouring around clock-heavy areas
The backend was technically improving, but physically it was becoming less stable with every iteration. That distinction matters more than most teams appreciate, because timing convergence and physical convergence don’t always move together.
What we changed
We decided to relax the skew targets slightly. Nothing dramatic, no major flow overhaul, just enough to stop CTS from aggressively chasing low-value skew improvements.
Immediate impact in the next cycle
- Routing pressure reduced
- Buffer growth stabilised
- Hold behaviour became more predictable
- Runtime started dropping again
- The backend stopped amplifying its own instability
That was the point at which timing behaviour finally started improving.
The clue that something deeper was wrong
One thing still bothered us. CTS was still overcompensating in exactly the same physical regions, no matter how much the skew behaviour improved. That was the clue that the deeper issue wasn’t timing at all. Something underneath the implementation itself was still fighting the clock tree.
Part three to follow …