Nice result.
On the argument about whether the loops are hidden refusal, I don't think your data settles it either way, because all 40 of your prompts are sensitive. Forty boring ones would do it. If the loops really were refusal in disguise, they'd run long on the sensitive prompts and stay short on the dull ones.
One thing to watch if you try that. Most refusal counters treat a cut-off thinking trace as the model complying, so a model that quietly obstructed scores the same as one that actually answered. I built mine to log those as "no answer" and read only what comes after the thinking. It's called senbonzakura (https://github.com/elementmerc/senbonzakura). It's open source, if it saves you writing your own.