FSL Quantum Experiment 001 comparing predictive and reactive quantum error correction strategies

We Tried to Prove Our Idea Worked. The Data Said No.

Forward Security Labs · Research Note · Experiment 001

Internal research · Result: negative

We tried to prove our idea worked. The data said no.

Most incident response plan testing is built to confirm the plan works. That is backwards. A good test is built to break it.

Hypothesis Not supported

4 falsification criteria set in advance · 3 triggered · idea closed

This is a record of what happens when you run a test the other way around. Forward Security Labs ran an internal research experiment in fault tolerant quantum computing, designed from the start to break its own hypothesis.

The hypothesis sounded reasonable. If you can predict that part of a quantum computer is becoming unreliable, moving critical quantum information before it fails should produce fewer errors than waiting for the failure and reacting to it.

So we tested it.

It failed.

And that failure ended up being more useful than the result we originally hoped to find.

The idea

Quantum computers have an enormous reliability problem. Qubits are fragile. Errors happen. Modern quantum error correction attempts to preserve logical information despite those physical errors.

Our question was simple: instead of waiting until degradation becomes a problem, could we forecast that degradation and move the logical information somewhere healthier first?

Think of it like evacuating a building before the fire reaches it instead of waiting for the alarm.

On paper, it made sense. But plans that make sense on paper are exactly the ones worth testing.

Our first result looked great

The first version of the experiment appeared to show roughly a 33 percent advantage for the predictive approach.

That would have been exciting. It was also wrong.

During review, we found that the comparison was unfair. The reactive system had been given a simple binary trigger. The predictive system had been given a more sophisticated probability based decision.

We were not measuring prediction versus reaction. We were partly measuring a sophisticated decision system versus a simpler one.

So we threw that result away. The experiment was rebuilt so both approaches received the same information, used the same decision rules and operated under the same conditions. The only meaningful difference left was whether one attempted to predict the future.

Then we ran it again.

Prediction mostly did not help

We tested five types of simulated hardware degradation. Lower is better in every row below. The percentage is the change in logical failures when prediction was added to a reactive system that had already been tuned to its own best setting.

Final result · 60 paired trials per scenario
Slow degradation
-1.1%
Negligible
p = 0.047
Sudden changes
+7.5%
No convincing advantage
p = 0.111
Burst failures
+37.5%
Prediction made it worse
p < 0.001
Random fluctuation
-18.2%
Prediction helped
p < 0.001
Mixed conditions
-12.9%
Prediction helped
p = 0.002

The predictive approach produced essentially no meaningful improvement during slow degradation, and no statistically convincing advantage during sudden changes.

During simulated cosmic ray burst events, something worse happened. Prediction produced 37.5 percent more logical failures than the reactive approach.

Why? Because the burst happened too quickly to meaningfully predict. The predictive system reacted late and could move information into a neighboring region that had also been affected by the same event.

The supposedly smarter system made the situation worse.

Then we found something more important

Prediction did appear to perform well under noisy and mixed conditions. Initially, that looked promising. Those are two of the five scenarios, and they favored the idea.

But when we examined forecasting performance at different time horizons, the improvement barely changed.

Forecast skill versus persistence · random fluctuation
+0.271 sec
+0.275 sec
+0.2710 sec
+0.2730 sec

Flat across a thirtyfold range of horizons. A useful forecast would normally decay as it reaches further out. This one did not.

That should not happen with a real forecast. Predicting something 30 seconds into the future should generally be harder than predicting it one second into the future. Instead, measured skill remained essentially flat across all four horizons.

The evidence suggests the apparent advantage came primarily from filtering present-state noise rather than useful forward prediction.

That is useful. But it does not appear to require a forecasting system. A much simpler filter may provide the same benefit at a fraction of the complexity.

The assumption underneath the whole idea was wrong

There was another surprise.

Our hypothesis assumed moving quantum information would be expensive enough that knowing when to move would be extremely important.

The experiment did not support that assumption. Under the modeled conditions, the direct error cost of relocation was tiny. We deliberately swept that cost across five orders of magnitude to test whether the conclusion was sensitive to the assumption. The answer barely moved.

The binding constraint was capacity. Every relocation needs somewhere healthy to go, and resources undergoing recalibration become temporarily unavailable.

That changes the problem. The important question was never how early should we move. It was closer to how efficiently can we manage the healthy resources we have.

Our hypothesis was built around timing. The experiment told us timing was not the primary lever.

So we stopped

We established four falsification criteria before accepting the idea. Three were triggered.

  • TriggeredPrediction failed to beat reaction where forecasting was supposed to matter.
  • TriggeredForecasting failed to beat simple persistence at useful horizons in several scenarios.
  • Not hitAdvantage should vanish under changed cost assumptions. It did not, but only because cost turned out to be irrelevant at any value, which is worse news.
  • TriggeredThe apparent wins came from filtering noise, not predicting degradation.

So we killed the idea. No moving the goalposts. No selecting only the favorable graphs. No rewriting the hypothesis after seeing the results.

The hypothesis failed.

What this means for incident response plan testing

This experiment was not ultimately about quantum computing. It demonstrated something much more relevant to the work organizations do every day.

A plan can be well researched, technically sophisticated, logical, supported by smart people, and completely wrong once reality starts pushing against its assumptions.

That is why incident response plan testing cannot happen on paper. Business continuity plans and security procedures have the same problem. The document is not the capability. Put people under pressure instead.

  • Remove a dependency they expected to have.
  • Introduce conflicting information.
  • Make the primary system unavailable.
  • Force someone to decide without complete information.

Then watch what happens.

This is not a new idea. NIST SP 800-84, the federal guide to test, training, and exercise programs, has said for years that exercises exist to validate whether people can actually perform their roles, not to confirm that a document exists. Most organizations still run the confirming version.

The objective of a good exercise is not to prove your organization is prepared. It is to discover that you are not prepared while the consequences are still simulated.

The worst calls I ever ran were not the ones with the most blood. They were the ones where the plan looked fine on paper and fell apart in the first two minutes.

Security is no different. If your incident response plan testing has never included a scenario built to break the plan, you do not have a plan. You have a draft.

Find out where your plan breaks

Forward Security Labs runs facilitated tabletop exercises built to stress a plan until it breaks, in a room where nothing is actually on fire.

See the tabletop exercises
About Experiment 001

FSL Quantum Experiment 001 used real surface code error correction simulation through Stim and PyMatching. Hardware degradation traces were synthetic, modeled from published research rather than measured on a commercial quantum computer. No vendor hardware data was used and none is claimed.

The primary comparison used 60 paired trials per degradation scenario. Tuning was performed on separate trials from final evaluation. Every policy in a given trial saw the same degradation trace and the same sensor data.

This should be read as a falsification test of one hypothesis under modeled conditions, not as a claim about the performance of any specific commercial quantum computing system. Full source code, raw results and a complete limitations disclosure are available on request.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *