Anyone can show you their one good week. We pre-commit to a hypothesis, measure it properly (multi-run, per engine, with a control group), and post the result whether it worked, did nothing, or backfired. This is that record.
No effect of ~38+ points was detected. At this sample size the test could only catch large effects, so smaller real gains can't be ruled out — we publish it as-is rather than bury it.
No effect of ~40+ points was detected. At this sample size the test could only catch large effects, so smaller real gains can't be ruled out — we publish it as-is rather than bury it.