Trial Registry · The honest record

We register every experiment before it has results — and publish every outcome.

Anyone can show you their one good week. We pre-commit to a hypothesis, measure it properly (multi-run, per engine, with a control group), and post the result whether it worked, did nothing, or backfired. This is that record.

2concluded experiments
2published nulls / non-wins
100%outcomes posted
No effect detectedstepped-wedge

Answer-first + FAQ/HowTo schema pages on asncheck.com targeting the TREATMENT how-do-I queries raise ASN appearance-rate on them vs held-back CONTROL queries (stepped-wedge; control treated later).

Engineschatgpt, gemini, perplexity, google_aio
Baseline? → 2026-07-23
After2026-07-28 → 2026-07-31
Effect (lift)0 pts
Min. detectable effect38 pts

No effect of ~38+ points was detected. At this sample size the test could only catch large effects, so smaller real gains can't be ruled out — we publish it as-is rather than bury it.

No effect detectedstepped-wedge

Weekly freshness refresh (visible last-updated date + one new sourced stat + revised copy block) on the asn-grader pages serving the TREATMENT queries raises ASN appearance-rate vs CONTROL queries whose pages stay static (causal test of the ~3.2x freshness correlate).

Engines
Baseline2026-07-23 → 2026-07-28
After2026-08-01 → 2026-08-06
Effect (lift)0 pts
Min. detectable effect40 pts

No effect of ~40+ points was detected. At this sample size the test could only catch large effects, so smaller real gains can't be ruled out — we publish it as-is rather than bury it.

How we run these

  • Pre-registration. The hypothesis, treatment/control queries, engines, and metric are locked before any work begins — so we can't move the goalposts.
  • Stepped-wedge design with a control. Some queries get the change, a matched set doesn't. If treated climbs and control doesn't over the same window and model version, the effect is real — not a model update.
  • Honest minimum detectable effect. Every result states the smallest effect the test could catch. A "no effect" at a high MDE means "no LARGE effect," not "nothing works" — and we say so.
  • Model-fingerprinted. If the underlying AI model changes mid-run, the experiment is voided rather than misreported.