CONTINUUM · SEQUENTIAL BLIND
3 SEALED BATCHES · 36 CHAINS · 540 OBSERVATIONS

Does verified memory improve the next unseen episode?

Three fresh Bedrock populations are sealed before any candidate runs. Stateless, raw-RAG, and Continuum then face the same five-episode provider chains. Only real GitHub and S3 receipts score future-episode success; false promotion, leakage, duplicates, and cleanup residuals are hard failures.

Public gatefixed three-batch plan
Target pairsper comparison
Continuum target successverified outcomes
Stateless target successno memory tools
Raw-RAG target successappend-all baseline
Lift vs statelesspercentage points
Lift vs raw-RAGpercentage points
False promotionsraw-RAG / Continuum
Memory-assisted winsContinuum targets
Promotion precisionContinuum canonical
Minimum spacingbatch start seconds
Residual effectsall arms
95% CI vs statelessbatch-cluster bootstrap, pp
95% CI vs raw-RAGbatch-cluster bootstrap, pp
Sequential e-valueContinuum vs raw-RAG
Batch gatesall three required

Three-arm outcome comparison

Target provider success uses the same 144 future episodes per comparison. All bars start at zero; the second panel reports exact canonical-promotion errors rather than mixing counts with rates.

Target provider success

Percent of later unseen target episodes with a verified provider outcome · n=144 paired targets per comparison.

Continuum
Stateless
Raw-RAG
0%50%100%

False canonical promotions

Provider-unverified memories admitted to canonical state. Zero is the hard gate.

Continuum
Raw-RAG
0max —

Causal memory-compounding contract

01SeedA clean provider outcome establishes the first receipt.
02ParaphraseA fresh wording tests useful transfer without label reuse.
03PoisonUnverified instructions pressure raw append-all memory.
04StaleProvider state changes while prior text stays plausible.
05ConflictOnly provider-verified outcome evidence may survive.

Checksum-bound lineage

Candidate workflow
Evaluator workflow
Source SHA
Campaign ID
Manifest SHA-256
Campaign seal receipt
Public result SHA-256

Recovery boundary: all 540 candidates and cleanup completed before the first evaluator failed at Python 3.10 import. A reviewed Python 3.12 workflow scored the exact digest-bound artifact once; no candidate was regenerated. Claim boundary: these are independently sealed time clusters separated by at least five minutes. They are not presented as independent people or three calendar days.