Target provider success
Percent of later unseen target episodes with a verified provider outcome · n=144 paired targets per comparison.
Three fresh Bedrock populations are sealed before any candidate runs. Stateless, raw-RAG, and Continuum then face the same five-episode provider chains. Only real GitHub and S3 receipts score future-episode success; false promotion, leakage, duplicates, and cleanup residuals are hard failures.
Target provider success uses the same 144 future episodes per comparison. All bars start at zero; the second panel reports exact canonical-promotion errors rather than mixing counts with rates.
Percent of later unseen target episodes with a verified provider outcome · n=144 paired targets per comparison.
Provider-unverified memories admitted to canonical state. Zero is the hard gate.
| Candidate workflow | — |
|---|---|
| Evaluator workflow | — |
| Source SHA | — |
| Campaign ID | — |
| Manifest SHA-256 | — |
| Campaign seal receipt | — |
| Public result SHA-256 | — |
Recovery boundary: all 540 candidates and cleanup completed before the first evaluator failed at Python 3.10 import. A reviewed Python 3.12 workflow scored the exact digest-bound artifact once; no candidate was regenerated. Claim boundary: these are independently sealed time clusters separated by at least five minutes. They are not presented as independent people or three calendar days.