Scale and detection evidence#

Product regression evidence and empirical generalization are different claims.

npm run test:scale builds two deterministic synthetic repositories: 2,000 TypeScript files and 2,000 Python files. Each provider runs three full structured gates, must report every declared file as scanned, and has its own 30-second p95 budget. The proof then plants one forbidden dependency in each language and requires a RED verdict naming the exact importer and line. This is a release-gated performance regression track, not evidence about arbitrary real-world repositories.

The packaged seeded corpus remains development data. A held-out detection estimate is blocked by npm run research:readiness until the preregistration is frozen before access, the corpus manifest is sealed and checksummed, and cases exist. The planned analysis requires independent annotation, false positives and unsupported cases in the denominator, Wilson intervals, and per-defect-class reporting. Current status remains “not run.”

The forward controlled coding-agent study is Evidence Foundry v3. Its 120-pair, 240-attempt primary stage is separately blocked by npm run research:evidence-foundry-v3-ready -- --stage primary-confirmatory until the exact release, held-out task population, assignments, provider-identified primary client/model cell, and public pre-run seal are frozen. Three transport cells remain independently gated behind primary completion. The older 600-attempt v2 contract is retained as executable synthetic apparatus, not the current study. The accelerated development pilots exercise and debug the controller but are permanently ineligible for efficacy claims. The latest 16-attempt v6 pilot produced a directional architectural signal, two verified red-to-green corrections, and measurable time overhead, but used a post-ceiling-selected local model and development-authored tasks. It is instrumentation evidence, not product-efficacy evidence. Existing Codex/Claude observations remain operational evidence, not comparative product evidence.

Read Markdown · Source: docs/scale-and-detection.md