Trust and evidence#
We use BCE to govern BCE.
Live checks not loadedFollow the observed main commit through our own gate, evidence checks and public deployment. Every stage links to GitHub.
- Source: mainNot checked
- Self-gate + doctorNot checked
- Tests + evidenceNot checked
- Node 22 / 24Not checked
- Pages deploymentNot checked
Live results require JavaScript and access to the public GitHub API.
Committed policy: bce-engine-architecture@0.1.1, ratified-enforced, self-ratified. Authenticated decision · How to read this pipeline.
First-party operational evidence; no independent replication or product efficacy claim. These lifecycle additions are implemented in source and v0.3.1; they are absent from the immutable npm 0.3.0 artifact.
Current decision: no product-efficacy claim. Accelerated v6 is a completed author-operated instrumentation pilot. Its observations are useful for designing the next study; they are not a recommendation to adopt BCE.
What the records support#
| Claim | State | Evidence |
|---|---|---|
| BCE deterministically detects its packaged seeded violations | Product-test supported | First-party mechanism tests |
| Accelerated v6 produced directional observations in one exact author-operated local model/client cell | Directional only | Author-operated instrumentation pilot |
What v6 observed#
- Design: 4 generated repositories, 8 repair/refactor tasks, 16 paired attempts; every randomized attempt remains in the result.
- Cell:
qwen3:8b@sha256:500a1f067a9f782620b40bee6f7b0c89e17ae61f686b92c24933e4ca4b2b8b41throughbce-ollama-tool-client 1.0.0. - Authority: tasks and machine oracles were written by the maintainer; execution was author-operated and machine-adjudicated.
| Recorded outcome | Baseline | BCE enabled | Paired record |
|---|---|---|---|
| Safe successful completion | 2/8 (25%) | 3/8 (37.5%) | +12.5 pp [0 pp, 37.5 pp] |
| Escaped defect, intention-to-treat | 2/8 (25%) | 0/8 (0%) | reduction 25 pp [0 pp, 50 pp] |
| Architecture conformance | 6/8 (75%) | 8/8 (100%) | descriptive only |
| Task success | 4/8 (50%) | 3/8 (37.5%) | descriptive only |
| Policy mutation | 0/8 (0%) | 0/8 (0%) | 0 pp |
The paired visible elapsed-time ratio was 1.317× [1.130, 1.727]. Cost was not measured. 1 baseline infrastructure timeout remains in the denominator.
Next claim-bearing study#
Evidence Foundry v3 is in design (design-draft). No stage execution is recorded.
Current public claim classes: no-efficacy-claim. This status does not state an effect magnitude. The real task manifest, exact release and client/model cells, assignments, and public pre-run seal remain unset. The first stage is deliberately bounded so a useful answer does not wait
for every transport cell.
| Registered scope | Value |
|---|---|
| Repository clusters | 40 |
| Task shapes per repository | 3 |
| Paired tasks | 120 |
| Retained attempts | 240 |
| Later transport stages | 3 |
Inspect the v3 protocol or run
npm run research:evidence-foundry-v3-ready -- --stage primary-confirmatory to see every current design blocker.
Verify the public record#
From a clean checkout, one command installs the locked verifier dependencies, checks the claim boundary, re-derives the sealed input bundle, and replays the public result:
npm ci --ignore-scripts && npm run evidence:verify
The command is local and requires no model, Ollama service, API key, or paid credential. It verifies the published record; it does not rerun model inference or make the pilot independent.
Integrity anchors#
This page is generated from the public claim index and sealed v6 result. Change a denominator, model identity, eligibility flag, or digest without changing its source record and the docs build turns red. The machine-readable source is research/claim-evidence-matrix.json.
- Study:
bce-accelerated-instrumentation-pilot-v6-2026-09-05 - Result SHA-256:
bb6ef2317d7e4df0d38d08b6e808d918cfa6f71c9a28db8d31457fb8eeed787c - Sealed-input root:
e39339e36f221108d40b06abf6863b248e8ee776dc8b2fe32c6d40ff7e34694b - Full record: accelerated-v6/RESULTS.md
The byte-immutable v6 record preserves its then-current v2 next-step language. That historical plan is superseded by the Evidence Foundry v3 registry and lifecycle shown above.
What remains unestablished#
- BCE generalizes to held-out repositories.
- BCE improves coding-agent outcomes.
- BCE reduces cost or iteration count.
- BCE governance is independently enforced on GitHub.
Other trust records#
- What is measured, and by whom — the landing page's Evidence and limits section states the position in full: every proof in this repository is produced by machinery its authors wrote, run on infrastructure its authors control, and that is not the same thing as independent confirmation.
- Independent witnesses: 0. ATTESTATIONS.md is the witness ledger, published at its honest count. The one-minute, offline procedure for adding a row — including a run that contradicts the documentation — is docs/launch/witness-kit.md.
- Citation metadata is software-only. CITATION.cff carries no provisional paper, arXiv, or DOI identifier. A preferred paper citation is added only after a real manuscript and archival record exist; scripts/check-release-citation.mjs refuses placeholder identifiers.