Evidence ledgerNumbers with names, dates and limits.
Rangabot reports numerator and denominator, suite version, model, hardware context and execution errors. Targeted reruns never become complete-suite claims.
Current source invariants
Deterministic checks for what the software does.
These checks cover recovery, local binding and the production dependency tree. They do not measure answer quality or prove a signed package.
InvariantResultEvidence scopeState
Manual failed-turn recovery24/24Independent source-level reconstruction, integrity, changed-resource and no-auto-run cases
Pass Conversation/resource binding differential5/5Protected cases; merged base was 0/5
Pass Production dependency audit0npm audit --omit=dev; development packaging advisories are disclosed separately
Pass Historical model quality
Useful results, with later invalidation kept visible.
The historical v1.0.11 run scored 59/60 and 22/22 critical. A later v1.0.12 run scored 58/60, but its scorer accepted a false premise, so it is invalid standalone release evidence. Exact v1.0.13 machine and blind-human evidence remains NOT RUN.
CapabilityResultEvidence scopeState
Historical core conversation result59/60v1.0.11 · llama3.2:3b · one complete structural run
Conditional Historical critical trust cases22/22v1.0.11 one-run result; repeated release review remains open
Conditional Reasoning cases5/5Same candidate and frozen rubric
Pass Memory relevance precision15/15Synthetic selection audit
Pass Memory relevance recall15/15Synthetic selection audit
Pass Teacher answer quality50/60Latest recorded public result
Below gate Teacher grounding54/6090%; target is at least 95%
Below gate Analytical transfer12/12Astronomy v4 and separately frozen theatre v5 regressions
Pass Trusted analytical narration44/44Canonical renders; 222/222 hostile mutations rejected
Pass SQL disclosed regression720/720Engineering regression after the 571/720 first look; not a hidden holdout
Pass Business-analyst disclosed regression624/64097.5%; 16 ambiguous target-variance prompts clarified
Conditional External BIRD Mini-Dev18/5003.6% with typed context, up from 0.2%; general SQL gate not met
Below gate Where Rangabot shines nowStrong when the work is grounded and the evidence is visible.
- 720/720Known SQL failures stayed fixed in the disclosed engineering regression.
- 624/640Business-analyst questions passed after schema disclosure; 16 ambiguous prompts were clarified rather than guessed.
- 44/44Canonical analytical narration passed, while 222/222 hostile mutations were rejected.
- Local firstPermissions, provenance and private-data boundaries remain part of the product—not an afterthought.
Where it is improvingOpen-world interpretation is still the honest frontier.
- 18/500External BIRD Mini-Dev accuracy is 3.6%, up from 0.2%, but far below a general SQL expert claim.
- 0/102Challenging BIRD questions remain unsolved in this measured run.
- 55.4 secMean BIRD latency is still too slow for a dependable everyday analyst workflow.
- Below gateTeacher answer quality and grounding still need stronger repetition and independent review.
Internal regressionShows that defects we found and repaired stay closed under disclosed schemas. It is engineering proof, not an unseen exam.
≠External BIRDTests transfer to unfamiliar schemas and natural-language questions. It is the release counterweight, so the two scores are never pooled.
What this proves
Source invariants and model quality are different claims.
Deterministic source tests can prove a recovery or permission boundary without proving a good answer or a reliable packaged app. Model behavior varies by model, quantization, context, hardware and run. Disclosed regression suites prove identified defects stayed closed; the external 18/500 BIRD result shows the remaining open-world SQL gap.