Evidence ledger

Numbers with names, dates and limits.

Rangabot reports numerator and denominator, suite version, model, hardware context and execution errors. Targeted reruns never become complete-suite claims.

CapabilityResultEvidence scopeState
Core conversation candidate59/60

v1.0.11 · llama3.2:3b · complete run

Conditional
Critical trust cases22/22

One complete run; repetition still required

Pass
Reasoning cases5/5

Same candidate and frozen rubric

Pass
Memory relevance precision15/15

Synthetic selection audit

Pass
Memory relevance recall15/15

Synthetic selection audit

Pass
Teacher answer quality50/60

Latest recorded public result

Below gate
Teacher grounding54/60

90%; target is at least 95%

Below gate
Analytical transfer10/12

Frozen astronomy holdout

Below gate
What this proves

A candidate, not a universal promise.

Model behavior varies by model, quantization, context, hardware and run. These results describe exact evaluated candidates. They do not imply that every Ollama model will match them.

Private fixtures stay private.

Full model answers, personal chats, saved memories, document titles and Knowledge Vault files remain Git-ignored. Public methodology and aggregate results are reviewable.

Read the frozen contract