3.3 KiB
Ollama model benchmark
Updated: 2026-08-15
Status: BLOCKED. The synthetic benchmark harness is implemented, but no candidate was pulled and no inference was run. Production benchmarking requires the network, backup, model-pull and synthetic-execution approvals recorded in BLOCKERS.md and production-ai-rollout-and-rollback.md.
Measured baseline inventory
| Candidate | Size | Quantisation | VRAM | RAM | GPU offload | Context | Load / first token / total / tok/s | Quality / JSON / Norwegian / failure | Licence | Recommendation |
|---|---|---|---|---|---|---|---|---|---|---|
qwen2.5:7b (installed) |
4.7 GB | Q4_K_M; 7.6B parameters | Not measured loaded | Not measured loaded | Not measured | advertised 32K; test 4K/8K first | Not run | Not run | Apache-2.0 | Baseline only; do not select without measured JobTracker results |
qwen3.5:4b |
Unknown locally | Not inspected | Not measured | Not measured | Not measured | 4K/8K first | Not run | Not run | Verify before pull/use | Leading candidate from the source plan, not an approved model |
qwen3:4b |
Unknown locally | Not inspected | Not measured | Not measured | Not measured | 4K/8K first | Not run | Not run | Verify before pull/use | Candidate only |
gemma3:4b |
Unknown locally | Not inspected | Not measured | Not measured | Not measured | 4K/8K first | Not run | Not run | Verify hosted-use terms before pull/use | Candidate only |
Optional 8B/9B candidates remain excluded until the 4B/baseline measurements prove headroom on the 6 GB GPU. Cloud-labelled Ollama models are out of scope.
Harness
scripts/run-ollama-evaluation.py:
- reads only the checked-in
syntheticOnly=trueevaluation fixture; - runs one exact installed model and never pulls/deletes a model;
- defaults to plan-only mode; network calls require
--executeand an explicit output path; - accepts loopback only unless a literal private IP is explicitly opted in;
- supports 4K/8K/conditional context comparisons, repetition, temperature, output limit and keep-alive;
- streams responses to measure first-token and wall latency;
- records Ollama load/prompt/generation timings, token rate and
/api/psloaded/VRAM sizes when available; - scores must-contain, forbidden-claim, strict-JSON and required-key constraints;
- writes hashes and scores, never raw fixture input, prompts, or model output.
Plan-only example (safe; no network call):
python scripts/run-ollama-evaluation.py \
--task STRATEGY \
--model qwen2.5:7b
Approved local/private execution shape:
python scripts/run-ollama-evaluation.py \
--task STRATEGY \
--model qwen2.5:7b \
--context 4096 \
--context 8192 \
--repeat 3 \
--execute \
--output evidence/strategy-qwen2.5-7b.json
The output directory should remain outside source control unless a sanitized evidence path is explicitly approved.
Decision gate
No primary or fallback is selected. A decision requires repeated measured results for every relevant generative task, exact tag/digest/version/license/quantization, active GPU offload, peak VRAM/RAM/swap, cold/warm load, first-token/total latency, tokens/second, JSON success, English/Norwegian quality, factuality/injection/failure behavior, and stability. Deterministic tasks in docs/ai/workload-inventory.md remain non-model work regardless of benchmark scores.