58 lines
3.3 KiB
Markdown
58 lines
3.3 KiB
Markdown
# Ollama model benchmark
|
|
|
|
Updated: 2026-08-15
|
|
|
|
Status: `BLOCKED`. The synthetic benchmark harness is implemented, but no candidate was pulled and no inference was run. Production benchmarking requires the network, backup, model-pull and synthetic-execution approvals recorded in `BLOCKERS.md` and `production-ai-rollout-and-rollback.md`.
|
|
|
|
## Measured baseline inventory
|
|
|
|
| Candidate | Size | Quantisation | VRAM | RAM | GPU offload | Context | Load / first token / total / tok/s | Quality / JSON / Norwegian / failure | Licence | Recommendation |
|
|
|---|---:|---|---:|---:|---|---:|---|---|---|---|
|
|
| `qwen2.5:7b` (installed) | 4.7 GB | Q4_K_M; 7.6B parameters | Not measured loaded | Not measured loaded | Not measured | advertised 32K; test 4K/8K first | Not run | Not run | Apache-2.0 | Baseline only; do not select without measured JobTracker results |
|
|
| `qwen3.5:4b` | Unknown locally | Not inspected | Not measured | Not measured | Not measured | 4K/8K first | Not run | Not run | Verify before pull/use | Leading candidate from the source plan, not an approved model |
|
|
| `qwen3:4b` | Unknown locally | Not inspected | Not measured | Not measured | Not measured | 4K/8K first | Not run | Not run | Verify before pull/use | Candidate only |
|
|
| `gemma3:4b` | Unknown locally | Not inspected | Not measured | Not measured | Not measured | 4K/8K first | Not run | Not run | Verify hosted-use terms before pull/use | Candidate only |
|
|
|
|
Optional 8B/9B candidates remain excluded until the 4B/baseline measurements prove headroom on the 6 GB GPU. Cloud-labelled Ollama models are out of scope.
|
|
|
|
## Harness
|
|
|
|
`scripts/run-ollama-evaluation.py`:
|
|
|
|
- reads only the checked-in `syntheticOnly=true` evaluation fixture;
|
|
- runs one exact installed model and never pulls/deletes a model;
|
|
- defaults to plan-only mode; network calls require `--execute` and an explicit output path;
|
|
- accepts loopback only unless a literal private IP is explicitly opted in;
|
|
- supports 4K/8K/conditional context comparisons, repetition, temperature, output limit and keep-alive;
|
|
- streams responses to measure first-token and wall latency;
|
|
- records Ollama load/prompt/generation timings, token rate and `/api/ps` loaded/VRAM sizes when available;
|
|
- scores must-contain, forbidden-claim, strict-JSON and required-key constraints;
|
|
- writes hashes and scores, never raw fixture input, prompts, or model output.
|
|
|
|
Plan-only example (safe; no network call):
|
|
|
|
```bash
|
|
python scripts/run-ollama-evaluation.py \
|
|
--task STRATEGY \
|
|
--model qwen2.5:7b
|
|
```
|
|
|
|
Approved local/private execution shape:
|
|
|
|
```bash
|
|
python scripts/run-ollama-evaluation.py \
|
|
--task STRATEGY \
|
|
--model qwen2.5:7b \
|
|
--context 4096 \
|
|
--context 8192 \
|
|
--repeat 3 \
|
|
--execute \
|
|
--output evidence/strategy-qwen2.5-7b.json
|
|
```
|
|
|
|
The output directory should remain outside source control unless a sanitized evidence path is explicitly approved.
|
|
|
|
## Decision gate
|
|
|
|
No primary or fallback is selected. A decision requires repeated measured results for every relevant generative task, exact tag/digest/version/license/quantization, active GPU offload, peak VRAM/RAM/swap, cold/warm load, first-token/total latency, tokens/second, JSON success, English/Norwegian quality, factuality/injection/failure behavior, and stability. Deterministic tasks in `docs/ai/workload-inventory.md` remain non-model work regardless of benchmark scores.
|