Evaluation
Evaluating Local AI Models
Open-weight AI models running on local hardware can support document and content work without sending data to outside services. Their suitability, however, depends on the task, and public leaderboard scores rarely predict performance on specialized work. Models are therefore evaluated against domain-specific test suites before they are adopted.
Method
- Domain-specific suites. Test cases are written from real working material: document text, defective interface markup, and the kinds of requests the models will actually receive. A separate suite covers cleanup of text produced by optical character recognition (OCR).
- Deterministic scoring. Each case has explicit pass criteria, such as a required label, a valid structured output, or a specific recommendation, so results do not depend on a second model's judgment.
- Enough cases to separate signal from noise. An early, smaller suite proved inconclusive: a single changed answer moved a model's score enough to reverse the ranking. The suite was expanded until differences between models held across runs.
- Controlled conditions. Models are loaded one at a time on the same machine, with identical prompts and settings, so no model competes with another for memory or compute.
- Cross-family comparison. A single bake-off compared a dozen models from several families (Gemma, Qwen, Granite, Nemotron, and OCR-specialized models) across sizes, from under one billion to over thirty billion parameters.
Findings
-
Size is not a reliable predictor of quality.
A mid-size mixture-of-experts model (Gemma 4 26B-A4B), which activates only a fraction of its parameters for each token, outscored the largest dense model tested (Gemma 4 31B) on the domain suite overall and on the remediation tasks specifically. It is also faster, since less computation is performed per token. It became the default for most work, with the larger model reserved for tasks that need longer chains of reasoning.
-
Following a strict output format requires a minimum model size.
On document triage, where the answer must be one label from a fixed set, small models frequently produced explanations, near-miss labels, or extra text instead of the required label. Models in the mid-size class and above followed the format reliably. Triage is therefore routed to the mid-size tier regardless of how simple each individual decision appears.
-
Specialists can underperform outside their narrow task.
Models built for OCR (such as GLM-OCR) were substantially worse at correcting OCR output than a small general-purpose model. OCR specialists are trained to convert images into text, not to repair text that is already extracted: fixing broken hyphenation, merged words, and misread characters calls for general language knowledge. A caveat applies: the test harness was text-only, so the specialists' core task of reading images was not measured.
-
Very small models make effective routers.
A sub-one-billion-parameter model (Qwen 0.8B) passed every case in the classification, structured-data extraction, prompt-safety screening, and decision-labeling categories, returning answers in under a second. Placed in front of larger models, it can sort incoming requests and pass each to the model best suited to it, reserving expensive computation for the requests that need it.
-
Some tasks defeat every model tested.
Certain design-judgment recommendations were missed across all models, including the largest. The clearest example: when a heading was presented as an image of text, models rarely recommended rebuilding it as real, styled text. Problems of skipped heading levels and of reading order differing from visual order also stayed difficult at every size. These tasks keep a human reviewer in the loop and are candidates for targeted prompting or fine-tuning rather than larger models.
-
Quantization is a meaningful design choice.
Two compressed (quantized) builds of the same mid-size model were compared: a community build and the vendor's own quantization-aware-trained (QAT) build, in which the model is trained to tolerate compression. Quality was effectively tied on both the domain and OCR suites, while the vendor build generated text noticeably faster. The vendor build was adopted. For the smallest models, the community builds matched or slightly exceeded the vendor builds, so those were kept.
Measurement pitfalls
- Prompt caching inflates speed.
- Local servers cache prompts by their opening text. A benchmark that sends the same prompt for warm-up and every timed run measures a cache lookup, not computation. In one run this produced an implausibly high apparent processing speed. A unique opening sentence on every request fixed it; varying only the end of the prompt did not, because the cache keys on the beginning. Generation speed was unaffected, so only time-to-first-token and prompt-processing figures were corrupted.
- Reasoning tokens consume the output budget.
- On reasoning models, the token limit covers internal reasoning as well as the visible answer. A vision model asked to describe images returned blank descriptions for most of them. Its usage statistics showed the whole allowance spent on reasoning before it produced any output. With a larger limit, the same model wrote correct descriptions. A too-small budget looks exactly like a missing capability.
- Comparisons should stay within one model family.
- When a model has already been chosen for its quality, the performance question is how to run that model better: a different quantization, runtime, or loading configuration. Substituting a different model answers a different question and discards the reasons the first one was chosen.
- The exact artifact must be recorded.
- A model name is a pointer, not a specification. The same name has shipped with different quantization, different vision components, and different context lengths across downloads, and in one case with a vision component the runtime could not load at all. Every run records the file hash or commit, quantization, capabilities, and effective context length.