Evaluation

Evaluating Local AI Models

Open-weight AI models running on local hardware can support document and content work without sending data to outside services. Their suitability, however, depends on the task, and public leaderboard scores rarely predict performance on specialized work. Models are therefore evaluated against domain-specific test suites before they are adopted.

Method

Findings

Measurement pitfalls

Prompt caching inflates speed.
Local servers cache prompts by their opening text. A benchmark that sends the same prompt for warm-up and every timed run measures a cache lookup, not computation. In one run this produced an implausibly high apparent processing speed. A unique opening sentence on every request fixed it; varying only the end of the prompt did not, because the cache keys on the beginning. Generation speed was unaffected, so only time-to-first-token and prompt-processing figures were corrupted.
Reasoning tokens consume the output budget.
On reasoning models, the token limit covers internal reasoning as well as the visible answer. A vision model asked to describe images returned blank descriptions for most of them. Its usage statistics showed the whole allowance spent on reasoning before it produced any output. With a larger limit, the same model wrote correct descriptions. A too-small budget looks exactly like a missing capability.
Comparisons should stay within one model family.
When a model has already been chosen for its quality, the performance question is how to run that model better: a different quantization, runtime, or loading configuration. Substituting a different model answers a different question and discards the reasons the first one was chosen.
The exact artifact must be recorded.
A model name is a pointer, not a specification. The same name has shipped with different quantization, different vision components, and different context lengths across downloads, and in one case with a vision component the runtime could not load at all. Every run records the file hash or commit, quantization, capabilities, and effective context length.