04 / project
GoodEnough
Local vs Hosted LLM Non-Inferiority Study · Research
The Question
"Just run it locally" is the reflexive answer to LLM cost. It is rarely tested honestly. GoodEnough asks the specific version of that question: can a quantized 1.7B model running on a commodity CPU hold answer quality within a fixed margin of a hosted 70B model, and what does it actually cost you in latency to find out?
The design is preregistered on purpose. The non-inferiority margin (10 points), the two primary slices, the statistical test, the three-way outcome classification, and an explicit falsification condition were all committed to a public repository before the first evaluation call, so the study cannot be talked into a flattering conclusion after the fact. The amendment log is empty for every one of them. The local model is Qwen3-1.7B (Q4_K_M, 1.03 GiB) served by llama.cpp with CPU execution only on a consumer laptop; the hosted baseline is Llama-3.3-70B-Instruct via Groq. Quality is measured on 8 MMLU subjects at 100 items each, with a separate 140-item held-out router split and a 150-item GSM8K split, at 95% confidence.
Why non-inferiority, and why paired
The wrong test here is "is local better than hosted." Nobody expects a 1.7B model to beat a 70B one. The right test is "is local not meaningfully worse" for a given task, which is exactly what a non-inferiority design measures: local minus hosted has to clear a predetermined lower bound, not zero.
Every item is run through both models, so the comparison is paired: the same question, scored the same way, differing only in the model. That lets the analysis use exact paired-item statistics (McNemar for the paired accuracy comparison, Clopper-Pearson for exact binomial intervals) instead of large-sample approximations that get shaky on per-domain slices. Splits and seeds are frozen, so a re-run reproduces the same items in the same order.
The apparatus
The eval harness is dependency-free Python by design: paired-item scoring, frozen splits and seeds, and a resumable, budget-aware runner that can stop and pick back up without double-spending, instrumenting every request for token count, cost, and latency. Every statistic is implemented from first principles in the standard library, with no third-party numerical dependency, and verified against a reference library to a maximum deviation of 1.25e-7. The output is a per-domain classification into non-inferior, below-margin, or inconclusive, rather than a single headline percentage.
The result: the hypothesis was rejected
The study failed, and saying so is the point of preregistering it.
On none of the eight MMLU slices does the quantized 1.7B model establish non-inferiority against the hosted 70B at the 10-point margin. Five slices fall clearly below it; three are inconclusive at 100 items, meaning the data cannot decide them in either direction. Both slices named as primary before collection fall below the margin. Widening the definition of "good enough" does not rescue it: the conclusion is unchanged at 5 points and at 15 points.
A margin chosen after seeing the data is not a margin. That is the whole reason it was fixed in advance, and it is why this null result is interpretable rather than just a failure to look hard enough.
The result worth carrying
The more useful question is not whether the small model is worse, which was always likely, but whether it is worse in a way a router could exploit. It is not.
Across the 140 held-out router items, the local model answered correctly where the hosted model failed on 5 items. That single number bounds the entire contribution the local tier can make, because it is the only place the local answer adds anything the hosted tier does not already have. The oracle policy, correct whenever either model is correct and not a deployable policy, scores 0.893 against 0.857 for always calling the hosted model: a gap of 3.6 points, 95% interval 0.7 to 7.1.
So the accuracy headroom available to any selector over those two already-generated answers is at most a few points, and it can be computed before a router is built. That is the cheap measurement this study argues you should take first.
A second result points the same way. The obvious cheap escalation trigger, sending anything the local model fails to parse to the hosted model, fires on only 4 of the local model's 65 errors: the model fails fluently, returning well-formed wrong answers rather than nothing. A cascade built on parse failure inherits almost all of the weak model's errors while capturing almost none of the strong model's advantage.
Cost and latency, stated carefully
- Cost: the hosted side consumed $0.2550 of list-price-equivalent tokens across 1,561 calls, but ran on a free tier, so incremental API spend was zero. Local inference has no incremental API spend, which is not the same as being free: electricity and hardware amortization are out of scope here, and I do not claim a percentage saving.
- Latency: local is 8.5x slower at the median (2,901 ms against 340 ms pooled, and 10.4 to 1 on the MMLU map split alone).
The cheaper option on one axis is the more expensive option on the other, and both axes have to be priced together. That is the part the "just run it locally" reflex usually skips.
What this does not show
Two pinned configurations cannot isolate model size: this comparison bundles quantization, execution location, CPU versus accelerator, network latency, and provider queueing. Both benchmarks are public and may sit in either model's training data, so this is a valid comparison between two deployments on these specific items, not an estimate of generalization. And none of it generalizes past the pair I pinned: a larger local model, a weaker hosted reference, or a narrower task distribution could each move the answer.
What I would claim generalizes is the procedure. Declare the margin before the data, classify every slice into three outcomes rather than two, report the oracle bound alongside the router, and price latency and cost on the same page as accuracy. The whole study cost $0.26, which suggests the barrier to doing this before committing to an architecture is not resources.
Status
Submitted to the NeurIPS 2026 workshop Who Verifies the Agents? Toward Reliable Agent Development. It is under review: zero reviews have been returned and there is no decision, so nothing here is peer-reviewed yet.
Tech
Python (dependency-free) · Qwen3-1.7B GGUF (Q4_K_M) local · Llama-3.3-70B via Groq · MMLU + GSM8K · exact paired-item statistics (McNemar / Clopper-Pearson) · resumable budget-aware runner with per-request cost and latency instrumentation.
next project
AI Chatbot & Agentic Copilot
T-Mobile for Business · Enidus