02 / project
semblance
Verified Error-Bounded LLM Caching Gateway (Go) · Personal
What it is
semblance is an OpenAI-compatible LLM gateway written in Go that adds verified semantic caching. Ordinary semantic caches (LiteLLM, Portkey, GPTCache) make you pick a similarity threshold and hope. semblance instead learns, per cache entry, when reuse is safe, and holds the overall rate of wrong cached answers under an error budget you choose.
The method is not mine. It is a clean-room implementation of the vCache algorithm (Schroeder et al., Verified Semantic Prompt Caching, ICLR 2026, arXiv:2502.03771, UC Berkeley Sky Computing Lab), which shipped as a Python research library. semblance's contribution is the production system: a real Go gateway on Kubernetes, with the concurrency, eviction, and cold-start problems the paper leaves open actually solved.
The honest claim
Semantic caching is commodity. The verified error bound is the paper's idea. What I built is the engineering around it: a drop-in gateway any unmodified OpenAI client can point at, plus the systems work to make the algorithm survive real traffic. I cite the paper because the value here is rigor and execution, not novelty.
How the decision works
Each cache entry fits a per-entry logistic model of P(correct given similarity), written by hand in Go via IRLS with a delta-method confidence band and no ML library. On a new request the gateway embeds the final user turn, finds the nearest neighbor inside an exact-match bucket (keyed by model, system prompt, temperature, and prior turns), and runs a randomized explore-or-exploit rule. Exploit serves the cached answer; explore calls the backend, returns it, then asynchronously labels the result and updates the model. The randomization is what lets the realized error rate provably respect the budget instead of drifting.
The systems work the paper skips
A research library does not have to survive concurrency or eviction. A gateway does. semblance adds a sharded thread-safe store (race-detector clean), LRU eviction with a bounded per-entry observation cap, a defined cold-start rule for entries with too few observations, and cached model fits so the hot path never refits per request. Around the cache sits the production surface: streaming (SSE) and non-streaming passthrough, static API-key auth, per-key spend budgets that return HTTP 429 on breach, and Prometheus metrics. It ships as a multi-stage distroless Docker image onto a local k3s cluster via a Helm chart.
The result
I benchmarked it by offline replay of the authors' own public benchmarks, feeding the datasets' precomputed embeddings through the real cache and policy. Two arms, static threshold (the GPTCache and LiteLLM baseline) versus the verified policy at four error budgets, 40,000 queries per run, fixed seed, no eviction.
The guarantee held in every run. Across both datasets and all four budgets (8 verified runs) the realized error rate came in at or under its budget.
- On LmArena (dense paraphrases) at a 5% budget, the verified cache served 7.9% of requests at 0.20% error, a sub-1% error regime no static threshold can reach (the best static floor is 1.62%).
- On SearchQueries (sparse) an untuned static threshold runs at 37.6% error, while the verified policy stays safely under budget (1.8% hit at 0.59% error). That gap is the argument for a guarantee: you do not have to find the right threshold per workload.
The limitation I will state plainly
The verified policy is conservative. At the 5% budget on LmArena it runs about 24x under budget (0.20% versus 5%), which leaves hit rate on the table. Lowering the cold-start floor from 5 observations to 3 roughly doubled hit rate while still holding every budget; the sensitivity study is in DECISIONS.md. Squeezing more of the budget safely is the main open direction. The project's value was never a big hit-rate number. It is a guaranteed error dial that threshold caching cannot offer, plus the engineering and the rigor to back it.
One design point I like
There is deliberately no live realized-error metric in production. An exploit serves the cache and never calls the model, so you can never observe the true error rate at serving time. Realized error is a benchmark-only quantity; production exposes the budget target plus honest proxies (a judge-observed miss rate and the hit rate). Knowing why the metric cannot exist is the point.
next project
AI Chatbot & Agentic Copilot
T-Mobile for Business · Enidus