Every other coordinate move is a step change. The cache is the only ratchet.

Quantize once: L₂→L₁, done. RAG once: facts→navigation, done. Each move is a one-time relocation. The cache is different. The cache compounds. Each query that hits the cache costs ~zero and makes the same query low-cost forever.


The LLM is not the system. It's the exception handler.

A model is a fixed function: same cost on day one and day one thousand. A memory is the opposite: it accumulates, amortizing cost over hits. The field built models. It should have built model+memory systems — where the memory is the primary path and the model is the fallback for cache misses.

This is an architectural inversion, not an optimization. In the current architecture, every query runs the model. The cache, when it exists, is a side-effect — an opportunistic short-circuit on the model's output. In the inverted architecture, every query first asks the memory; the model runs only when the memory misses. The model becomes a function of cache misses, not a function of users.

Under the inverted architecture, the workload's cost is dominated by the cache, not the model. As the cache fills, the model engages less. The cost curve bends down because the *mechanism mix* shifts toward the lower-cost coordinate. The model is still there, doing what only it can do — on the questions the memory hasn't seen.


The cache's coverage is the efficient-coordinate fraction of the workload.

Cache hits require that the same input map to the same output. The condition is trivially true in some zones and basically false in others. Once you name what's cacheable per zone, the architectural question stops being "should we cache?" and starts being "what is the right granularity to cache at?"

Closed-form · Z₁

trivially cacheable

Math = words. Same input, same output, forever. A hash of the input is the cache key; the value is the deterministic answer. Lookup is O(1) on a memory read. ~10⁹× lower-cost than recomputing.

Constraint-bound · Zc

cacheable with invalidation

Formal rules, combinatorial answers. The Sudoku solved today is the Sudoku solved tomorrow. The cache invalidates when the rules change — dependency manifest updates, schedule constraint added, type signature shifted. Versioning the substrate is mandatory.

Pattern-bound · Z₂

retrieval yes, judgment no

The retrieval layer is cacheable (the statute, the regulation, the codebase doesn't change between queries). The judgment layer is not — the same statute applied to two different fact patterns yields two answers. Cache the source. Don't cache the application.

Pattern-rich · Z₂↔Z₃

embeddings yes, output no

In-style synthesis — drafting, summarizing, classification. The intermediate embeddings cache well (same input chunks produce the same embeddings). The final stylized output rarely repeats; caching it overfits to one phrasing.

Open-ended · Z₃

basically not

Math ≠ words. "Write me a poem about loss" must not return the cached poem. Novelty is the point. The only cacheable artifact here is the system prompt — everything downstream is unique synthesis.

Perceptual · temporal

window-cacheable

Real-time sensing. The preprocessing caches in short windows (the same frame, the same audio chunk). The decision usually doesn't — the world has moved by the next tick.

Seam discipline: cache the parse, never the judgment. At every Z₁↔Z₂ and Z₂↔Z₃ boundary, the lower-cost side caches and the higher-cost side doesn't. Pipeline composition is exactly the architecture that lets each stage cache at its right granularity.


A hit is low-cost, not free.

"Caching is free" overclaims and invites the rebuttal that breaks the argument. A cache hit is a memory read (L₁ measurement) plus a hash lookup (~L₀). It costs picojoules. It's low-cost.

The cache itself has thermodynamic overhead: storage (idle power on the memory holding entries), eviction (decisions about what to drop), invalidation (knowing when entries are stale). These costs are bounded and amortized; they don't scale with query volume the way model inference does.

The ratio is the whole game. An L₀ memory read at ~10 pJ vs. an L₂ LLM inference at ~3 J. That's 3 × 10⁸ times lower-cost per answer. Even at low cache hit rates — 30%, 40% — the workload's average cost is dominated by the misses. At higher hit rates — 80%, 90% — the average cost approaches the cost of the cache lookup itself.

E(workload) = (1 − h) · Emiss + h · Ehit

Workload energy · weighted by hit rate h

With Emiss ≈ 3 J and Ehit ≈ 10 pJ, the workload's average cost at hit rate h is ~(1−h) · 3 J. The arithmetic is brutal: every percentage point of cache hit is worth 30 mJ per query.


A stale cache is worse than no cache.

A cached legal answer that's now wrong because the statute changed is confidently wrong, with a citation. A cached schedule that's now infeasible because a constraint was added still gets shipped. A cached fact that the world has updated still gets returned. The cache's correctness depends entirely on its invalidation discipline.

Invalidation requires substrate versioning. The cached answer is keyed not just on the question but on the version of the source. When the source updates, all dependent cache entries invalidate. When a regulation changes, the legal cache invalidates the affected entries; the rest stay warm. When a dependency manifest updates, the resolved-version cache invalidates the affected packages; the rest stay cached.

A versioned substrate is a cache that knows when it's stale. The cache and the live-versioned LUT are the same artifact at two scales — one is the technique, one is the product.

The frontier cell on the Stack named "live-versioned regulatory LUT" is exactly the cache-with-invalidation discipline, productized. Whoever builds that for a domain owns the substrate every regulated-AI product in that domain has to license. The cache is the engineering practice. The versioned substrate is the company.


The only mechanism that bends cost down over time.

Workloads are power-law — a small set of queries account for a large fraction of volume. As the cache samples the distribution, the hit rate climbs faster than linear. After 1M queries on a typical workload, the cache covers ~70−85% of volume. After 1B, ~95%. The arithmetic of the average query gets lower-cost monotonically as usage grows.

Workload size Hit rate Avg cost / query Vs. baseline
First 1K queries ~20% ~2.4 J 1.25× lower-cost
First 100K ~55% ~1.35 J 2.2× lower-cost
First 1M ~78% ~660 mJ 4.5× lower-cost
First 100M ~92% ~240 mJ 12.5× lower-cost
First 1B ~96% ~120 mJ 25× lower-cost

Illustrative power-law workload with Emiss = 3 J, Ehit = 1 mJ. Real numbers depend on access patterns, cache size, and invalidation rate — but the shape is reliable.

Compare to a traditional LLM system: ~3 J per query at every scale. After 1B queries, the cached system is 25× lower-cost on average than the flat-cost system, and the gap is still widening. Every other coordinate move is a step function. The cache is the slope.


Caches make unit economics asymmetric.

A vendor with a real cache has monotonically improving unit economics. A vendor without one has flat unit economics. After three years on the same workload, the cached vendor's cost per decision has dropped by an order of magnitude; the uncached vendor's has dropped by exactly nothing.

In a long-term procurement contract, this asymmetry is decisive. The buyer who locks in a cache-equipped vendor for ten years gets a unit cost trajectory the uncached competitor cannot match without throwing away their architecture. The first vendor with a real cache wins the long-term contracts, then locks in the customer base, then funds the next round of improvements from the compounded margin.

Capability marketing doesn't survive this. A frontier-model competitor with no cache discipline gets undercut on year-three pricing by a less capable but cached system. The cache is the moat capability can't dig around.


The mechanism that bends the curve.

The cost function names the choice. Composition gives every stage its own cache. The receipt records what each hit cost. The flywheel turns.