A cost function for information synthesis
What if we could estimate the cost of a piece of work before doing it — and pick the right tool?
The capability of a model is irrelevant if it has nothing to do with the work. Every AI provider sells modes — fast, balanced, expert — but these tiers don't tell you which one fits the task in front of you. The choice the user is being asked to make is the wrong choice.
Estimate the cost. Read the shape of the work. Pick the right tool. Repeat.
A picture of the current question
Three providers. Three tiers each. One question to the user.
Each provider asks the user to pick a capability tier before the work runs.
The user does not know which mode fits the work. The model picker does not know either — because there isn't one. There is a UI for a question that should not have been asked of the user at all.
Part 01 · The choice that actually matters
A model is the wrong unit of choice.
Capability is one number. Fit is many. A task has a shape — what kind of answer it needs, how much input, how strict the determinism, what counts as correct. Capability tells you how powerful the engine is. Fit tells you whether the engine is the right one for the road.
A frontier model running a closed-form transform is a Formula 1 car driving to the grocery store. It will do the job. The energy spent is absurd. The task requires correctness, which a 300M-parameter specialized model can also provide.
The honest question is not how powerful a model. It is which mechanism, at what cost, with what confidence of fit. The set of candidate mechanisms is broader than three buttons. It includes deterministic transforms. It includes retrieval. It includes specialized small models. It includes the frontier model. The question is which one this task wants.
Part 02 · The cost function
Estimate before you run.
The point of a measurement framework is to know what something will cost before you do it. The cost function is the artifact that does this.
E ( task, candidate ) → cost, fit, confidence
Given a task and a candidate mechanism, return cost, fit, confidence
- task shape, modality, determinism, latency, reuse
- candidate mechanism, hardware, cost per unit work
- history past performance on similar work
- cost estimated joules per task
- fit expected quality of output
- confidence how sure the fit estimate is
This is the same E from the thesis — evaluated before the work runs instead of after. The thesis equation tells you what something cost. The cost function tells you what it will cost — low-cost enough to run on every candidate, every time.
Run the function for every candidate. Rank by fit and cost. The router picks. The user submits work. The user does not pick the model. The model picker has been removed from the user interface because it was never the user's question to answer.
The function itself is small. The router that uses it is small. Neither runs a large model. They are upstream of every decision about whether to run a large model at all.
Part 03 · The shapes the function reads
Six shapes of work.
The function reads the task and identifies its shape. Six categories cover the work people ask computers to do.
The answer is fully specified by the input. Formatters, parsers, hash functions, lookups, deterministic transforms. The mechanism is software that already exists.
The rules are formal, but the answer space is combinatorial. Sudoku, formal proofs, scheduling, EDA, theorem search. The mechanism is search with a verifier — an energy-based model that scores candidate states against the constraints, plus a checker that certifies the result.
The answer is constrained by an authoritative source — a codebase, a regulation, a framework. The work is retrieval plus adaptation. The mechanism is a small model with retrieval.
Recognizable shapes, no single authoritative source. Drafting in a style, summarizing, paraphrasing, classifying. The mechanism is a specialized small model.
Genuinely novel synthesis. Architecture, design, complex reasoning. The mechanism is the large model.
Real-time sensing and control. Microsecond budgets. Embedded. The mechanism is deterministic preprocessing with a specialized small model where needed.
The mapping
These six shapes are the product-facing projection of the formal coordinate axes from the seven-axis stack. Closed-form is Z₁ direct evaluation. Constraint-bound is also Z₁ — same formal rules — but the answer space is combinatorial, so the mechanism is search-with-verifier instead of look-up. Pattern-bound is Z₂ with a retrieval anchor against an authoritative source. Pattern-rich is the Z₂↔Z₃ transition — recognizable shapes without an authoritative anchor. Open-ended is Z₃ (math ≠ words). Perceptual · temporal is what happens when the interface axis (signals, bodies) dominates the zone axis — microsecond budgets, embedded contexts, where the loop closes against the physical world.
The shape comes first. The mechanism follows. The mode buttons invert this — they ask the user to pick a mechanism without ever surfacing the shape. The cost function works the right way around.
Part 04 · Two worked examples
Two tasks. One function. Different picks.
Two tasks of different shapes. The cost function runs against the same candidate set for both. The output is a ranked list — the router picks the top of each list, and the picks are different mechanisms.
The router picks Prettier. The task is closed-form. The fit is perfect. The cost is six orders of magnitude lower than the frontier model. The output is identical.
The router picks the energy-based model + verifier. The task is constraint-bound — the rules are formal, the answer space is combinatorial. Prettier returns zero fit. The general small models drift. The frontier model without tools is the most expensive and the least reliable. The frontier model with a Python tool solves it — but reveals the limit: it isn't reasoning about constraints, it's writing a brute-force search script and running it. Same answer, different mechanism, and the cost reflects that.
The EBM + verifier wins on three axes at once: cost two orders of magnitude lower than the frontier, latency one order lower, and checkable output — the only candidate whose answer comes with a proof the answer is correct.
Same function. Same candidate set. Different winners — because the tasks are different shapes. The cost function doesn't care which mechanism wins. It cares which one fits.
Part 05 · What this replaces
Remove the model picker.
The mode buttons disappear. The "is this a Haiku question or an Opus question" anxiety disappears. The capability marketing — flash, pro, ultra, deep think — becomes background information the router uses, not foreground information the user has to interpret.
What the user does instead is what they always wanted to do: submit the work. The router reads the shape, runs the cost function against every candidate, picks the one that fits. The energy spend reflects the work, not the user's guess about what the work needs.
A 300M-parameter energy-optimized open model becomes a first-class citizen of the candidate set. So does Prettier. So does grep. So does a SQL query. The frontier model is still there, doing what only it can do. It is one of many tools, not the default tool.
Part 06 · What pricing in joules costs
The honest unit produces an honest floor.
When the receipt is real, a vendor can price the work in the unit the work actually costs — cost-of-compute in joules + a markup denominated in the same unit. Not tokens. Not seats. Joules. The price is verifiable against the receipt. The buyer's question — show me what this cost — has an answer for the first time.
A vendor pricing in joules has a floor competitors who can't measure their own waste cannot reach. They're pricing in proxies — tokens, seats, requests — that hide variance. A 10-token answer and a 10,000-token reasoning chain bill the same per-token but cost wildly different amounts. The token-priced vendor can't undercut a joule-priced vendor without knowing where their own waste is, and they can't know that without a receipt.
Procurement becomes a comparison of receipts rather than a comparison of marketing. Vendor A: 0.3 Wh per decision, V₄, no citation. Vendor B: 0.0003 Wh per decision, V₂, citation-anchored. Both correct. The buyer can verify both claims against the runtime receipts. There's nothing to discuss.
This isn't a strategy. It's what happens when the receipt exists and a buyer asks for it. The proxy-priced market becomes the unit-priced market the moment the unit becomes legible.
Energy-floor pricing isn't a moral position about efficiency. It's the natural pricing model that emerges when you can measure. The vendors who measure first set the floor; the vendors who can't measure compete above it. The receipt produces the market.
Continue
See it run.
The cost function names the choice. The receipt is what the choice leaves behind. Composition is what stages add up to. Verifiability is the gate. The cache is the slope.