What if we could estimate the cost of a piece of work before doing it — and pick the right tool?

The capability of a model is irrelevant if it has nothing to do with the work. Every AI provider sells modes — fast, balanced, expert — but these tiers don't tell you which one fits the task in front of you. The choice the user is being asked to make is the wrong choice.

Estimate the cost. Read the shape of the work. Pick the right tool. Repeat.


Three providers. Three tiers each. One question to the user.

Each provider asks the user to pick a capability tier before the work runs.

Claude
Haiku fast · efficient
Sonnet balanced
Opus expert · expensive
Gemini
Flash fast · efficient
Pro balanced
Deep Think expert · expensive
DeepSeek
V3 fast · efficient
R1 reasoning
R1 high-think deep · expensive

The user does not know which mode fits the work. The model picker does not know either — because there isn't one. There is a UI for a question that should not have been asked of the user at all.


A model is the wrong unit of choice.

Capability is one number. Fit is many. A task has a shape — what kind of answer it needs, how much input, how strict the determinism, what counts as correct. Capability tells you how powerful the engine is. Fit tells you whether the engine is the right one for the road.

A frontier model running a closed-form transform is a Formula 1 car driving to the grocery store. It will do the job. The energy spent is absurd. The task requires correctness, which a 300M-parameter specialized model can also provide.

The honest question is not how powerful a model. It is which mechanism, at what cost, with what confidence of fit. The set of candidate mechanisms is broader than three buttons. It includes deterministic transforms. It includes retrieval. It includes specialized small models. It includes the frontier model. The question is which one this task wants.


Estimate before you run.

The point of a measurement framework is to know what something will cost before you do it. The cost function is the artifact that does this.

E ( task, candidate ) → cost, fit, confidence

Given a task and a candidate mechanism, return cost, fit, confidence

Inputs
  • task shape, modality, determinism, latency, reuse
  • candidate mechanism, hardware, cost per unit work
  • history past performance on similar work
Outputs
  • cost estimated joules per task
  • fit expected quality of output
  • confidence how sure the fit estimate is

This is the same E from the thesis — evaluated before the work runs instead of after. The thesis equation tells you what something cost. The cost function tells you what it will cost — low-cost enough to run on every candidate, every time.

Run the function for every candidate. Rank by fit and cost. The router picks. The user submits work. The user does not pick the model. The model picker has been removed from the user interface because it was never the user's question to answer.

The function itself is small. The router that uses it is small. Neither runs a large model. They are upstream of every decision about whether to run a large model at all.


Six shapes of work.

The function reads the task and identifies its shape. Six categories cover the work people ask computers to do.

Closed-form

The answer is fully specified by the input. Formatters, parsers, hash functions, lookups, deterministic transforms. The mechanism is software that already exists.

Constraint-bound

The rules are formal, but the answer space is combinatorial. Sudoku, formal proofs, scheduling, EDA, theorem search. The mechanism is search with a verifier — an energy-based model that scores candidate states against the constraints, plus a checker that certifies the result.

Pattern-bound

The answer is constrained by an authoritative source — a codebase, a regulation, a framework. The work is retrieval plus adaptation. The mechanism is a small model with retrieval.

Pattern-rich

Recognizable shapes, no single authoritative source. Drafting in a style, summarizing, paraphrasing, classifying. The mechanism is a specialized small model.

Open-ended

Genuinely novel synthesis. Architecture, design, complex reasoning. The mechanism is the large model.

Perceptual · temporal

Real-time sensing and control. Microsecond budgets. Embedded. The mechanism is deterministic preprocessing with a specialized small model where needed.

The mapping

These six shapes are the product-facing projection of the formal coordinate axes from the seven-axis stack. Closed-form is Z₁ direct evaluation. Constraint-bound is also Z₁ — same formal rules — but the answer space is combinatorial, so the mechanism is search-with-verifier instead of look-up. Pattern-bound is Z₂ with a retrieval anchor against an authoritative source. Pattern-rich is the Z₂↔Z₃ transition — recognizable shapes without an authoritative anchor. Open-ended is Z₃ (math ≠ words). Perceptual · temporal is what happens when the interface axis (signals, bodies) dominates the zone axis — microsecond budgets, embedded contexts, where the loop closes against the physical world.

The shape comes first. The mechanism follows. The mode buttons invert this — they ask the user to pick a mechanism without ever surfacing the shape. The cost function works the right way around.


Two tasks. One function. Different picks.

Two tasks of different shapes. The cost function runs against the same candidate set for both. The output is a ranked list — the router picks the top of each list, and the picks are different mechanisms.

Example A · Closed-form
format this JSON file
Candidate mechanism Cost Fit Latency Confidence
Prettier (deterministic) ~2 μJ 1.00 <1 ms high
Specialized small model · 300M (local) ~5 mJ 0.99 ~25 ms high
General small model · 1B (local) ~30 mJ 0.95 ~80 ms high
Mid model (Sonnet-class) ~1.5 J 0.99 ~600 ms high
Frontier model (Opus-class) ~2.8 J 0.99 ~1.2 s high
Frontier deep-think mode ~12 J 0.99 ~8 s over-fit

The router picks Prettier. The task is closed-form. The fit is perfect. The cost is six orders of magnitude lower than the frontier model. The output is identical.

Example B · Constraint-bound
solve this Sudoku puzzle
Candidate mechanism Cost Fit Latency Confidence
Prettier (deterministic) ~2 μJ 0.00 n/a does not apply
Specialized small model · 300M (local) ~10 mJ 0.40 ~50 ms medium
General small model · 1B (local) ~40 mJ 0.30 ~120 ms low
Frontier model (no tools) ~6 J 0.72 ~25 s low — often wrong
Frontier model (with Python tool) ~3 J 0.99 ~5 s high — via brute-force
Energy-based model + verifier ~80 mJ 0.99 ~0.4 s high — checkable

The router picks the energy-based model + verifier. The task is constraint-bound — the rules are formal, the answer space is combinatorial. Prettier returns zero fit. The general small models drift. The frontier model without tools is the most expensive and the least reliable. The frontier model with a Python tool solves it — but reveals the limit: it isn't reasoning about constraints, it's writing a brute-force search script and running it. Same answer, different mechanism, and the cost reflects that.

The EBM + verifier wins on three axes at once: cost two orders of magnitude lower than the frontier, latency one order lower, and checkable output — the only candidate whose answer comes with a proof the answer is correct.

Same function. Same candidate set. Different winners — because the tasks are different shapes. The cost function doesn't care which mechanism wins. It cares which one fits.


Remove the model picker.

The mode buttons disappear. The "is this a Haiku question or an Opus question" anxiety disappears. The capability marketing — flash, pro, ultra, deep think — becomes background information the router uses, not foreground information the user has to interpret.

What the user does instead is what they always wanted to do: submit the work. The router reads the shape, runs the cost function against every candidate, picks the one that fits. The energy spend reflects the work, not the user's guess about what the work needs.

A 300M-parameter energy-optimized open model becomes a first-class citizen of the candidate set. So does Prettier. So does grep. So does a SQL query. The frontier model is still there, doing what only it can do. It is one of many tools, not the default tool.


The honest unit produces an honest floor.

When the receipt is real, a vendor can price the work in the unit the work actually costs — cost-of-compute in joules + a markup denominated in the same unit. Not tokens. Not seats. Joules. The price is verifiable against the receipt. The buyer's question — show me what this cost — has an answer for the first time.

A vendor pricing in joules has a floor competitors who can't measure their own waste cannot reach. They're pricing in proxies — tokens, seats, requests — that hide variance. A 10-token answer and a 10,000-token reasoning chain bill the same per-token but cost wildly different amounts. The token-priced vendor can't undercut a joule-priced vendor without knowing where their own waste is, and they can't know that without a receipt.

Procurement becomes a comparison of receipts rather than a comparison of marketing. Vendor A: 0.3 Wh per decision, V₄, no citation. Vendor B: 0.0003 Wh per decision, V₂, citation-anchored. Both correct. The buyer can verify both claims against the runtime receipts. There's nothing to discuss.

This isn't a strategy. It's what happens when the receipt exists and a buyer asks for it. The proxy-priced market becomes the unit-priced market the moment the unit becomes legible.

Energy-floor pricing isn't a moral position about efficiency. It's the natural pricing model that emerges when you can measure. The vendors who measure first set the floor; the vendors who can't measure compete above it. The receipt produces the market.


See it run.

The cost function names the choice. The receipt is what the choice leaves behind. Composition is what stages add up to. Verifiability is the gate. The cache is the slope.