Skip to content
EssayAugust 2026

The Token Is Not the Unit

Companies can see what AI costs. Most cannot see what it produces, what that production is worth, or what happens to their workforce when they restrict it.

The whole argument in one sentence

Don’t judge AI by what it consumes. Judge it by the useful, checked work people and AI get done together.

Marcio Sete23 sections12,514 words63 min read

The essay in five minutes

Serious AI tooling comes with a serious bill. It is precise, it arrives monthly, and it is growing. So the instinct is to treat it the way finance treats any growing cost: cap it, default everyone to the cheaper option, and ask the heavy users to justify themselves.

That instinct answers the wrong question. This essay shows what to ask instead.

Start with one engineer. Say they cost A$200,000 a year, fully loaded. Now give them serious AI tooling for another A$60,000 a year. Finance sees a 30% cost increase, and Finance is right.

But suppose measurement shows that, working with AI, they now produce four times as much verified, integrated software as they did before. Their cost went up 30%. What they produce went up 300%. That creates roughly the same production capacity as three additional engineers — about A$600,000 of capacity — for A$60,000 of additional cost. On paper, ten to one, before discounts.

The essay calls the “four times” the Force Multiplier: how much more engineers working with AI — the essay calls them agentic engineers — produce against the pre-AI baseline. Everything in the essay exists to measure that number honestly, and to decide what it is worth.

Three rules make the measurement honest.

First, measure the result, not the consumption. Licences assigned, prompts sent and tokens burned tell you about adoption. None of them tell you whether more real software got built.

Second, count only work that survives. Code that fails review, or never reaches the product, is not production — however fast it was written. The quality gates are the admission test, not a discount applied afterwards.

Third, count conservatively. Extra capacity is not automatically extra cash. Some becomes more work, some better work, some faster work — and some evaporates in the queue for approvals. The essay discounts capacity honestly before asking Finance to believe it — and under those discounts, clearing the bar below takes a multiplier closer to seven than four.

Then hold AI spend to a hard test. Break-even is a surprisingly low bar — a 30% cost increase is repaid the moment the multiplier reaches 1.3 — but it is the wrong bar. The essay's rule is a 10:1 Compute Return Ratio: every extra dollar of AI spend must create at least ten dollars of additional engineering capacity, measured on the discounted numbers. Where a team proves that return, its budget should grow. Where it can't, the spend should shrink. The budget stops being a cap someone argues about and becomes a number the team earns.

Two ways to lose the benefit. A company can leave the new speed trapped in old structures — code produced in an hour queuing three days for approval, until the multiplier dies in the calendar. Or it can restrict capable AI so aggressively that its engineers never learn the new way of working. That is the essay's sharpest warning: a company can save on tokens today while quietly depreciating the capability of its engineering workforce for years.

What to do on Monday: establish the pre-AI baseline from history you already have; measure verified production against it for a fixed window, with rules frozen in advance; then move budgets — up where the return is proven, down where it is not.

The same argument as eight claims, each linked to the part of the essay that proves it:

  1. The AI bill answers the wrong question. It records spend; it cannot say whether the investment converted into more, better, faster work. Part I →
  2. One number carries the argument: the Force Multiplier — how much agentic engineers produce against the pre-AI baseline. Part II →
  3. Only work that survives the strictest quality gates and is integrated into the product counts as production. Part II →
  4. Break-even is not enough. The hurdle is a 10:1 Compute Return Ratio, risk-adjusted — and a rising AI bill can coexist with dramatically improving economics. Part III →
  5. Capping compute is not automatically financial discipline. Part V →
  6. The cheapest model can be the most expensive system — weak tools cost engineering time now and engineering skill later. Part IV →
  7. A 50× engineer inside a 1× operating model does not create a 50× organisation. Part IV →
  8. Saving on AI today can make your workforce less capable tomorrow. Engineers need enough exposure to learn the new production model, not just enough AI to make the old one easier. Part IV →

The thesis

The token is a metering unit. It is not an economic unit.

A friend working at one of Australia’s major banks sent me a message:

“People have started discovering that Claude isn’t free. From now on, Sonnet will be the default for everyone, and Opus will only be available to ‘power users’.” 😜

It is funny because it is true.

The AI bill has arrived.

Now comes the entirely predictable response.

  1. 01Use the cheaper model by default.
  2. 02Restrict access to the expensive model.
  3. 03Introduce per-person limits.
  4. 04Ask why one engineer consumed A$5,000 of inference in a month.
  5. 05Create a category called “power users” and make access to the frontier model an exception requiring approval.

From a conventional cost-control perspective, that response is perfectly rational.

It may also be catastrophically wrong.

Not because Claude is free.

It is not.

Not because every engineer should have unlimited access to the most expensive model.

They should not.

And not because AI expenditure should escape financial scrutiny.

It absolutely should not.

The problem is more fundamental.

Most organisations can see the cost of AI, but they cannot yet see the production function it is creating.

The token bill arrives before the force multiplier becomes visible. Leadership sees a new cost centre. Engineers see rapidly increasing capability. Finance sees weak attribution. The resulting proof-of-value gap naturally produces token caps, cheaper-model defaults and pressure to control consumption before the organisation has built a credible way to measure what that consumption returns.

That is the category error at the centre of enterprise AI economics:

The token is a metering unit. It is not an economic unit.

The economically relevant question is not:

How many tokens did this person consume?

It is:

How much real, tested software did this engineer deliver with AI? What did it cost to produce, including the engineer and the AI costs combined? And did the additional AI spend buy enough more, better or faster software to justify the additional cost?

This essay proposes a different way of thinking about the economics of agentic engineering: measure the cost of AI and the leverage it creates together, not separately.

Today, most organisations can measure the first far more easily than the second. Tokens, licences and model costs arrive neatly on an invoice. The resulting change in engineering capacity, quality and speed is much harder to see.

The argument that follows connects four things organisations often manage separately:

  1. The cost of models, tokens, tools and platforms.
  2. The probability that AI-assisted work produces an acceptable result.
  3. The change in the engineering production function.
  4. The economic value of the capacity created.

The argument is deliberately demanding.

It does not ask Finance to believe AI vendors.

It does not ask executives to accept self-reported productivity.

It does not confuse activity with output.

It does not treat raw lines of code as value.

It does not assume that more software is automatically better.

It does not claim that a 10× engineering force multiplier means 10× business value.

And it does not ask anyone to write a blank cheque.

Instead, it proposes a hard financial hurdle:

For every A$1 of incremental AI compute, require at least A$10 of risk-adjusted additional production capacity.

That is a 10:1 Compute Return Ratio.

At that hurdle, Finance can be extraordinarily conservative without accidentally destroying the upside.

The economics in six lines

Token cost tells you what large language models consumed.

Product Surface tells you what verified engineering production emerged.

Force Multiplier tells you how far the production function moved.

α, the economic realisation factor, tells you how much of that additional capacity the organisation can actually capture.

The 10:1 Compute Return Ratio tells Finance how much compute can rationally be funded.

Organisation design and workforce capability determine whether the multiplier compounds into enterprise value or dies at the individual contributor.

Tokens → accepted outcomes → Product Surface → Force Multiplier → organisational capture → economic value

Everything else in this essay exists to make that chain measurable, auditable and governable. Tokens are useful telemetry. The economic question is what happens after them.

But first, AI needs its equivalent of horsepower.

Part I · for everyone

The idea

Section 01AI needs horsepower

Rory Sutherland tells a powerful story about the commercialisation of the steam engine.

An invention, he argues, does not become an innovation merely because it works. It becomes an innovation when it changes behaviour. A technically impressive machine that nobody adopts remains an invention.

Mine owners did not want to hear about boiler capacity, piston length or mechanical specifications.

They wanted to know something much simpler:

How many horses will I no longer need to feed?

James Watt translated the unfamiliar capabilities of the steam engine into a unit that buyers could connect to an existing cost.

That unit was horsepower.

A mine owner could estimate the cost of acquiring, feeding, housing and maintaining horses. Once the engine was expressed in horsepower, the buyer could perform the economic calculation on the back of an envelope.

Horsepower was not primarily an engineering unit created to impress other engineers.

It was a translation unit.

It connected technical capacity to commercial consequence.

In the same conversation, the hosts make the connection to AI directly. Modern AI products present people with model families, version numbers, context windows and technical benchmarks, but do not give buyers an intuitive equivalent of horsepower. They contrast that with Apple’s much more comprehensible promise of “a thousand songs in your pocket”.

Enterprise AI now has the same problem the steam engine had.

A CFO does not fundamentally care that one model consumed 100 million input tokens, generated 12 million output tokens, scored a particular percentage on a coding benchmark or used adaptive reasoning.

Those facts may matter operationally.

They do not answer the investment question.

The CFO wants to know:

What existing production cost, constraint or delay does this intelligence remove?

The modern equivalent of horsepower is the Agentic Engineering Force Multiplier.

The full translation is:

Steam economyAgentic engineering economy
Boiler, piston and coalModel, tokens and inference
Mechanical outputVerified Product Surface
HorsepowerAgentic Engineering Force Multiplier
Horses no longer requiredAdditional engineering capacity released
Cost of feeding horsesCost of conventional engineering production
Share of coal savingsEarned compute envelope tied to return

The components have distinct roles:

Product Surface is the production unit.

Force Multiplier is AI’s horsepower.

Compute Return Ratio is the CFO’s purchase rule.

The earned compute envelope is the governance mechanism.

This is the central economic logic of the essay.

Section 02A token is fuel, not production

At the time of writing, Anthropic lists Claude Sonnet 5 at US$2 per million input tokens and US$10 per million output tokens. Claude Opus 5 is listed at US$5 per million input tokens and US$25 per million output tokens.

On the price sheet, Opus is therefore 2.5 times as expensive per token.

That appears to make model selection easy.

Use Sonnet unless someone can justify Opus.

But price per token is several layers removed from economic cost.

A token is not a standard unit of intelligence, effort, text, correctness or value.

Different models can use different tokenisers. Anthropic notes, for example, that Sonnet 5’s newer tokeniser produces approximately 30% more tokens for the same text than Sonnet 4.6, meaning the change in effective request cost is not proportional to the published change in per-token price.

Input and output tokens are priced differently.

Cached input is priced differently again.

Cache writes, batch processing, fast service tiers, long contexts, reasoning, retries and tool loops all change the real cost.

Prompt caching can materially reduce the cost of repeatedly loading stable context, and batch processing can change the economics again.

Even inside one model, two tasks consuming the same number of tokens may have completely different economic consequences.

One may produce a correct, integrated capability on the first attempt.

The other may produce plausible-looking output that consumes three hours of human repair, introduces a defect and never reaches production.

The invoice can be identical.

The economics are not.

This gives us the first economic rule:

Never confuse the price of inference with the cost of an accepted outcome.

The total cost of an AI-assisted task is closer to:

Total task cost =
  inference cost
+ human steering cost
+ context preparation cost
+ verification cost
+ retry and rework cost
+ failure remediation cost
+ cost of delay

That is why the more expensive model can be the cheaper production resource.

Suppose one model costs three times as much per run but:

  • completes the task in one attempt rather than four;
  • requires five minutes of human steering rather than forty;
  • produces tests that pass;
  • understands the architecture;
  • avoids a review cycle;
  • reduces the probability of a production defect;
  • finishes a critical task a day earlier.

The expensive model may have the lower cost per accepted outcome.

Therefore:

The right default is not the cheapest model. It is the cheapest model that reliably clears the acceptance criteria for the task.

Sonnet as a default can be sensible.

Sonnet as a doctrine is not.

Opus as a controlled escalation can be sensible.

Opus as a prestige benefit for senior “power users” is not.

The choice should be determined by task economics, not organisational status.

Section 03The four layers of agentic engineering economics

Most organisations stop at the first layer of measurement.

They count tokens.

A useful economic view needs four layers.

LayerPrimary questionUseful measure
MeteringWhat did machine intelligence cost?Input, output, cache, tool, platform and model costs
ExecutionHow did the agentic work behave?Run cost, elapsed time, retries, failures and human interventions, used diagnostically within comparable task classes
ProductionDid the engineering production function change?Product Surface per effective engineer coding day, Force Multiplier
EnterpriseDid the new capacity create or avoid economic value?Cost avoidance, capacity release, acceleration, risk reduction, revenue and option value

Each layer is necessary.

None is sufficient by itself.

Token cost without execution outcome is consumption telemetry.

Execution success without production measurement may describe isolated wins.

Production multiplication without product judgement can create more software pointed at the wrong problems.

Business outcomes without a production measure are multi-causal and difficult to attribute honestly.

This is why efficiency, effectiveness and value must remain separate.

Section 04Efficiency, effectiveness and value are different truths

The three questions are:

Efficiency

Can the same amount of human engineering capacity produce more verified software?

Effectiveness

Did leadership direct that extra capacity towards the right customer problems, product bets and technical investments?

Value

Did the resulting capability create revenue, avoid real cost, reduce risk, improve customer outcomes or advance strategy?

The Agentic Engineering Force Multiplier lives in the first layer.

It measures efficiency.

It does not claim to prove effectiveness or business value.

That separation is not a weakness.

It is what keeps the argument intellectually honest.

This makes the production-capacity case deliberately conservative: improvements in quality, reliability, UX, security or delivery speed can create additional value beyond the Force Multiplier itself.

A 50× production system aimed at the wrong work can produce 50× more waste.

A 2× system aimed at a critical commercial bottleneck may create extraordinary value.

Agentic engineering can multiply the rate at which intent becomes verified software.

It cannot guarantee that the intent was correct.

The pathway is:

Prove the multiplier → unlock new engineering capacity → direct it through established product disciplines → expand the range of economically solvable problems → create and capture value.

This also explains why it is usually impossible to calculate “AI business value” directly from a feature outcome.

A feature’s result depends on:

  • customer need;
  • product strategy;
  • design;
  • distribution;
  • pricing;
  • timing;
  • sales;
  • adoption;
  • operations;
  • change management;
  • competitive response.

Trying to isolate exactly how much revenue was caused by Claude usually produces false precision.

The narrower claim is more credible:

Did the augmented engineering system materially change the amount of verified production created per unit of human capacity?

That question can be measured.

You now understand the central argument. Continue only if you need the measurement, the mathematics, the implementation model or the objections.

Part II · for people who need evidence

The proof

Section 05The one metric worth proving

The wrong executive metrics are not always useless.

They are merely being asked to answer questions they cannot answer.

Common AI dashboards include:

  • licences assigned;
  • monthly active users;
  • prompts submitted;
  • tokens consumed;
  • suggestions shown;
  • suggestions accepted;
  • AI-generated lines accepted;
  • pull requests created.

These measures can help diagnose adoption, training, configuration, workflow constraints and vendor utilisation.

They do not prove that the engineering production function changed.

The primary proof metric is:

Positive net Product Surface per effective engineer coding day, relative to the relevant historical baseline.

In plain English:

How much real, tested product surface can one engineer land in a coding day compared with the pre-agentic baseline?

Product Surface is:

Selected real product and test source code that becomes part of the product after deterministic exclusions and the required quality gates.

It excludes obvious sources of fake volume:

  • generated files;
  • lock files;
  • dependency and vendored code;
  • build output;
  • bundles and source maps;
  • configuration;
  • documentation;
  • images and binaries;
  • caches;
  • tooling artefacts.

It is net, not gross:

Net Product Surface =
  selected lines added
- selected lines removed

Net measurement matters because gross additions reward verbosity, rewrites, churn and code movement.

A 50,000-line rewrite that removes 49,900 lines has created 100 net lines of surface, not 50,000.

Negative days do not erase positive production elsewhere. They record zero Product Surface produced and retain deleted surface separately as surface retired.

The denominator is not naive headcount. It uses effective contributors so that a day where one person produced almost all the work and two people made tiny changes is not treated as three equal engineer-days.

The Force Multiplier is:

Force Multiplier =
  Product Surface per effective engineer coding day (agentic engineering)
  ---------------------------------------------------------------------
  historical baseline Product Surface per engineer coding day (pre-AI)

The detailed measurement system uses the Git record, aggregates at the engineer-day level, applies deterministic exclusions, calculates effective contributors, compares current production with a frozen historical baseline and subjects current output to strict quality admission.

In a proprietary analysis across more than 5,000 enterprise repositories, I observed an initial anchor of approximately 120 selected net lines per engineer coding day, with a broader range of roughly 100 to 200.

That is not a universal constant.

Each organisation should calculate its own benchmark from its own engineering record.

Against a 120-line benchmark:

Product Surface per effective engineer coding dayForce Multiplier
1201×
6005×
1,20010×
6,00050×
12,000100×

The claim is not that every line contains equal value.

The claim is that if the same organisation, measured through the same processor and exclusions, moves from 120 to 1,200 selected, net, quality-gated lines per effective engineer coding day, its engineering production function has changed.

That is a narrower claim than “AI generated business value”.

It is also more falsifiable.

Section 06Quality is admission, not an adjustment factor

Raw lines of code are a terrible productivity metric.

They have historically rewarded exactly the wrong behaviours:

  • unnecessary verbosity;
  • copy and paste;
  • churn;
  • artificial decomposition;
  • boilerplate;
  • generated artefacts;
  • “busy” engineering.

Product Surface does not make raw lines of code respectable.

It defines a different object.

Code does not count because a model generated it.

Code counts because the engineering system verified and integrated it.

Quality is therefore not inferred from volume.

It is a condition of admission.

A production-grade system can include:

  • engineer-approved intent and plans;
  • explicit acceptance criteria;
  • architectural and dependency constraints;
  • strict linting and type checking;
  • secret scanning;
  • dependency and vulnerability scanning;
  • architecture enforcement;
  • minimum code coverage;
  • duplication thresholds;
  • static application security testing;
  • semantic analysis such as CodeQL;
  • infrastructure-as-code scanning;
  • licence controls;
  • integration tests;
  • performance tests;
  • container scanning;
  • mutation testing;
  • reliability and regression suites.

The current agentic system should be held to a stronger, more automated bar than the historical system.

The thesis is not:

AI can produce more code.

The thesis is:

An agentically augmented engineering system can produce orders of magnitude more verified Product Surface while increasing the degree of machine-enforced quality.

The detailed framework treats quality as a defence-in-depth admission system, combining human control of intent and architecture with automated verification of implementation.

At high production rates, this is not optional.

A human review system designed for 1× output cannot linearly absorb 50× output.

If every additional line requires the same human review effort as before, the local speedup simply moves the bottleneck downstream.

The operating model must change with the production function.

Section 07What the evidence does and does not say

Public research on AI coding productivity remains mixed.

DORA’s 2025 research describes AI as an amplifier. It magnifies the strengths of capable organisations and the dysfunctions of struggling ones. Individual coding gains can disappear into downstream disorder in testing, review, security and deployment.

METR’s early-2025 randomised study found that experienced open-source developers working on familiar repositories took 19% longer when using the AI tools available at the time, despite believing that AI had made them faster. METR’s 2026 update found signals moving towards possible speedups, but with broad uncertainty and important selection effects.

These findings are not contradictions to be resolved with a slogan.

They tell us that “Does AI make engineers faster?” is too broad a question.

The answer depends on:

  • model capability;
  • task type;
  • repository familiarity;
  • harness design;
  • context quality;
  • human skill;
  • autonomy;
  • latency;
  • quality automation;
  • architecture;
  • team design;
  • downstream constraints;
  • the measurement unit.

Every organisation must measure its own production system.

In my own work, I have measured weekly Force Multipliers of 84.1× and 92.9×.

The 92.9× week included approximately 69,645 lines of net functional TypeScript integrated over five coding days, 163 commits, 15 database migrations and 31 releases.

It was predominantly net-new product capability, not a refactor.

The output passed an engineering system that included zero-warning linting, strict type checking, architecture enforcement, changed-file coverage thresholds, duplication limits, full tests and builds, secret and dependency scanning, vulnerability checks, static analysis and deeper scheduled verification.

Those numbers do not mean:

  • 92.9× business value;
  • one engineer should replace 92.9 engineers;
  • every organisation will reproduce the result;
  • every week will sustain the same multiplier;
  • all types of engineering are represented equally.

They mean that, in the measured work class and operating context, the production ratio was too large to dismiss as ordinary variance.

The published 84.1× and 92.9× calculations used the active baseline of approximately 150 selected net functional lines per engineer-day at the time.

Make the denominator explicit. Pre-register it. Freeze it for the proof window. Version it. Never change it retrospectively to improve the result.

The objective is not to declare a universal multiplier.

It is to build a credible system in which a multiplier can be proved, challenged, audited and either accepted or rejected.

Part III · for finance leaders

The mathematics

Section 08The CFO mathematics

Now we can translate engineering horsepower into financial terms.

Let:

 = annual fully loaded human engineering cost

 = annual incremental AI cost
    including inference, licences and allocated platform cost

 = verified Agentic Engineering Force Multiplier

For simplicity, assume:

 = A$200,000 per year

Using salary alone would understate the true organisational cost. A mature model should include superannuation, leave, equipment, management, recruitment, facilities, insurance and other relevant overheads.

The augmented engineer costs:

Augmented cost =  + 

The relative cost per unit of Product Surface is:

Relative unit cost =
       + 
  -------------
       × 

                  1 + ( / )
                = -----------
                       

The first form compares what the engineer + AI system costs, H + A, with what the historical human-only system would have cost to produce the same amount of software, H × M. A result of 0.26, for example, means the augmented system costs 26% as much to produce the same amount of Product Surface. The simplified form expresses the same ratio differently. The 1 represents the engineer's original cost before any AI spend is added.

The pure financial break-even multiplier is:

Break-even multiplier =
  1 + ( / )

For a A$200,000 engineer:

Monthly AI spendAnnual AI spendIncrease over human costPure break-even multiplier
A$1,000A$12,0006%1.06×
A$5,000A$60,00030%1.30×
A$10,000A$120,00060%1.60×
A$15,000A$180,00090%1.90×

This is the first important surprise.

An engineer consuming A$5,000 per month does not need to become 5× or 10× as productive merely to break even.

They need to move from 1× to 1.3×.

If they produce 5× Product Surface at a 30% cost uplift, the relative unit cost becomes:

1.30 / 5 = 0.26

The organisation is producing each unit of Product Surface at approximately 26% of the previous cost.

That is roughly a 74% reduction in unit production cost.

At 10×:

1.30 / 10 = 0.13

The unit cost is approximately 87% lower.

The total engineering spend has increased.

The cost per unit of production has collapsed.

This distinction is fundamental:

A rising AI bill can coexist with dramatically improving economics.

A CFO who watches only the total token bill will see cost growth.

A CFO who watches cost per verified unit of production may see an industrial transformation.

The economic question is therefore not whether AI makes engineering spend go up. It is whether, after adding the AI cost, each unit of real, tested software becomes cheaper to produce.

Section 09Break-even is not enough

Finance should not fund AI merely because it crosses break-even.

The hurdle should be much higher.

There are two different financial tests here, and they answer different questions.

Break-even is a unit-economics test. It asks whether, after adding the AI cost, each unit of verified software production is cheaper than it was under the historical human-only production model.

The 10:1 hurdle is an incremental investment test. It asks a harder question: does the additional money spent on AI create enough additional economic capacity to make that incremental investment compelling?

Let us require:

A$10 of additional capacity-equivalent value for every A$1 of incremental AI expenditure.

In conventional finance language, this resembles a 10:1 gross benefit-cost hurdle: ten dollars of capacity-equivalent benefit for every dollar of incremental cost.

If the full A$10 were realised as cash, that would correspond to a 900% net ROI: A$9 of net benefit on A$1 invested. But that is deliberately not the claim being made here. At this stage, the numerator is capacity-equivalent value, not necessarily realised cash savings. Later, the economic realisation factor α will apply a further haircut to that theoretical capacity before Finance recognises it economically.

I use the term Compute Return Ratio rather than simply ROI because conventional ROI is often calculated as net benefit divided by cost.

That is a severe hurdle.

Good.

Let:

 = required Compute Return Ratio

For this article:

 = 10

The gross additional capacity-equivalent value is:

Gross additional capacity =
   × ( - 1)

Before applying the economic realisation factor introduced later, the gross Compute Return Ratio is:

Gross Compute Return Ratio =
   × ( - 1)
  -----------
       

The denominator is incremental AI spend, so the numerator must also represent incremental benefit above the original human-only baseline. That is why the formula uses M - 1 rather than M.

Before any economic-realisation haircut, the maximum annual AI spend supported by the 10:1 hurdle is:

Maximum annual AI spend =
   × ( - 1)
  -----------
       10

The multiplier required for a given AI budget is:

Required multiplier =
  1 + (10 ×  / )

Assuming full recognition of the additional capacity at this stage, the thresholds for a A$200,000 engineer are:

Monthly AI computeAnnual AI costPure break-evenMinimum multiplier for 10:1
A$1,000A$12,0001.06×1.6×
A$5,000A$60,0001.30×4.0×
A$10,000A$120,0001.60×7.0×
A$15,000A$180,0001.90×10.0×

These are the gross thresholds. Section 12 introduces α, allowing Finance to recognise only a proportion of the theoretical capacity. If α is below 1, the multiplier required to preserve a 10:1 risk-adjusted return rises accordingly.

Now we have the basic capital-allocation rule.

For a single engineer with a A$200,000 annual cost basis, consuming A$5,000 per month in AI, the gross 10:1 hurdle requires a 4× Force Multiplier.

At 4×:

Additional capacity equivalent =
  A$200,000 × (4 - 1)
= A$600,000

Annual AI cost =
  A$60,000

Gross Compute Return Ratio =
  A$600,000 / A$60,000
= 10:1

This assumes α = 1, meaning Finance recognises the full capacity-equivalent value. The later risk adjustment deliberately tests what happens when it does not.

At A$10,000 per month, the required multiplier rises to 7×.

At A$15,000 per month, it rises to 10×.

These thresholds are specific to the A$200,000 human cost basis. For a pod, team or other unit, use the total human engineering cost and total AI spend for that same unit.

Break-even asks: After adding the AI cost to the engineer’s cost, does each unit of verified production cost no more than it did before?

10:1 asks: Does each additional dollar of AI spend support at least ten dollars of additional capacity-equivalent value?

This is not permissive.

It is not “AI at any cost”.

It is an extraordinarily demanding return requirement.

Yet it still shows why arbitrary A$1,000 or A$5,000 token caps can be economically irrational.

Section 10What a 10× engineer is worth

At a 10× Force Multiplier and a A$200,000 cost basis:

Gross additional capacity equivalent =
  A$200,000 × (10 - 1)
= A$1,800,000 per year

At a 10:1 Compute Return Ratio:

Maximum annual AI spend =
  A$1,800,000 / 10
= A$180,000

Maximum monthly AI spend =
  A$15,000

This does not mean the engineer should deliberately spend A$15,000 per month.

That would confuse an economic ceiling with a spending target.

It means:

Preventing a demonstrably 10× production system from consuming A$6,000 instead of A$5,000 because it exceeded an arbitrary monthly cap is not necessarily financial discipline.

It may be local optimisation.

The business could be protecting A$12,000 of annual budget while constraining hundreds of thousands of dollars of additional production capacity.

At 84.1×, the unadjusted 10:1 calculation produces a theoretical compute envelope of approximately A$138,500 per month.

At 92.9×, approximately A$153,000 per month.

Those are not proposed budgets.

They are diagnostic numbers.

They expose the scale mismatch between the inference cost being debated and the production ratio being claimed.

When a result produces an apparently absurd economic envelope, the responsible reaction is:

  1. Audit the baseline.
  2. Audit the exclusions.
  3. Audit the denominator.
  4. Audit the quality gates.
  5. Audit whether the result persists.
  6. Audit whether the organisation can absorb the output.

If the multiplier fails scrutiny, reject it.

If it survives scrutiny, stop managing the system as though it were an ordinary software licence.

Section 11Capacity equivalent is not automatically cash saving

This distinction is essential.

The formula:

 × ( - 1)

calculates additional engineering capacity equivalent.

It does not automatically calculate cash saved.

Cash cost avoidance occurs only when the organisation avoids expenditure it would otherwise have incurred, such as:

  • planned hires;
  • contractors;
  • outsourced delivery;
  • programme extensions;
  • additional teams;
  • duplicated platforms;
  • future support costs.

If the organisation keeps the same people and spends the released capacity on additional work, the result is not immediate payroll saving.

It is capacity release.

That capacity may still be extraordinarily valuable, but it must be described honestly.

There are at least five distinct value pools.

1. Cost avoidance

Planned expenditure is not incurred.

2. Capacity release

The existing team produces more without proportional headcount growth.

3. Acceleration value

Capabilities reach customers or operations earlier.

A simple expression is:

Acceleration value =
  cost of delay per day
× days brought forward

4. Risk reduction

Stronger tests, faster remediation, improved reliability, better controls or reduced operational exposure create value.

5. Option value

The organisation can run more experiments and solve long-tail problems that were previously too small, expensive or slow to justify.

These should not all be added together indiscriminately.

A credible investment case should:

  • choose a conservative primary value pool;
  • identify additional value separately;
  • avoid double counting;
  • disclose assumptions;
  • distinguish capacity from cash;
  • reconcile forecasts with realised outcomes over time.

The biggest long-term effect may not be doing the old backlog faster.

It may be changing which work is economically possible.

When the cost of producing verified software falls by an order of magnitude, problems that once affected too few users to justify a team can become viable.

A new integration for one important customer.

An internal workflow used by twelve people.

A modernisation task that always lost against feature work.

A support tool for a narrow operational process.

A new channel that could never earn its place in the old portfolio.

The feasible problem set expands.

Recent METR work makes a related distinction between uplift on old tasks, uplift on new tasks and uplift in value. When the cost of completing some tasks collapses, people substitute towards different tasks, so productivity measured only against the old task portfolio may miss part of the change.

This is why capacity-equivalent production matters even when it does not immediately remove a salary line.

Section 12A risk-adjusted realisation factor

Finance may reasonably reject the assumption that every dollar of additional engineering capacity will be absorbed and converted into useful output.

The economic model can accommodate that conservatism.

Let:

 = economic realisation factor

The factor represents the proportion of additional capacity that Finance is willing to recognise for budget purposes after considering:

  • absorption constraints;
  • product effectiveness;
  • portfolio quality;
  • adoption;
  • substitution;
  • value capture;
  • measurement uncertainty.

Then:

Risk-adjusted capacity value =
   ×  × ( - 1)

The Compute Return Ratio becomes:

Compute Return Ratio =
   ×  × ( - 1)
  -----------------
          

The maximum AI budget at a 10:1 hurdle becomes:

Maximum annual AI spend =
   ×  × ( - 1)
  -----------------
          10

At 10×, a A$200,000 engineer creates A$1.8m of theoretical additional capacity. The question is how much of that Finance is prepared to recognise.

Economic realisation factorRisk-adjusted additional capacityMaximum annual AI spend at 10:1Monthly envelope
25%A$450,000A$45,000A$3,750
50%A$900,000A$90,000A$7,500
75%A$1,350,000A$135,000A$11,250
100%A$1,800,000A$180,000A$15,000

This is an excellent mechanism for CFOs who want a margin of safety.

Even if Finance recognises only 25% of the capacity-equivalent result, a genuinely 10× system still supports A$3,750 per month of AI spend at a 10:1 hurdle.

At 50% realisation, it supports A$7,500.

The organisation does not need to agree that every unit of Product Surface becomes cash.

It only needs to agree on a conservative realisation factor and then apply it consistently.

The most important determinant of α may therefore not be model capability at all. It may be whether the organisation is designed to capture the capacity the model releases.

Part IV · for engineering and people leaders

The organisation and the workforce

Section 13The multiplier must escape the individual contributor

There is a failure mode appearing across enterprise AI that is easy to mistake for success.

The engineer is dramatically more productive.

The organisation is not.

Recent conversations with several engineers working at large software companies crystallised it for me. They described developers using coding agents to start a piece of work, let the agent run, step away from the computer, spend time with their family, go for a walk, do something around the house, come back later, inspect the result, provide feedback and send the agent around the loop again.

I have heard versions of this story repeatedly.

And to be clear, this is a real benefit.

If an engineer can produce the same outcome with less cognitive load, more flexibility, better focus and a better life, that has value. There is nothing inherently wrong with an employee capturing some of the productivity dividend.

But it is not the same thing as an organisational force multiplier.

If the engineer previously needed eight hours of concentrated work to complete a task and can now direct an agent for one hour while achieving the same result, their personal labour efficiency may have improved enormously.

But if the organisation still ships the same number of outcomes, on the same cadence, through the same queues, with the same team structure, then the enterprise production function may have barely moved.

The employee has experienced the multiplier.

The organisation has not captured it.

That distinction may explain one of the great puzzles of enterprise AI adoption so far: organisations can simultaneously have thousands of employees saying AI makes them more productive and executives struggling to find corresponding economic results.

This is not merely anecdotal. DORA’s recent research has documented precisely this tension. Individual developers report productivity and wellbeing improvements, while gains can be swallowed by downstream constraints in review, testing, security and deployment.

The implication is profound:

A 50× engineer inside a 1× operating model does not create a 50× organisation.

The individual contributor trap

Most enterprise AI programmes begin by sprinkling licences over an operating model designed for human production.

The unit of organisation does not change.

Work still arrives as tickets.

Capacity is still planned around humans.

Projects are still estimated against human implementation rates.

Teams are still shaped around specialised roles and handoffs.

Progress is still discussed through sprint commitments.

Pull requests still queue for human review.

Testing may still require another team.

Security may still arrive as a gate.

Deployments may still wait for release windows.

Governance may still assume that every consequential action needs synchronous human approval.

Managers may still allocate approximately the same amount of work because the planning system has no concept of machine-created capacity.

So the engineer accelerates dramatically inside their part of the value stream and then encounters the next system boundary.

The agent produces in minutes.

The ticket waits until tomorrow.

The implementation takes an hour.

The review waits six.

The code is ready on Tuesday.

The release train leaves Friday.

The engineer can manage five concurrent agentic workstreams.

The team is still allocated one ticket at a time.

At that point, machine speed terminates at the next human-shaped queue.

This is classic local optimisation expressed through a new technology. Making one stage of a system dramatically faster does not create equivalent system throughput when the constraint simply moves somewhere else.

The organisation has made the engineer faster without making the system of work faster.

This is multiplier leakage

Think of the force multiplier as capacity entering a pipeline.

Some of it reaches the organisation.

Some of it leaks away.

It can leak into waiting.

It can leak into larger batches.

It can leak into review queues.

It can leak into meetings.

It can leak into dependency management.

It can leak into approval processes.

It can leak into work that was already committed and therefore cannot expand.

And yes, some of it can become personal time returned to the engineer.

That last category should not automatically be described as waste. Improved wellbeing, reduced burnout and greater flexibility are legitimate benefits.

But Finance should distinguish them from enterprise production economics.

If the organisation pays for AI intending to create additional delivery capacity, while the operating model has no mechanism for absorbing that capacity, it should not be surprised when the financial return is difficult to see.

This is where the economic realisation factor, α, becomes more than a conservative finance assumption.

It becomes partly an organisational design variable.

A company with an extraordinary local multiplier and a poor ability to absorb it may have a low α.

Another company using exactly the same model with an operating system designed around agentic production may convert a much larger proportion of the multiplier into shorter lead times, more Product Surface, fewer external hires, smaller teams, more experiments or more ambitious products.

Same model.

Same token price.

Completely different economics.

Sprinkling licences is not transformation

A licence changes what an individual can do.

An operating model changes what an organisation can do.

Those are different interventions.

The first era of enterprise AI was largely a distribution exercise: buy licences, make tools available, train people, encourage adoption and measure usage.

That was reasonable when AI behaved primarily as autocomplete and chat.

Agentic engineering is different.

An agent can inspect a repository, plan a change, modify dozens of files, create tests, execute tools, diagnose failures, call other agents, iterate against acceptance criteria and continue working while the human is elsewhere.

That is no longer simply a faster typing tool.

It is a new production mechanism.

And a new production mechanism eventually demands a compatible system around it.

Sprinkling agent licences across a human-shaped operating model gives you AI-assisted humans. It does not give you an agentic organisation.

The belief system has to change first

Operating-model change begins with a change in what leadership believes is possible.

If leaders fundamentally believe software is still produced by humans typing code, with AI helping around the edges, every management mechanism will continue to be designed around the scarcity of human implementation effort.

Team sizes will assume it.

Planning will assume it.

Estimation will assume it.

Funding will assume it.

Role definitions will assume it.

Review processes will assume it.

Governance will assume it.

Portfolio economics will assume it.

The first change is therefore conceptual:

Human typing is no longer necessarily the rate-limiting step in software production.

Once that becomes true, many mechanisms designed around that constraint need to be reconsidered.

The human role does not disappear.

It moves.

Humans increasingly own intent, architecture, context, judgement, taste, trade-offs, risk and acceptance.

Agents increasingly perform implementation, investigation, mechanical verification, repetitive transformation and longer-running execution loops.

The transition is not from human coding to AI coding.

It is from human-led production to human-directed machine production.

That is a different operating model.

You cannot review 50× output with a 1× review system

This is one of the clearest examples.

If agents can produce 10×, 50× or 100× more implementation surface, asking humans to compensate by manually reviewing 10×, 50× or 100× more code is not governance.

It is a bottleneck.

The quality model has to evolve from predominantly human inspection towards engineered verification.

Humans should continue to own the high-level things machines cannot safely decide for the organisation: intent, architecture, experience, important trade-offs, material risk and final accountability.

But implementation properties that can be expressed as policy should increasingly be machine enforced: types, tests, coverage, mutation, architectural boundaries, security, dependency policy, licence policy, performance constraints and integration behaviour.

The point is not to trust the agent more.

It is to make trust less necessary.

The same principle applies beyond review.

If agents can work asynchronously, work allocation has to tolerate asynchronous execution.

If one engineer can orchestrate several workstreams, team design has to acknowledge orchestration capacity.

If implementation becomes dramatically cheaper, portfolio thresholds have to change.

If agents can test continuously, testing cannot remain primarily a downstream phase.

If governance rules are deterministic, they should increasingly become policy-as-code rather than meetings.

If production becomes continuous, a release system designed around scarce human implementation may become the constraint.

The organisation has to be redesigned around the new bottleneck, not the old one.

This is not an argument for squeezing engineers harder

There is an ugly interpretation of all of this that should be rejected explicitly.

The answer to an engineer saving four hours with AI is not to install surveillance software and make sure Finance extracts four more hours of keyboard activity.

That would preserve the old mental model while making the workplace worse.

The objective is not utilisation.

It is throughput, learning and value.

An agentic operating model should allow humans to spend less time on mechanical production and more time on the work where humans remain scarce: understanding customers, framing problems, making architectural decisions, designing experiences, exercising judgement, exploring possibilities and directing increasingly capable machines.

Some of the productivity dividend should manifest as lower cognitive load and better working lives.

Some should manifest as additional organisational capacity.

Leadership needs to decide deliberately how that dividend is shared.

What should not happen is for the organisation to pay for the new production mechanism, leave the old production system untouched, observe little movement at the enterprise level and conclude that the models did not create leverage.

The real transformation question

The question is no longer:

How do we get more developers using AI?

It is:

How do we redesign the organisation so that locally created machine leverage can flow all the way to customer and business outcomes?

That requires a different belief system, different engineering controls, different team structures, different work allocation, different planning assumptions, different review mechanisms, different platform capabilities and, eventually, different portfolio economics.

A licence is access.

Access can create individual leverage.

Individual leverage can create local capacity.

But only a compatible operating model turns that capacity into an organisational force multiplier.

That gives us the progression that matters:

Model capability → individual leverage → quality-admitted production → organisational throughput → economic realisation.

If the chain breaks at the individual contributor, the employee may still have a better day.

That is valuable.

But the CFO will still see a token bill.

The job of agentic organisation design is to make sure the multiplier does not stop there.

Section 14The hidden cost of cheap AI is human-capital depreciation

There is another cost missing from most enterprise AI budgets.

It does not appear on the Anthropic invoice.

It does not appear in the Copilot licence report.

It may not become visible in the P&L for several years.

It is human-capital depreciation.

AI access policy is not merely a technology procurement decision.

It is increasingly a workforce development decision.

An organisation deciding which models, harnesses and levels of autonomy its engineers can access is also influencing what kind of engineers those people will become.

That matters because software engineering is currently moving through several different production paradigms at once.

Three engineers can now have fundamentally different jobs

Consider three engineers with the same title.

The first is still primarily a human-first engineer.

They understand the problem, design the solution and manually implement most of it. AI may occasionally help with search, explanation or small fragments of code, but human typing remains the primary production mechanism.

The second is an AI-assisted engineer.

They use autocomplete, chat and perhaps constrained coding assistance. They type less. They receive suggestions. They ask AI to explain unfamiliar code, generate tests or scaffold implementations.

But the underlying operating model remains largely unchanged.

Humans still decompose the work.

Humans still produce most implementation incrementally.

Humans still move ticket by ticket.

Humans still inspect the result through processes designed around human-authored code.

Then there is the agentic engineer.

Their primary skills increasingly sit one level above implementation:

intent → context → architecture → constraints → orchestration → verification → acceptance

They routinely delegate coherent bodies of implementation.

They let agents inspect repositories, use tools, execute tests, diagnose failures and iterate.

They learn when to intervene and when not to.

They run multiple workstreams.

They create reusable skills, context and automation.

They restructure repositories and delivery systems so agents can operate safely.

They design acceptance criteria that machines can verify.

They spend progressively less of their scarce human cognition mechanically producing code and more of it directing a software production system.

These are not simply three productivity levels.

They are increasingly three different professional paradigms.

The dangerous middle

The greatest workforce risk may be the second category.

An engineer can receive enough AI assistance to reduce how often they practise some traditional implementation skills, while receiving too little model capability, autonomy, harness sophistication or organisational permission to develop the skills of the agentic paradigm.

That creates an uncomfortable transition state.

The old craft begins to atrophy.

The new craft does not fully develop.

Autocomplete makes the engineer type less, but does not necessarily teach them orchestration.

Chat answers questions, but does not necessarily teach them how to design long-horizon autonomous work.

A constrained assistant may produce fragments, but never expose the engineer to what becomes possible when an agent can own a coherent outcome, use tools, recover from errors and continue working over a long horizon.

The engineer becomes more dependent on AI while remaining inside the mental model that preceded it.

That is not transformation.

It can become professional limbo.

The most dangerous AI environment may be one powerful enough to erode the old skill set, but too constrained to force development of the new one.

This tension is beginning to appear in research as well. DORA’s 2026 study of computing students found them recognising AI literacy as an essential capability for their future while simultaneously worrying that over-reliance could erode their foundational skills — concerns DORA notes mirror those of professional developers. For this argument, the answer is not retreating wholesale to manual coding. It is ensuring that reduced implementation practice is replaced by deeper capability in reasoning, verification, architecture and orchestration rather than by passive dependence on AI.

The new paradigm does not make software engineering fundamentals less important. It makes them the control plane: the engineer may author less implementation directly, but must become better at architecture, correctness, failure modes, security, verification and judgement.

You cannot learn the paradigm from a presentation

Agentic engineering is experiential.

Someone can explain context engineering.

They can demonstrate Claude Code.

They can present slides about agents, subagents, skills, MCPs, tools and long-running workflows.

That is useful.

It is not equivalent to hundreds of hours of operating the system yourself.

The important knowledge is tacit.

What should I delegate?

What should remain human?

How large can the task be?

What context does the agent actually need?

When should I interrupt it?

When should I let it continue?

How do I recognise that it is confidently heading in the wrong direction?

How do I structure acceptance criteria?

How do I make the repository legible to machines?

How do I turn an architectural rule into an executable constraint?

How do I run three or five streams of work without becoming the bottleneck myself?

How do I change my own role once implementation is no longer the scarce resource?

Those capabilities develop through repeated exposure.

And they compound.

The frontier is creating a career bifurcation

This creates a second-order labour-market effect.

Some engineers are waiting for their employers to provide the environment in which they will learn the future of their profession.

Others are not.

The latter group is disproportionately entrepreneurial.

They buy their own frontier-model subscriptions.

They use powerful coding agents on personal projects.

They build an application, an open-source project, an automation, a tiny product or an idea that may never become commercially successful.

The success of the side project is almost beside the point.

The project gives them something more valuable:

unrestricted practice in the emerging production model.

They experience the failures.

They learn the limits.

They redesign around them.

They accumulate reusable patterns.

They build intuition that cannot be acquired through annual training modules.

Now imagine two equally talented engineers spending the next three years in different environments.

One works inside a company where AI means constrained autocomplete and chat, frontier models require exceptional approval, agentic harnesses are restricted and the operating model remains built around human implementation.

The other spends those three years routinely orchestrating capable agents and building real systems end to end.

Both may eventually have “three more years of experience”.

Those years may no longer represent remotely equivalent professional development.

That is a career issue for the engineer.

It is also a strategic workforce issue for the enterprise.

An organisation that systematically prevents its engineers from learning the frontier may save on tokens while depreciating the capability of its engineering workforce.

The saving appears now. The cost appears later.

This is particularly dangerous for Finance because the two sides of the equation arrive at different times.

Restricting model access produces an immediate and measurable saving.

The organisation can report that inference expenditure fell.

Human-capital depreciation arrives slowly.

It appears later as:

  • weaker internal agentic capability;
  • dependence on a small number of frontier practitioners;
  • greater reliance on external specialists;
  • slower adoption of new operating models;
  • difficulty attracting engineers who want frontier experience;
  • difficulty retaining those who have developed it;
  • a widening gap between internal practices and external technical capability;
  • lower future Force Multipliers.

The saving is visible.

The opportunity cost is not.

That asymmetry is exactly the kind of situation in which local optimisation becomes dangerous.

Access should therefore have a learning objective as well as a production objective

This does not mean every engineer requires unlimited access to the most expensive model.

It means model policy should recognise two legitimate reasons to consume compute.

The first is production.

Production compute should be governed against outcomes, Product Surface and the 10:1 Compute Return Ratio described in this essay.

The second is capability development.

Some compute should deliberately fund experimentation and professional learning before a 10:1 production return can reasonably be expected.

That belongs inside the discovery envelope.

The organisation is not merely discovering whether a model works.

It is discovering what its people can become when they learn to operate it.

A company that demands demonstrated agentic leverage before allowing engineers enough access to learn agentic engineering has created a circular policy:

prove you can produce the multiplier before we give you the environment in which you can learn to produce the multiplier.

That is not prudent governance.

It is a capability deadlock.

There is an individual responsibility too

Engineers should not interpret any of this as permission to put employer source code, confidential information or proprietary data into personal AI accounts contrary to company policy.

The boundary matters.

But another boundary matters too.

Your employer’s tooling policy should not become the ceiling of your professional development.

An engineer can learn on personal projects, open-source software, synthetic environments and their own products.

The tools are increasingly accessible.

The responsibility for a career ultimately remains with the person whose career it is.

Enterprises have a responsibility to create environments in which their people can develop.

Engineers have a responsibility not to outsource their entire professional future to the procurement decisions of their current employer.

At scale, this becomes more than a career issue.

It becomes a societal one.

If access to capable AI systems determines who gets enough practical exposure to cross from AI-assisted work into agentic work, access to compute increasingly becomes access to career capital.

That should concern employers, educators and policymakers as much as it concerns engineers.

The transition is already under way.

The organisations that understand it will not merely buy better AI.

They will develop a different generation of engineers.

And the engineers who understand it will not wait for someone else to tell them when their profession has changed.

Section 15The expensive model may be the cheaper system

The phrase “Opus is more expensive” is incomplete.

More expensive per what?

Per million tokens?

Per run?

Per completed task?

Per accepted pull request?

Per production release?

Per unit of Product Surface?

Per business outcome?

These are not equivalent.

The correct comparison is:

Expected cost per accepted outcome =
  inference
+ human steering
+ verification
+ retry and rework
+ failure risk
+ delay

A lower-capability model may create hidden costs through:

  • more prompting;
  • more context reconstruction;
  • repeated failures;
  • smaller task horizons;
  • more human intervention;
  • more fragmented changes;
  • review fatigue;
  • shallow tests;
  • architectural inconsistency;
  • incomplete execution;
  • higher defect risk.

A more capable model may cost more per token but require:

  • fewer attempts;
  • less steering;
  • less repair;
  • less context repetition;
  • fewer handovers;
  • fewer review cycles;
  • less elapsed time;
  • less human attention.

Therefore:

The model with the lowest procurement price may not be the model with the lowest production cost.

Nor is current production cost the only denominator. Model policy also influences the future capability of the people using it. A cheaper model can therefore be locally cheaper while contributing to a more expensive workforce capability problem over time.

The governing question should be:

Which approved model, harness and effort level produces the lowest risk-adjusted cost per accepted outcome for this class of work?

A model that cannot legally or safely process the relevant data is not economically eligible, regardless of its theoretical return.

Security, privacy, sovereignty, regulatory and contractual constraints remain hard boundaries.

Economics operates inside those boundaries.

Part V · for implementers

The governance

Section 16“Power user” is the wrong category

A “power user” is an entitlement category.

It tells us that someone:

  • uses the tool frequently;
  • has seniority;
  • is considered sophisticated;
  • has received permission;
  • may consume more.

It says nothing about return on capital.

A better category is:

Power producer

A power producer is not someone who enjoys using the premium model.

It is an engineer, pod, workflow or operating pattern that demonstrably converts machine intelligence into disproportionately more verified production.

The distinction changes the question.

Power user framing:

Why is this person allowed to consume so much?

Power producer framing:

What verified return is this production system creating, and what compute should be allocated to it?

Rory Sutherland gives several examples of small wording changes altering behaviour. In one American Express example, changing the framing from “apply for your card” to what someone needed to do to “receive your card” reduced the psychological fear of rejection.

Language changes what people attend to.

Consider the difference:

Consumption languageProduction language
Power userPower producer
Token allowanceEarned compute envelope
AI usageAugmented production investment
Expensive modelHigher-capability production resource
Cost per tokenCost per accepted outcome
AI-generated codeQuality-admitted Product Surface
Productivity claimMeasured Force Multiplier
Unlimited accessReturn-governed compute

This cannot be empty relabelling.

The production language must be supported by a ledger, a baseline and evidence.

But when the evidence exists, the language should represent the economics accurately.

One caution matters:

Do not turn “power producer” into an individual employee leaderboard.

The Force Multiplier is an operating-model metric.

It should not become a quota for lines per person or a ranking of human worth.

The measured unit may be:

  • a team;
  • a pod;
  • a product area;
  • a workflow;
  • a repository family;
  • a defined agentic cohort.

Use the smallest unit with enough signal, but never convert the metric into surveillance theatre.

Section 17The earned compute envelope

The logical alternative to an arbitrary token cap is an earned compute envelope.

The envelope expands or contracts according to demonstrated leverage.

It should have three parts.

1. Discovery envelope

A bounded, time-limited budget for exploring:

  • models;
  • harnesses;
  • agent patterns;
  • context strategies;
  • skills;
  • tools;
  • task classes;
  • quality controls.

Not every experiment must individually return 10:1.

The discovery portfolio exists to find the high-leverage operating patterns.

It also funds capability development. Engineers cannot be expected to demonstrate frontier-level leverage before they have had enough frontier exposure to learn the operating model that produces it.

Production compute earns its envelope through demonstrated return. Discovery compute creates the conditions under which future return can be discovered and future capability can develop.

2. Earned production envelope

Once a team or workflow demonstrates sustained production leverage, its available compute expands according to the 10:1 formula.

Earned compute envelope =
   ×  × ( - 1)
  -----------------
          10

The envelope is not a spending target.

It is the maximum amount that still preserves the required return.

3. Critical-path exception

Some work should be governed by cost of delay rather than normal monthly limits.

Examples include:

  • material incidents;
  • security remediation;
  • regulatory deadlines;
  • critical bids;
  • production outages;
  • time-sensitive launches;
  • customer commitments.

For these tasks, the value of an hour or a day may dwarf the token cost.

The decision should still be recorded and reviewed, but forcing a critical task through a low-cost model because an individual exhausted a monthly allowance can be irrational.

The envelope should also be elastic.

If the multiplier falls, the earned envelope falls.

If quality deteriorates, output stops being admitted.

If the active benchmark changes, the measurement must be rerun consistently.

If the multiplier rises and survives audit, the envelope rises.

This is governance through demonstrated economics rather than status.

Section 18Token anxiety is range anxiety with a cloud invoice

Rory tells a story about driving an electric car whose battery display showed 16%.

The remaining absolute range was similar to what he regularly tolerated in another vehicle without concern.

But 16% felt alarming.

The physical situation was not materially different.

The presentation changed the psychological response.

Enterprise AI has the same problem.

A dashboard shows:

A$5,000 of Claude usage this month.

The number is isolated.

It feels uncontrolled.

Now display:

A$5,000 of AI compute produced A$75,000 of risk-adjusted additional engineering capacity, after quality admission, for a 15:1 Compute Return Ratio.

The invoice has not changed.

The decision information has improved.

This is not an argument for psychological manipulation.

It is an argument for completing the denominator.

Token anxiety is rational when production is invisible.

The answer is not to hide the cost.

It is to place the cost next to the return.

Tokens, prompts, active users and suggestions can remain available as diagnostic telemetry.

They should not be the headline.

A useful test for any executive metric is:

If this number doubled tomorrow, what material decision would we make differently?

If token consumption doubles, Finance investigates.

If Product Surface rises from 1× to 10× while cost per surface collapses and quality remains green, leadership may reconsider:

  • model caps;
  • team size;
  • hiring;
  • contractor dependence;
  • portfolio scope;
  • programme duration;
  • product economics;
  • operating model;
  • funding.

That is the difference between telemetry and a production metric.

Section 19Reverse benchmarking for AI

Rory calls one of his methods reverse benchmarking.

Instead of studying everything competitors already measure and trying to become marginally better, find the important dimension the entire category has neglected.

Then become disproportionately good at it.

Enterprise AI is currently focused on the obvious dimensions:

  • model price;
  • benchmark scores;
  • token volume;
  • context window;
  • active users;
  • adoption;
  • suggestion acceptance;
  • latency.

These measures are available and familiar.

The neglected metric is harder:

Cost per unit of verified production.

Anyone can read the Anthropic invoice.

It is much harder to:

  • define Product Surface;
  • inspect Git history;
  • exclude non-product artefacts;
  • normalise by effective engineer-days;
  • establish a valid baseline;
  • freeze the denominator;
  • version the processor;
  • enforce strict quality gates;
  • link inference spend to accepted outcomes;
  • separate efficiency from value;
  • maintain the ledger continuously.

That difficulty is precisely why the metric matters.

The best enterprise opportunities often sit in dimensions competitors do not measure because the measurement is inconvenient.

Force Multiplier is reverse benchmarking applied to AI production.

Section 20Measurement cadence

The metric should be continuously produced but not impulsively managed.

A practical cadence is:

CadenceWhat it covers
DailyCapture source facts, identify anomalies and confirm quality status.
Weekly or fortnightlyReview where the operating model is producing leverage and where agents are getting stuck.
MonthlyReview the sustained Force Multiplier, total augmented cost, cost per Product Surface and earned compute envelope.
QuarterlyDecide on model access, budgets, caps, team structures, hiring, contractor strategy, portfolio ambition, platform investment and operating-model redesign.

Do not make executive decisions from one exceptional day.

Use rolling windows with enough coding-day signal.

Do not change the baseline every time the current system improves.

The baseline answers:

What was normal before?

The current system answers:

What are we producing now?

The multiplier answers:

How far did the production function move?

Part VI · for sceptics — and CFOs

The stress test and the decision

Section 21The strongest objections

The objections to this argument are important.

They should not be dismissed.

They should shape the measurement design.

“Lines of code are a terrible metric.”

Raw lines of code are a terrible productivity metric.

Product Surface is selected, net, quality-admitted and integrated software surface measured at the operating-system level.

The claim is narrow.

If the same organisation moves from approximately 120 to 1,200 selected net lines per effective engineer-day under the same processing rules and stronger quality gates, the production system has changed.

The metric does not rank engineers or claim that every line has equal value.

“More code is not better.”

Correct.

If a problem can be solved clearly in 100 lines, solving it in 1,000 is waste.

That is why the measure is net, excludes generated artefacts, constrains duplication and maintainability, and operates over broader windows rather than judging one change.

The opportunity is not to make one feature 50 times larger.

It is to make 50 times more economically viable work possible.

“AI code is verbose.”

It can be.

Human code also varies by engineer, language, framework and style.

A 10% or 20% verbosity difference cannot explain a 10× or 100× shift.

Unnecessary verbosity should be constrained by architecture rules, duplication limits, maintainability gates, tests and human judgement.

“This will create AI slop.”

It will if the organisation measures generation instead of admission.

The model does not count what the model produced.

It counts what the engineering system accepted.

If low-quality output passes every configured gate, the quality system is not strict enough.

The answer is stronger admission, not pretending that unmeasured output does not exist.

“The result does not prove causality.”

Correct.

The hypothesis concerns the combined operating model:

  • models;
  • harnesses;
  • humans;
  • workflows;
  • context;
  • quality automation;
  • architecture;
  • organisation design.

The objective is not to isolate the causal contribution of every component.

It is to determine whether the current production system materially outperforms the historical one.

“Refactoring and deletion can be valuable.”

Correct.

Positive net Product Surface can undercount a major refactor, simplification or decommissioning.

That is why surface retired is recorded separately and transformation outcomes remain part of operational context.

The framework accepts this limitation rather than rewarding gross churn.

“Not every kind of engineering becomes code.”

Correct.

Product Surface is strongest where capability is expressed in application and test source.

Configuration-heavy platforms, data engineering, infrastructure, model development and other work may require separately defined surface families and relevant historical benchmarks.

Do not force unlike artefacts into one metric.

Use the same principle:

Define a stable production surface, process it consistently and compare it with the appropriate historical record.

“The metric can be gamed.”

Any metric can be gamed when converted into an individual target.

Safeguards include:

  • net rather than gross;
  • deterministic exclusions;
  • day-level aggregation;
  • quality admission;
  • effective contributors;
  • frozen benchmarks;
  • minimum signal thresholds;
  • coding-day-weighted roll-ups;
  • processor versioning;
  • separate retirement accounting;
  • no individual quotas.

The metric should evaluate the operating model.

It should never determine which human is worth more.

“A 10:1 hurdle is arbitrary.”

It is a policy choice, not a law of nature.

That is intentional.

Finance can choose 5:1, 10:1, 15:1 or another ratio according to risk appetite.

The economic model makes the choice explicit.

A 10:1 ratio is useful because it is demanding enough to create a large margin of safety while still revealing how economically small inference cost can become at genuine multipliers.

“Not everyone will produce 84×.”

Correct.

They do not need to.

At A$5,000 per month and a A$200,000 cost basis, the 10:1 threshold is 4×.

At A$1,000 per month, it is 1.6×.

The objective is not universal 90× performance.

It is compute allocation proportional to demonstrated leverage.

“This sounds like headcount reduction.”

It can support cost avoidance, but that is not the only or necessarily the best use.

The capacity may be directed towards:

  • more products;
  • faster modernisation;
  • improved security;
  • stronger tests;
  • new integrations;
  • smaller programme teams;
  • shorter delivery windows;
  • reduced contractor growth;
  • more experimentation;
  • previously uneconomic long-tail problems.

Workforce decisions remain human, strategic and social decisions.

The Force Multiplier describes production capacity.

It does not dictate what leadership must do with it.

“Regulated organisations cannot just use any model.”

Correct.

Security, privacy, data residency, record keeping, explainability, legal, model-risk and operational controls are not optional.

The earned compute envelope operates only across approved models, data classes, deployment patterns and use cases.

A high economic return does not override a regulatory boundary.

Section 22The policy in one equation

The entire financial policy can be expressed as:

Permitted annual AI spend =
  economic realisation factor
× annual human engineering cost
× (verified Force Multiplier - 1)
÷ required Compute Return Ratio

Or:

Amax = ( - 1) / 

Where:

Amax = maximum economically supported annual AI spend

    = economic realisation factor

    = fully loaded annual human engineering cost

    = verified Force Multiplier

    = required Compute Return Ratio

For this article:

 = 10

The matching access rule is:

Allow the cheapest approved model that reliably clears the task’s acceptance criteria. Escalate automatically when a more capable model lowers the risk-adjusted cost per accepted outcome. Expand compute only where the measured production system continues to clear the 10:1 hurdle.

That is what it means to govern AI cost and engineering leverage together.

Not free tokens.

Not unlimited inference.

Not executive enthusiasm.

Not procurement by sticker price.

Measured intelligence under an aggressive return constraint.

Section 23A letter to Australian CFOs

Dear CFOs,

You are right to challenge the AI bill.

You are right to ask whether thousands of licences are being used.

You are right to ask why one engineer consumed A$5,000 in a month.

You are right to demand evidence when someone claims a 10×, 50× or 90× multiplier.

Audit the baseline.

Audit the exclusions.

Audit the denominator.

Audit the quality gates.

Audit the cost allocation.

Audit persistence.

Audit whether the organisation can absorb the capacity.

Audit whether it becomes value.

But please do not stop at the invoice.

The token is not the unit.

A token tells you how a vendor meters the machine.

It does not tell you how much engineering production the machine created.

It does not tell you how much human time it released.

It does not tell you how many retries it avoided.

It does not tell you whether the output passed security, testing and architecture controls.

It does not tell you whether the work reached production.

It does not tell you whether a critical release arrived three months earlier.

It does not tell you whether a A$5,000 inference bill avoided a A$500,000 delivery cost.

Claude is not free.

Opus is not free.

Sonnet is not free.

Neither are engineers.

Neither is allowing the capability of your engineering workforce to depreciate because the token policy optimised this quarter’s invoice at the expense of next year’s production system.

Neither are contractors.

Neither are delivery teams.

Neither are handovers.

Neither are review queues.

Neither are delayed programmes.

Neither are missed market windows.

Neither are defects.

Neither is waiting six months to learn what the customer could have told you next week.

So keep the hard financial hurdle.

Make it 10:1.

For every dollar of incremental AI compute, demand ten dollars of risk-adjusted additional production capacity.

If the system cannot prove it, constrain the spend.

If it can prove it, do not optimise the token while destroying the multiplier.

Do not ask who deserves access to Opus because they are a “power user”.

Ask which approved model, task and operating pattern produces the lowest cost per verified outcome.

Do not ask only whether AI expenditure increased.

Ask whether cost per unit of production fell.

Do not ask the model how productive the engineer feels.

Read the engineering record.

Do not assume 90×.

Make the system prove it.

But once it does, act on the evidence.

Force Multiplier is AI’s horsepower.

Product Surface is the instrument that measures it.

The 10:1 Compute Return Ratio is the purchase rule.

The earned compute envelope is the governance mechanism.

The token is fuel.

Verified production is the unit.

Written by

Marcio Sete is the Head of AI Engineering and Platforms at Kinetic IT, a founding engineer and AI builder with 25+ years shipping products and teams. He writes and mentors on agentic engineering and engineering leadership.