The AI Iceberg

Margin Architecture: AI Economics Designed Into the Workflow

AI Value is set less by model pricing than by how a workflow handles context, execution, and quality.

September 2026  ·  21 min read  ·  The AI Iceberg


Many organizations start AI economics at the pricing page.

They negotiate model prices. They limit token budgets. They choose a cheaper model. They ask users to be more efficient with prompts. All of that can help. None of it changes the underlying system that determines whether AI gets more valuable or more expensive as usage grows.

Much of the unit economics of AI is decided inside the workflow.

Every AI-powered workflow makes a set of architectural decisions before it produces useful work:

Those are not implementation details. They are the architecture of value.

A strong AI workflow produces more useful work from each unit of compute and each minute of human attention. A weak one turns falling model prices into expanding token volume, brittle automation, and review queues that keep growing, and the bill grows with them.

This is Margin Architecture: a three-part operating system for designing AI economics into the workflow.

Context Foundation — what the system knows, retrieves, carries forward, and can reuse. Execution Design — how work is decomposed, routed, performed, and escalated. Quality Governance — how outcomes are verified, measured, improved, and promoted into trusted state.

The key insight is simple:

The economics of AI are designed into the workflow. AI Value is how you measure whether that architecture is working.

The workflow economic system. Context Foundation, covering context, state and retrieval, feeds Execution Design, covering routing and tools, which feeds Quality Governance, covering validation, escalation and learning. Quality Governance produces trusted workflow state, and that state returns to Context Foundation to improve the next run. The result is more useful work per unit of compute.
Quality Governance is the only stage that writes back. Verified outcomes become trusted workflow state, which returns to the Context Foundation as reusable input, so the next run carries less and routes better.

Defining AI Value

AI economics are often reduced to a token calculation:

AI cost=input tokens × input price  +  output tokens × output price

That is a billing equation, not a business equation.

The business measure is AI Value:

AI Value=Value of completed work  −  (Compute + Tooling + Human review + Rework + Failure cost)

AI Value measures whether a workflow produces an outcome worth more than the resources, controls, corrections, and consequences required to produce it.

Compute covers metered API spend and the cost of owned or reserved infrastructure alike. Failure cost is the expected consequence of work that passes through and turns out to be wrong. The expression works per completed task or totaled over a period. Divide it by the value of completed work and you get the workflow's AI margin, which is where this framework gets its name.

This changes the design objective.

The question is not, "How do we use the cheapest model?" The question is:

How do we create more AI Value at the required quality, reliability, security, and governance bar?

That distinction matters because the apparently expensive part of an AI system is often not the economically consequential part.

A frontier-model call may be visibly expensive. But a workflow that routes every case to human review, repeats the same long context on every turn, retrieves irrelevant documents, or generates outputs that must be corrected later can destroy value far faster than model pricing alone.

Take an illustrative contract-intake workflow handling 10,000 contracts a month. Decomposing it and routing each step to the cheapest tier that can do the work cuts monthly token spend from about $1,980 to about $810. That is a 59% reduction, and it is worth roughly $1,200 a month.

Sending only flagged cases to a human reviewer, rather than every output, cuts review cost from about $116,700 to about $17,500. That saving is roughly $99,000 a month, about 85 times the token saving.

The full parameters are in the worked example below, and the conclusion does not rest on them. Across 81 combinations of contract length, reviewer cost, review time and flag rate, the review saving exceeds the token saving by between 19 and 600 times. The ordering never reverses.

The point is not that tokens do not matter. It is that quality and escalation design determine whether token savings translate into AI Value.

Four terms worth separating

Term Definition Role in the operating model
Margin Architecture The three-part workflow system: Context Foundation, Execution Design, and Quality Governance The strategic design framework
AI Value Value of completed work less compute, tooling, human-review, rework, and failure costs The business measure
Cost per completed task The all-in cost of producing an outcome that meets its quality bar The core efficiency metric
Quality bar The standard an outcome must meet to count as completed work The constraint that makes the measurement meaningful

AI Value is not a standardized accounting measure. Each workflow must define the value of its completed work, what quality threshold applies, and which costs and failure consequences are material. The point is to make the economics explicit rather than treating token spend as a proxy for value.

Why lower prices do not automatically create AI Value

Model prices are falling. Capability is diffusing into smaller and cheaper models. Caching, batching, open-weight models, local inference, and model portfolios continue to improve the supply side of the equation.

But declining unit prices do not guarantee declining total spend, or greater AI Value.

AI workloads can expand in response to lower cost. Teams add users and workflows, stretch context windows, deploy more agents, and accept more retries and reasoning steps. A system can get cheaper per token and more expensive per completed task at the same time.

Economists know the pattern as the Jevons paradox: efficiency gains can raise total consumption instead of lowering it. Whether that happens here depends on how the workflow is built.

If the workflow is designed as… Then lower model prices tend to create…
A chat interface attached to old processes More usage, more tokens, more inconsistency
A frontier model as the default Commodity work performed at premium cost
An uncontrolled agent loop More retries, tool calls, and hidden failure modes
A governed operating system More completed work and more AI Value per unit of compute

The economic goal is not to suppress volume. It is to ensure that volume scales through lower-cost, reliable paths rather than through uncontrolled compute and human-review growth.

That is what Margin Architecture does.

1. Context Foundation

The first part of Margin Architecture is the Context Foundation:

Context, state, retrieval, and reusable inputs are designed as an economic layer, not assembled ad hoc at runtime.

Most AI systems treat context as a technical matter: concatenate a system prompt, chat history, retrieved documents, tool definitions, and a user request; then send it to the model.

That is convenient. It is also often expensive, slow, and unreliable.

Every token included in a request is something the system must process, pay for, and potentially expose across a model boundary. Every stale or irrelevant document increases cost without improving the answer. Every missing fact creates rework. Every unverified memory can create compounding error.

A better system distinguishes four kinds of information.

Context element What it contains Design objective
Cacheable foundation Stable policy, instructions, tool schemas, reusable knowledge, shared templates Keep stable and reusable
Workflow state Verified facts, prior decisions, completed steps, approved intermediate outputs Keep concise, durable, and trusted
Retrieval Current documents, records, policies, and evidence needed for this task Retrieve selectively and citeably
Dynamic request data The current user's request, case details, latest event, or changing inputs Introduce only when relevant

Stable context is an AI Value asset

Prompt caching exists because repeated computation is wasteful. When requests share an identical beginning, providers can reuse prior work rather than recomputing the full prefix. Microsoft's prompt-caching guidance explicitly ties identical beginning content to lower latency and cost, and recommends keeping stable content before dynamic content. Microsoft Learn

This is not merely a prompt-engineering trick. It is a workflow-design principle.

A high-volume workflow should make its durable elements explicit:

Then it should separate those stable elements from what changes by user, task, or moment.

The result is less repeated compute, lower latency, and a clearer boundary between reusable enterprise knowledge and dynamic case data.

The discount is material and varies by provider. Anthropic prices cache reads at a fraction of the standard input rate, a tenth for most current models and less for some, charges a premium to write a cache entry, and expires entries after five minutes or one hour depending on the option chosen. OpenAI documents discounts of up to 95% on cached input for some models, and newer models add a cache-write charge. Minimum prompt lengths, short lifetimes, and exact-prefix matching all apply, so a change near the start of a prompt can forfeit the saving. Cached tokens are discounted, not free. Check current pricing for the models you use before building a business case on it. (Anthropic, OpenAI)

State is not just memory

State is the compact record of what the workflow has already established.

It may include:

The important word is verified.

A system that retains every model output as durable memory is not building intelligence. It is accumulating ungoverned claims. A system that preserves only reliable, relevant workflow state reduces future work while protecting the quality of subsequent decisions.

Context should become more useful over time, not merely larger.

Retrieval should add evidence, not noise

Retrieval is valuable when it supplies the evidence the workflow needs now. It is counterproductive when it becomes a reflexive "search everything" step that adds cost, latency, and ambiguity.

Good retrieval design asks:

The goal is not maximum context. It is sufficient context for a reliable decision.

The economic test

For each context component, ask:

Does this information increase AI Value by reducing error, rework, or future compute by more than it costs to carry and process?

If the answer is no, it does not belong in the default context path.

2. Execution Design

The second part of Margin Architecture is Execution Design:

Every unit of work should follow the least costly path that can meet its required quality, latency, security, and reliability bar.

This is where decomposition, routing, tools, model selection, deterministic code, and human escalation belong.

A common failure in enterprise AI is to treat every request as if it were the same kind of work.

It is not.

A customer-message classification, a SQL lookup, a policy comparison, a research synthesis, a complex exception, and a regulated approval all have different requirements. They should not automatically use the same model, the same context window, the same tool chain, or the same level of human involvement.

The system should decompose a workflow into meaningful steps, then choose the right execution path for each one.

Routing is the execution decision

Routing asks:

Given the task, the available context, the required quality bar, the sensitivity of the data, and the expected cost, what should perform this next step?

The answer may be deterministic code. It may be a small local model. It may be an intranet-hosted model, a cloud model, a frontier reasoning system, an external tool, or a human reviewer.

The best execution path is not always the cheapest immediate path. It is the least costly path that produces a trustworthy completed result.

Execution path Appropriate work AI Value logic
Deterministic code Rules, validation, transformations, SQL, calculations Avoid model calls where software is more reliable
Small or local model Routine classification, extraction, simple drafting, privacy-sensitive tasks Keep low-risk work inexpensive and close to the data
Intranet model Internal knowledge work and repeatable enterprise tasks Capture owned-inference economics and data control
Mid-tier cloud model General reasoning, synthesis, common generation tasks Balance capability with variable cost
Frontier model Difficult reasoning, novel problems, high-value exceptions Spend premium capability only where it changes the outcome
Human reviewer Material risk, ambiguity, approvals, regulated or consequential exceptions Reserve scarce expert attention for cases that merit it

The organization's advantage does not come from guessing the "best" model once. It comes from owning the decision rules that match work to an execution path repeatedly.

Decompose before you optimize

Most workflows should not be routed as one undifferentiated request.

A contract-intake process, for example, may involve:

  1. Extracting fields from a document
  2. Checking them against known policy
  3. Identifying nonstandard clauses
  4. Comparing clauses against precedent
  5. Producing a summary
  6. Determining whether a reviewer must approve it

Those are different computational problems.

Treating all six steps as one frontier-model prompt is simple. It is not architecture.

Tools are part of the execution economy

Tools are often presented as a capability upgrade: the model can browse, query databases, trigger workflows, call APIs, or write records.

But each tool call also has an economic and reliability profile:

Tool design therefore needs the same discipline as model design.

Use stable interfaces. Limit tools to the tasks that need them. Carry forward validated tool results as state. Define permission boundaries. Instrument failure and retry rates. Do not let an agent call a tool repeatedly because the system failed to preserve a result it already obtained.

Human review is a routed tier

In many workflows, the largest cost is not compute. It is indiscriminate human oversight, as the modeled contract example above suggests.

There are two weak models:

The first destroys the automation business case. The second creates risk, error, and distrust.

A strong design treats human review as an execution tier, invoked when the workflow detects a meaningful exception:

This is where routing becomes a management system rather than a technical feature.

3. Quality Governance

The third part of Margin Architecture is Quality Governance:

Quality is not the last step. It is the control system that decides what can be trusted, what must escalate, and what the workflow learns from next.

Teams often frame quality as a tradeoff against efficiency. In practice, poor quality is often one of the most expensive forms of inefficiency.

A cheap answer that must be corrected is not cheap. A fast workflow that produces an unsupported recommendation is not fast. An autonomous tool call that triggers a recovery process is not automated in any economically meaningful sense.

Quality governance protects AI Value by preventing low-quality work from flowing downstream.

NIST's Generative AI Profile calls for documented test, evaluation, validation, and verification practices across the lifecycle (action IDs MP-2.1-002 and GV-1.5-003). The AI Risk Management Framework asks that human-oversight processes be defined, assessed, and documented (MAP 3.5) and that roles for human-AI oversight be assigned (GOVERN 3.2). NIST AI 600-1 · NIST AI RMF 1.0

Quality governance has three jobs

Assure

Determine whether the output meets the standard required for its task.

That may include:

The right quality bar varies by task. A low-stakes draft may need basic validation. A recommendation that affects pricing, eligibility, employment, credit, health, safety, or legal obligations needs a much higher bar.

Escalate

When the workflow cannot meet its quality bar, it should not pretend otherwise.

It should:

Escalation is not failure. It is a controlled response to uncertainty.

The expensive failure is not escalation. The expensive failure is sending low-confidence work into a downstream process that creates more rework, risk, or human intervention later.

Learn

This layer is often the least developed in AI architectures.

Quality governance should feed back into the workflow. Verified outcomes become durable state. Errors become evaluation cases. Escalation patterns identify weak retrieval, poor tool interfaces, insufficient policy context, or incorrect routing thresholds.

That is why the feedback loop from Quality Governance to State is central to the Margin Architecture model.

A workflow should not blindly store everything it produces. It should promote only verified outcomes into durable state.

Outcome What the system should do
High-confidence, validated result Preserve as trusted workflow state when useful
Successful human correction Add to evaluation data; consider updating policy, retrieval, or routing
Repeated model failure Investigate context, tool interface, decomposition, and model-fit assumptions
Uncertain or conflicting evidence Escalate or request clarification; do not promote to durable state
Tool failure or retry pattern Improve tool contract, permissions, observability, or fallback path

The economics are compounding. A system that learns from verified outcomes reduces repeat work. A system that repeats unverified outputs spreads costly uncertainty.

The feedback loop is the advantage

Margin Architecture is not a linear pipeline.

It is a loop:

Context Foundation→Execution Design→Quality Governance→Trusted State→Better Context Foundation

Each loop should improve one or more of the following:

This is how an AI system becomes economically better over time without becoming carelessly autonomous.

A worked example: contract intake

Consider a workflow that receives 10,000 contracts per month.

The naive design sends every document to a frontier model, asks it to interpret the contract, produces a summary, and sends every output to a legal reviewer.

That design may appear safe because a human sees every result. It is also expensive and difficult to scale.

Now apply Margin Architecture.

Context Foundation

Execution Design

Quality Governance

What the model assumes

Every figure quoted above comes from these parameters. They are illustrative, chosen to be ordinary rather than favourable. Substitute your own and the arithmetic still runs.

Parameter Value
Contracts per month 10,000
Contract length 30,000 tokens, roughly 22,000 words
Stable foundation: policy, clause taxonomy, output schema, tool definitions 12,000 tokens
Naive path Whole document plus the full foundation to a frontier model on every contract, nothing cached
Routed path Deterministic checks, then small-model extraction, mid-tier clause comparison on the 25% that need it, mid-tier summary, frontier model on the hardest 8%, foundation served from cache
Reviewer loaded cost $70 per hour
Review time per contract 10 minutes
Share reviewed, naive 100%
Share reviewed, routed 15%
Token prices Frontier $4 / $20, mid tier $2 / $10, small $1 / $5, cache reads $0.20, per million tokens

Prices are the published rates for Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5, checked on 30 September 2026. Model prices move; re-run the numbers rather than trusting these past their date.

What the model produces

Monthly cost Naive Routed Saved
Tokens $1,980 $812 $1,168 (59%)
Human review $116,667 $17,500 $99,167 (85%)
Total $118,647 $18,312 $100,335 (85%)

Two things are worth reading off that table.

The token column is the one everyone optimises, and it is 1.7% of the naive total. The review column is 98.3% of it. A team that negotiated a 20% discount on model pricing would save $396 a month. A team that changed who reviews what saved $99,167.

And the naive design is not obviously wrong. It routes everything to the best available model and puts a human in front of every output. It looks careful. That careful-looking version costs about six and a half times what the designed version costs, and almost none of the difference is tokens.

The aim is not merely lower token cost. It is a workflow that scales review capacity, protects quality, and holds or grows AI Value as volume rises.

That is the difference between adding AI to a contract process and designing an AI-native contract operation.

What to measure

A team cannot manage Margin Architecture using "AI adoption" or "tokens consumed" as its primary metric.

Those are activity measures. They do not tell you whether the workflow creates value.

The unit of measurement should be the AI Value of completed work at its required quality bar.

Context Foundation metrics

Execution Design metrics

Quality Governance metrics

These measurements transform the conversation.

Instead of asking, "Why did we use so many tokens?" leaders can ask:

Which part of the workflow is reducing AI Value, and what is the least costly reliable design change?

The practical design sequence

Margin Architecture does not require a fully autonomous multi-agent system. It begins with disciplined workflow design.

1. Choose one high-volume workflow

Start where there is enough repetition to learn:

Avoid starting with a broad, undefined "copilot" deployment. Choose a workflow with a clear beginning, end, owner, volume, quality bar, and cost baseline.

2. Map the real work

Document the inputs, the decisions, the handoffs, the sources of truth, the tools used, the exceptions, the human-review triggers, the outputs, the failure modes, and the cost and time of current execution.

This often reveals that the apparent "AI task" is actually several different tasks with different requirements.

3. Design the Context Foundation

Identify:

4. Route each workflow step

For every step, specify the cheapest viable execution tier, the required quality bar, the latency requirement, the data classification, the fallback path, the escalation condition, and the cost owner.

5. Define the quality system before scaling

Create:

6. Measure completed outcomes

Track AI Value, cost, latency, quality, escalation, rework, and human minutes per completed task.

Then improve the system where it matters:

The organizational implication

Margin Architecture is not owned by a prompt engineer or a procurement team alone.

It crosses product, platform, security, operations, finance, legal, and the business process owner.

That requires clear ownership.

Area Primary owner Core responsibility
Context Foundation Product and AI platform State model, retrieval design, data boundaries, reusable knowledge
Execution Design AI platform and process owner Decomposition, routing policy, tool interfaces, model portfolio
Quality Governance Process owner, risk, and evaluation function Quality bars, escalation policy, evaluation, review design
AI Value measurement AI FinOps and business owner Cost per completed task, value of completed work, budget, chargeback, value realization

The organization that owns these decisions owns the value. The organization that treats them as vendor defaults rents the economics from someone else.

The strategic choice

The AI market has been making individual model calls cheaper, and that may well continue.

That is not the same thing as creating more AI Value.

A company can use lower-priced models and still increase its AI bill, expand its review queues, weaken its controls, and create more work than it eliminates. Or it can use the same underlying models to build a system that gets more useful, faster, cheaper, and more reliable as volume grows.

The difference is not just the model.

It is the architecture.

Context Foundation makes information reusable, current, and trustworthy. Execution Design directs each unit of work to the right mix of code, models, tools, and people. Quality Governance protects the outcome and converts verified work into better state for the next cycle.

That is Margin Architecture.

The economics of AI are designed into the workflow. AI Value is how you measure whether the architecture is working.

Sources and notes

Provider caching prices were checked against the vendor documentation on September 30, 2026 and change often. Prompt-caching mechanics vary by provider and model family; the architectural principle is stable, in that reusable content should be separated from dynamic case data on purpose. State promotion requires security, privacy, data-retention, and model-risk controls appropriate to the workflow, and verified does not automatically mean permissible to retain. "Least costly" throughout means least costly while meeting the defined quality, reliability, latency, data-handling, and governance bar. AI Value is a decision framework rather than a standardized accounting ratio, so teams must define value, acceptable quality, rework, and failure cost for their own workflow.

Provider caching prices in this piece were checked against the vendor documentation on September 30, 2026. The contract-intake figures are modelled, not measured, and are labelled as such where they appear.

← All writing Not sure which thread is yours? Reading guide →