Many organizations start AI economics at the pricing page.
They negotiate model prices. They limit token budgets. They choose a cheaper model. They ask users to be more efficient with prompts. All of that can help. None of it changes the underlying system that determines whether AI gets more valuable or more expensive as usage grows.
Much of the unit economics of AI is decided inside the workflow.
Every AI-powered workflow makes a set of architectural decisions before it produces useful work:
- What information does the system carry forward?
- What evidence does it retrieve?
- What instructions, tools, and knowledge repeat from task to task?
- Which model, code path, tool, or human should handle each step?
- What happens when the answer is uncertain or wrong?
- Which verified outcomes should improve the next run?
Those are not implementation details. They are the architecture of value.
A strong AI workflow produces more useful work from each unit of compute and each minute of human attention. A weak one turns falling model prices into expanding token volume, brittle automation, and review queues that keep growing, and the bill grows with them.
This is Margin Architecture: a three-part operating system for designing AI economics into the workflow.
Context Foundation — what the system knows, retrieves, carries forward, and can reuse. Execution Design — how work is decomposed, routed, performed, and escalated. Quality Governance — how outcomes are verified, measured, improved, and promoted into trusted state.
The key insight is simple:
The economics of AI are designed into the workflow. AI Value is how you measure whether that architecture is working.
Defining AI Value
AI economics are often reduced to a token calculation:
That is a billing equation, not a business equation.
The business measure is AI Value:
AI Value measures whether a workflow produces an outcome worth more than the resources, controls, corrections, and consequences required to produce it.
Compute covers metered API spend and the cost of owned or reserved infrastructure alike. Failure cost is the expected consequence of work that passes through and turns out to be wrong. The expression works per completed task or totaled over a period. Divide it by the value of completed work and you get the workflow's AI margin, which is where this framework gets its name.
This changes the design objective.
The question is not, "How do we use the cheapest model?" The question is:
How do we create more AI Value at the required quality, reliability, security, and governance bar?
That distinction matters because the apparently expensive part of an AI system is often not the economically consequential part.
A frontier-model call may be visibly expensive. But a workflow that routes every case to human review, repeats the same long context on every turn, retrieves irrelevant documents, or generates outputs that must be corrected later can destroy value far faster than model pricing alone.
Take an illustrative contract-intake workflow handling 10,000 contracts a month. Decomposing it and routing each step to the cheapest tier that can do the work cuts monthly token spend from about $1,980 to about $810. That is a 59% reduction, and it is worth roughly $1,200 a month.
Sending only flagged cases to a human reviewer, rather than every output, cuts review cost from about $116,700 to about $17,500. That saving is roughly $99,000 a month, about 85 times the token saving.
The full parameters are in the worked example below, and the conclusion does not rest on them. Across 81 combinations of contract length, reviewer cost, review time and flag rate, the review saving exceeds the token saving by between 19 and 600 times. The ordering never reverses.
The point is not that tokens do not matter. It is that quality and escalation design determine whether token savings translate into AI Value.
Four terms worth separating
| Term | Definition | Role in the operating model |
|---|---|---|
| Margin Architecture | The three-part workflow system: Context Foundation, Execution Design, and Quality Governance | The strategic design framework |
| AI Value | Value of completed work less compute, tooling, human-review, rework, and failure costs | The business measure |
| Cost per completed task | The all-in cost of producing an outcome that meets its quality bar | The core efficiency metric |
| Quality bar | The standard an outcome must meet to count as completed work | The constraint that makes the measurement meaningful |
AI Value is not a standardized accounting measure. Each workflow must define the value of its completed work, what quality threshold applies, and which costs and failure consequences are material. The point is to make the economics explicit rather than treating token spend as a proxy for value.
Why lower prices do not automatically create AI Value
Model prices are falling. Capability is diffusing into smaller and cheaper models. Caching, batching, open-weight models, local inference, and model portfolios continue to improve the supply side of the equation.
But declining unit prices do not guarantee declining total spend, or greater AI Value.
AI workloads can expand in response to lower cost. Teams add users and workflows, stretch context windows, deploy more agents, and accept more retries and reasoning steps. A system can get cheaper per token and more expensive per completed task at the same time.
Economists know the pattern as the Jevons paradox: efficiency gains can raise total consumption instead of lowering it. Whether that happens here depends on how the workflow is built.
| If the workflow is designed as… | Then lower model prices tend to create… |
|---|---|
| A chat interface attached to old processes | More usage, more tokens, more inconsistency |
| A frontier model as the default | Commodity work performed at premium cost |
| An uncontrolled agent loop | More retries, tool calls, and hidden failure modes |
| A governed operating system | More completed work and more AI Value per unit of compute |
The economic goal is not to suppress volume. It is to ensure that volume scales through lower-cost, reliable paths rather than through uncontrolled compute and human-review growth.
That is what Margin Architecture does.
1. Context Foundation
The first part of Margin Architecture is the Context Foundation:
Context, state, retrieval, and reusable inputs are designed as an economic layer, not assembled ad hoc at runtime.
Most AI systems treat context as a technical matter: concatenate a system prompt, chat history, retrieved documents, tool definitions, and a user request; then send it to the model.
That is convenient. It is also often expensive, slow, and unreliable.
Every token included in a request is something the system must process, pay for, and potentially expose across a model boundary. Every stale or irrelevant document increases cost without improving the answer. Every missing fact creates rework. Every unverified memory can create compounding error.
A better system distinguishes four kinds of information.
| Context element | What it contains | Design objective |
|---|---|---|
| Cacheable foundation | Stable policy, instructions, tool schemas, reusable knowledge, shared templates | Keep stable and reusable |
| Workflow state | Verified facts, prior decisions, completed steps, approved intermediate outputs | Keep concise, durable, and trusted |
| Retrieval | Current documents, records, policies, and evidence needed for this task | Retrieve selectively and citeably |
| Dynamic request data | The current user's request, case details, latest event, or changing inputs | Introduce only when relevant |
Stable context is an AI Value asset
Prompt caching exists because repeated computation is wasteful. When requests share an identical beginning, providers can reuse prior work rather than recomputing the full prefix. Microsoft's prompt-caching guidance explicitly ties identical beginning content to lower latency and cost, and recommends keeping stable content before dynamic content. Microsoft Learn
This is not merely a prompt-engineering trick. It is a workflow-design principle.
A high-volume workflow should make its durable elements explicit:
- Policy and safety instructions
- Output formats and schemas
- Stable tool definitions
- Reusable knowledge bases
- Standard operating procedures
- Workflow-specific quality bars
- Persistent but validated reference context
Then it should separate those stable elements from what changes by user, task, or moment.
The result is less repeated compute, lower latency, and a clearer boundary between reusable enterprise knowledge and dynamic case data.
The discount is material and varies by provider. Anthropic prices cache reads at a fraction of the standard input rate, a tenth for most current models and less for some, charges a premium to write a cache entry, and expires entries after five minutes or one hour depending on the option chosen. OpenAI documents discounts of up to 95% on cached input for some models, and newer models add a cache-write charge. Minimum prompt lengths, short lifetimes, and exact-prefix matching all apply, so a change near the start of a prompt can forfeit the saving. Cached tokens are discounted, not free. Check current pricing for the models you use before building a business case on it. (Anthropic, OpenAI)
State is not just memory
State is the compact record of what the workflow has already established.
It may include:
- A verified customer identity
- A completed classification
- An approved plan
- A tool result that does not need to be queried again
- A decision and its rationale
- A known exception
- A structured summary of a prior interaction
- A confidence score or quality result
The important word is verified.
A system that retains every model output as durable memory is not building intelligence. It is accumulating ungoverned claims. A system that preserves only reliable, relevant workflow state reduces future work while protecting the quality of subsequent decisions.
Context should become more useful over time, not merely larger.
Retrieval should add evidence, not noise
Retrieval is valuable when it supplies the evidence the workflow needs now. It is counterproductive when it becomes a reflexive "search everything" step that adds cost, latency, and ambiguity.
Good retrieval design asks:
- What does this task require that the system does not already know?
- Which sources are authoritative?
- Which source version is current?
- What is the minimum evidence needed to support the action?
- Should retrieved material be summarized, cited, or preserved in state?
- What should never leave the intranet or protected data boundary?
The goal is not maximum context. It is sufficient context for a reliable decision.
The economic test
For each context component, ask:
Does this information increase AI Value by reducing error, rework, or future compute by more than it costs to carry and process?
If the answer is no, it does not belong in the default context path.
2. Execution Design
The second part of Margin Architecture is Execution Design:
Every unit of work should follow the least costly path that can meet its required quality, latency, security, and reliability bar.
This is where decomposition, routing, tools, model selection, deterministic code, and human escalation belong.
A common failure in enterprise AI is to treat every request as if it were the same kind of work.
It is not.
A customer-message classification, a SQL lookup, a policy comparison, a research synthesis, a complex exception, and a regulated approval all have different requirements. They should not automatically use the same model, the same context window, the same tool chain, or the same level of human involvement.
The system should decompose a workflow into meaningful steps, then choose the right execution path for each one.
Routing is the execution decision
Routing asks:
Given the task, the available context, the required quality bar, the sensitivity of the data, and the expected cost, what should perform this next step?
The answer may be deterministic code. It may be a small local model. It may be an intranet-hosted model, a cloud model, a frontier reasoning system, an external tool, or a human reviewer.
The best execution path is not always the cheapest immediate path. It is the least costly path that produces a trustworthy completed result.
| Execution path | Appropriate work | AI Value logic |
|---|---|---|
| Deterministic code | Rules, validation, transformations, SQL, calculations | Avoid model calls where software is more reliable |
| Small or local model | Routine classification, extraction, simple drafting, privacy-sensitive tasks | Keep low-risk work inexpensive and close to the data |
| Intranet model | Internal knowledge work and repeatable enterprise tasks | Capture owned-inference economics and data control |
| Mid-tier cloud model | General reasoning, synthesis, common generation tasks | Balance capability with variable cost |
| Frontier model | Difficult reasoning, novel problems, high-value exceptions | Spend premium capability only where it changes the outcome |
| Human reviewer | Material risk, ambiguity, approvals, regulated or consequential exceptions | Reserve scarce expert attention for cases that merit it |
The organization's advantage does not come from guessing the "best" model once. It comes from owning the decision rules that match work to an execution path repeatedly.
Decompose before you optimize
Most workflows should not be routed as one undifferentiated request.
A contract-intake process, for example, may involve:
- Extracting fields from a document
- Checking them against known policy
- Identifying nonstandard clauses
- Comparing clauses against precedent
- Producing a summary
- Determining whether a reviewer must approve it
Those are different computational problems.
- Extraction may require a small model or structured document system.
- Policy checks may be deterministic.
- Clause comparison may require retrieval.
- A summary may need a mid-tier model.
- High-risk exceptions may require a frontier model or human reviewer.
- Final approval may be a governance decision, not a language-generation task.
Treating all six steps as one frontier-model prompt is simple. It is not architecture.
Tools are part of the execution economy
Tools are often presented as a capability upgrade: the model can browse, query databases, trigger workflows, call APIs, or write records.
But each tool call also has an economic and reliability profile:
- Invocation cost
- Latency
- Failure probability
- Data exposure
- Permission scope
- Retry behavior
- Downstream operational effect
- Human recovery cost if it acts incorrectly
Tool design therefore needs the same discipline as model design.
Use stable interfaces. Limit tools to the tasks that need them. Carry forward validated tool results as state. Define permission boundaries. Instrument failure and retry rates. Do not let an agent call a tool repeatedly because the system failed to preserve a result it already obtained.
Human review is a routed tier
In many workflows, the largest cost is not compute. It is indiscriminate human oversight, as the modeled contract example above suggests.
There are two weak models:
- Review everything
- Review nothing
The first destroys the automation business case. The second creates risk, error, and distrust.
A strong design treats human review as an execution tier, invoked when the workflow detects a meaningful exception:
- Confidence falls below a defined threshold
- Source evidence conflicts
- The task is classified as high stakes
- The output would trigger a consequential external action
- A policy, legal, safety, or compliance threshold is crossed
- The expected cost of an error exceeds the cost of review
This is where routing becomes a management system rather than a technical feature.
3. Quality Governance
The third part of Margin Architecture is Quality Governance:
Quality is not the last step. It is the control system that decides what can be trusted, what must escalate, and what the workflow learns from next.
Teams often frame quality as a tradeoff against efficiency. In practice, poor quality is often one of the most expensive forms of inefficiency.
A cheap answer that must be corrected is not cheap. A fast workflow that produces an unsupported recommendation is not fast. An autonomous tool call that triggers a recovery process is not automated in any economically meaningful sense.
Quality governance protects AI Value by preventing low-quality work from flowing downstream.
NIST's Generative AI Profile calls for documented test, evaluation, validation, and verification practices across the lifecycle (action IDs MP-2.1-002 and GV-1.5-003). The AI Risk Management Framework asks that human-oversight processes be defined, assessed, and documented (MAP 3.5) and that roles for human-AI oversight be assigned (GOVERN 3.2). NIST AI 600-1 · NIST AI RMF 1.0
Quality governance has three jobs
Assure
Determine whether the output meets the standard required for its task.
That may include:
- Structured-output validation
- Source-grounding checks
- Policy checks
- Unit and integration tests for tool use
- Evaluation sets
- Constraint checks
- Model-graded or rule-based scoring
- Human sampling
- Domain-expert review for critical workflows
The right quality bar varies by task. A low-stakes draft may need basic validation. A recommendation that affects pricing, eligibility, employment, credit, health, safety, or legal obligations needs a much higher bar.
Escalate
When the workflow cannot meet its quality bar, it should not pretend otherwise.
It should:
- Retry with a revised instruction
- Retrieve additional evidence
- Use a more capable model
- Switch to a different tool or model family
- Ask for missing information
- Route to a human reviewer
- Stop a consequential action
Escalation is not failure. It is a controlled response to uncertainty.
The expensive failure is not escalation. The expensive failure is sending low-confidence work into a downstream process that creates more rework, risk, or human intervention later.
Learn
This layer is often the least developed in AI architectures.
Quality governance should feed back into the workflow. Verified outcomes become durable state. Errors become evaluation cases. Escalation patterns identify weak retrieval, poor tool interfaces, insufficient policy context, or incorrect routing thresholds.
That is why the feedback loop from Quality Governance to State is central to the Margin Architecture model.
A workflow should not blindly store everything it produces. It should promote only verified outcomes into durable state.
| Outcome | What the system should do |
|---|---|
| High-confidence, validated result | Preserve as trusted workflow state when useful |
| Successful human correction | Add to evaluation data; consider updating policy, retrieval, or routing |
| Repeated model failure | Investigate context, tool interface, decomposition, and model-fit assumptions |
| Uncertain or conflicting evidence | Escalate or request clarification; do not promote to durable state |
| Tool failure or retry pattern | Improve tool contract, permissions, observability, or fallback path |
The economics are compounding. A system that learns from verified outcomes reduces repeat work. A system that repeats unverified outputs spreads costly uncertainty.
The feedback loop is the advantage
Margin Architecture is not a linear pipeline.
It is a loop:
Each loop should improve one or more of the following:
- The system needs less redundant context
- Retrieval becomes more precise
- Routing choices improve
- More work resolves on lower-cost paths
- Exceptions become easier to identify
- Human reviewers spend time where judgment matters
- Rework declines
- Cost per completed task falls
- AI Value increases
- Trust increases
This is how an AI system becomes economically better over time without becoming carelessly autonomous.
A worked example: contract intake
Consider a workflow that receives 10,000 contracts per month.
The naive design sends every document to a frontier model, asks it to interpret the contract, produces a summary, and sends every output to a legal reviewer.
That design may appear safe because a human sees every result. It is also expensive and difficult to scale.
Now apply Margin Architecture.
Context Foundation
- Stable contract policy, clause taxonomy, approved templates, tool definitions, and output schema are held as reusable context
- The workflow retrieves only the relevant policy clauses and comparable approved language
- A structured record carries verified party data, document metadata, prior negotiation decisions, and already-extracted terms
- Sensitive raw documents remain within the appropriate data boundary where possible
Execution Design
- Deterministic checks identify document type and mandatory fields
- A lower-cost extraction path parses standard terms
- Retrieval compares unusual clauses with policy and precedents
- A mid-tier model drafts a normalized summary
- A frontier model is reserved for ambiguous, high-value, or novel clauses
- A human reviewer receives only flagged exceptions and final high-stakes approvals
Quality Governance
- Structured fields are validated
- Missing or contradictory terms trigger retrieval or reprocessing
- Policy deviations receive a defined risk score
- Low-confidence cases escalate
- Reviewer corrections become evaluation examples
- Approved patterns become trusted workflow state for future similar contracts
What the model assumes
Every figure quoted above comes from these parameters. They are illustrative, chosen to be ordinary rather than favourable. Substitute your own and the arithmetic still runs.
| Parameter | Value |
|---|---|
| Contracts per month | 10,000 |
| Contract length | 30,000 tokens, roughly 22,000 words |
| Stable foundation: policy, clause taxonomy, output schema, tool definitions | 12,000 tokens |
| Naive path | Whole document plus the full foundation to a frontier model on every contract, nothing cached |
| Routed path | Deterministic checks, then small-model extraction, mid-tier clause comparison on the 25% that need it, mid-tier summary, frontier model on the hardest 8%, foundation served from cache |
| Reviewer loaded cost | $70 per hour |
| Review time per contract | 10 minutes |
| Share reviewed, naive | 100% |
| Share reviewed, routed | 15% |
| Token prices | Frontier $4 / $20, mid tier $2 / $10, small $1 / $5, cache reads $0.20, per million tokens |
Prices are the published rates for Claude Opus 5.5, Sonnet 5.5 and Haiku 4.5, checked on 30 September 2026. Model prices move; re-run the numbers rather than trusting these past their date.
What the model produces
| Monthly cost | Naive | Routed | Saved |
|---|---|---|---|
| Tokens | $1,980 | $812 | $1,168 (59%) |
| Human review | $116,667 | $17,500 | $99,167 (85%) |
| Total | $118,647 | $18,312 | $100,335 (85%) |
Two things are worth reading off that table.
The token column is the one everyone optimises, and it is 1.7% of the naive total. The review column is 98.3% of it. A team that negotiated a 20% discount on model pricing would save $396 a month. A team that changed who reviews what saved $99,167.
And the naive design is not obviously wrong. It routes everything to the best available model and puts a human in front of every output. It looks careful. That careful-looking version costs about six and a half times what the designed version costs, and almost none of the difference is tokens.
The aim is not merely lower token cost. It is a workflow that scales review capacity, protects quality, and holds or grows AI Value as volume rises.
That is the difference between adding AI to a contract process and designing an AI-native contract operation.
What to measure
A team cannot manage Margin Architecture using "AI adoption" or "tokens consumed" as its primary metric.
Those are activity measures. They do not tell you whether the workflow creates value.
The unit of measurement should be the AI Value of completed work at its required quality bar.
Context Foundation metrics
- Reusable-context ratio
- Cache-hit or reuse rate where supported
- Average context size per completed task
- Retrieval precision and evidence-use rate
- Share of state items that are verified
- Egress ratio: the share of data or tokens leaving the intended boundary
- Frequency of repeated retrieval or repeated tool calls
Execution Design metrics
- Cost per completed task
- Latency per completed task
- Tier mix by workflow step
- Share of work resolved through deterministic code or lower-cost model tiers
- Frontier-model use reserved for exceptions
- Tool-call success and retry rate
- Human-review minutes per completed task
Quality Governance metrics
- Task success rate at the defined quality bar
- Escalation rate and escalation reason
- Rework rate
- Error recurrence rate
- Human-correction rate
- False-accept and false-escalation rate
- Time to detect and correct drift
- State-promotion rate: how much verified output becomes durable workflow knowledge
These measurements transform the conversation.
Instead of asking, "Why did we use so many tokens?" leaders can ask:
Which part of the workflow is reducing AI Value, and what is the least costly reliable design change?
The practical design sequence
Margin Architecture does not require a fully autonomous multi-agent system. It begins with disciplined workflow design.
1. Choose one high-volume workflow
Start where there is enough repetition to learn:
- Contract intake
- Claims processing
- Customer support triage
- Sales research
- Compliance review
- Procurement analysis
- Internal knowledge support
- Incident management
Avoid starting with a broad, undefined "copilot" deployment. Choose a workflow with a clear beginning, end, owner, volume, quality bar, and cost baseline.
2. Map the real work
Document the inputs, the decisions, the handoffs, the sources of truth, the tools used, the exceptions, the human-review triggers, the outputs, the failure modes, and the cost and time of current execution.
This often reveals that the apparent "AI task" is actually several different tasks with different requirements.
3. Design the Context Foundation
Identify:
- What should remain stable?
- What should be retrieved?
- What becomes verified state?
- What must remain within a protected boundary?
- What information is repeated without creating value?
- What needs a structured representation rather than raw conversational history?
4. Route each workflow step
For every step, specify the cheapest viable execution tier, the required quality bar, the latency requirement, the data classification, the fallback path, the escalation condition, and the cost owner.
5. Define the quality system before scaling
Create:
- Evaluation cases
- Automated checks
- Human-review criteria
- Escalation thresholds
- Sampling plans
- Exception taxonomies
- A process for changing policy or routing rules
- A rule for what can enter durable workflow state
6. Measure completed outcomes
Track AI Value, cost, latency, quality, escalation, rework, and human minutes per completed task.
Then improve the system where it matters:
- Shrink unnecessary context
- Improve retrieval
- Compile recurring reasoning into code
- Move routine work down to lower-cost execution tiers
- Tighten tool interfaces
- Improve quality checks
- Reduce false escalations
- Turn verified corrections into better state and evaluation data
The organizational implication
Margin Architecture is not owned by a prompt engineer or a procurement team alone.
It crosses product, platform, security, operations, finance, legal, and the business process owner.
That requires clear ownership.
| Area | Primary owner | Core responsibility |
|---|---|---|
| Context Foundation | Product and AI platform | State model, retrieval design, data boundaries, reusable knowledge |
| Execution Design | AI platform and process owner | Decomposition, routing policy, tool interfaces, model portfolio |
| Quality Governance | Process owner, risk, and evaluation function | Quality bars, escalation policy, evaluation, review design |
| AI Value measurement | AI FinOps and business owner | Cost per completed task, value of completed work, budget, chargeback, value realization |
The organization that owns these decisions owns the value. The organization that treats them as vendor defaults rents the economics from someone else.
The strategic choice
The AI market has been making individual model calls cheaper, and that may well continue.
That is not the same thing as creating more AI Value.
A company can use lower-priced models and still increase its AI bill, expand its review queues, weaken its controls, and create more work than it eliminates. Or it can use the same underlying models to build a system that gets more useful, faster, cheaper, and more reliable as volume grows.
The difference is not just the model.
It is the architecture.
Context Foundation makes information reusable, current, and trustworthy. Execution Design directs each unit of work to the right mix of code, models, tools, and people. Quality Governance protects the outcome and converts verified work into better state for the next cycle.
That is Margin Architecture.
The economics of AI are designed into the workflow. AI Value is how you measure whether the architecture is working.
Sources and notes
- OpenAI: Prompt caching. Shared prompt prefixes can reduce repeated compute when request structures are suitable for caching, with discounted cached-input pricing on supported models.
- Microsoft Learn: Prompt caching with Azure OpenAI in Microsoft Foundry Models. Describes reduced latency and cost for requests with identical content at the beginning and recommends placing stable content before dynamic content.
- Anthropic: Prompt caching. Describes cache-control behavior, cache-read and cache-write pricing, and cache lifetimes.
- NIST AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. Guidance on managing generative-AI risks, including testing, evaluation, validation, and verification (MP-2.1-002, GV-1.5-003).
- NIST AI RMF 1.0. Framework guidance on governance, measurement, and defined human-oversight processes (MAP 3.5, GOVERN 3.2).
- The Routed Enterprise: An AI-Native Operating Model Built on Decomposition, Tiered Models, and Per-Call Cost Mapping. Internal project research. The contract-intake figures are modeled, not public data.
Provider caching prices were checked against the vendor documentation on September 30, 2026 and change often. Prompt-caching mechanics vary by provider and model family; the architectural principle is stable, in that reusable content should be separated from dynamic case data on purpose. State promotion requires security, privacy, data-retention, and model-risk controls appropriate to the workflow, and verified does not automatically mean permissible to retain. "Least costly" throughout means least costly while meeting the defined quality, reliability, latency, data-handling, and governance bar. AI Value is a decision framework rather than a standardized accounting ratio, so teams must define value, acceptable quality, rework, and failure cost for their own workflow.