Skip to content
Try CosmosGet Started
Back to Guides

Budgeting the Agent Fleet: Tokens, Compute, and Headcount

Sep 18, 2026
Ani Galstian
Ani Galstian
Budgeting the Agent Fleet: Tokens, Compute, and Headcount

Budgeting AI agents vs headcount means pricing the fleet as variable labor per merged change, because session-driven invoices outrun a seat budget while agent-opened pull requests keep expanding the salaried review queue.

TL;DR

Agent cost follows session volume and retry-driven turns, and each agent-authored change that requires human merge approval adds verification work. Budgeting AI agents vs headcount therefore carries separate fleet and supervision lines under a platform ceiling. Seat budgets miss retry-driven consumption; spend per merged change connects provider usage to delivered work.

A Vice President of Engineering (VP Engineering) signs a tooling line sized by seat count in January. The April invoice lands well past it because the vendor bills by consumption for the cloud agent a team turned on in March. Every retry adds sessions and turns. The seat count did not move. Sessions and retry-driven turns did, and a seat license has no field for either. The salary line now carries the time two senior engineers spend each morning reviewing agent-opened pull requests.

This guide gives the chief technology officer (CTO) and the VP Engineering a budget model for capacity bought per token alongside hired capacity. The two lines do not cancel, and that is what separates a fleet budget from a general AI tool ROI calculation. FinOps tooling sits outside this piece; forecasting mechanics and deflators, with supervision bandwidth as the binding constraint, live on the capacity planning guide, and engineering leaders set supervision capacity per engineer through the agents-per-engineer ratio. Spend per merged change, the mixed budget line, and governance that shapes consumption without starving the engineers who ship the most are covered here.

Budget the Fleet by Output

Agent capacity behaves like variable labor. The budget line should be sized by the work the fleet completes. Seats distribute access. A seat license buys access; an agent invoice buys inference and compute, including retries, and it moves with how many tasks the organization hands to agents across the software development lifecycle (SDLC) and how many turns each task takes. Running a fleet of specialized agents turns agent spend into a software factory problem measured at the fleet level.

A finance partner can defend a fleet running defined tasks under measurable limits; a terminal session is too variable and too personal to carry a budget of its own. Uber budgets this way at scale. More than 70% of its pull requests are now attributed to local or cloud agents, distinguished engineer Uday Kiran Medisetty wrote on August 27, 2026. Optimizing a fleet of specialized managed agents paired with dedicated benchmarks, he argued, is "inherently more cost-effective and scalable than optimizing individual terminal sessions". The operational shift to a hosted fleet is a separate argument; this piece takes the fleet as given and budgets it.

Cosmos, Augment Code's unified cloud agents platform available on all paid plans, gives the organization a named product line for this fleet consumption after seat budgeting stops describing the workload. Because Cosmos compute appears in paid-plan usage, finance can place that consumption on the fleet line. Human review goes on the supervision line that grows alongside it.

Budgeting AI Agents vs Headcount: The Two Cost Lines

Agent spend and headcount spend do not offset each other when an agent-authored change requires human merge approval because the verification work lands on a salaried engineer. A fully loaded full-time equivalent (FTE) carries salary, benefits, employer taxes, recruiting, and equipment. A fully loaded agent carries model inference and compute plus integration work, human-in-the-loop review, monitoring, and rework on rejected output. The human work rises as the fleet opens more changes. Booking agent spend on one line and a matching reduction on the salary line counts the review hours zero times.

The delegation evidence bounds the substitution a workforce plan can assume. More than half of 132 Anthropic engineers and researchers said in an August 2025 internal survey that they can fully delegate only 0 to 20% of their work to Claude. Respondents interpreted “fully delegate” as a range from no verification to light oversight. The survey covers staff at the company building the model, so its boundary is narrow, but the direction matches what a review queue shows: agents move work into the verification stage, and the stage stays staffed.

Supervision hours belong on their own line, costed by task risk and attributed to the fleet work that created the change.

Unit Economics Per Merged Change

The defensible number is spend per merged change. That number divides total fleet spend in a period, including provider consumption and service fees, by the count of agent-attributed changes merged in that period, with supervision hours costed at a loaded rate and added to the numerator. A token price is a vendor input; a merged change is an output the board recognizes. Rising token spend beside flat delivery is the productivity paradox in its budget form, and a per-change number is what shows whether delivery moved with it.

Uber's engineering team decomposes total agent spend into six terms that multiply: users, sessions per user, turns per session, requests per turn, tokens per request, and price per token. A budget owner's leverage sits in the middle terms, on the overhead the agent generates between an engineer's request and the answer. Adoption is meant to grow, and the vendor sets each model's list price.

  • Adoption and engagement: Users and sessions per user should grow when the fleet handles more useful work.
  • Interaction overhead: Turns per session measure work beyond the request an engineer made.
  • Request overhead: Requests per turn and tokens per request capture the remaining agent-generated workload.
  • Routing price: The vendor sets price per token. Model selection sits with the platform team.

Working the middle terms is what kept Uber's total spend flat while adoption climbed. Between February and mid-August 2026, its weekly active users across agentic tools grew 7x and weekly agent requests grew 9.4x while total AI spend stayed relatively flat from April onward. Holding one model constant from February to July, cost per 1,000 model requests fell almost 34% from its peak and cost per session fell 52% from its June peak, on a series that begins at the end of May. Those reductions are self-reported from one company and one monorepo environment, and Uber says plainly that another team's mileage will vary.

Rework affects the cost equation only when it creates additional sessions, turns, requests, or tokens. Human-equivalent hours measured against spend produce the agent hourly rate a team can use in this budget.

A Mixed Model: Fleet Spend, Supervision Hours, Platform Ceiling

The budget combines two operating lines with one exposure limit. Fleet spend covers provider consumption and the vendor's service fee. Supervision hours cover the review and monitoring work the fleet's output demands, including rework on rejected output. The VP Engineering sets the platform ceiling as the maximum exposure the organization accepts in a period, using the plan's allowance and top-up policy. For Cosmos, this structure separates paid consumption from the engineer time required to approve fleet output.

ComponentWhat it holdsUnitWho owns itThe lever that moves it
Fleet spendInference and compute at list price; service feeDollars per merged changePlatform engineering leadModel and cache defaults; payload size per request
Supervision hoursReview and monitoring; rework of rejected outputEngineer hours per merged changeGate owner, usually the review leadAutomated review before human review; task risk tiering
Platform ceilingIncluded allowance plus approved top-upsDollars per month, pooledVP Engineering with financeSpend tiers and manager sign-off on upgrades

Two flat self-serve plans, read on September 17, 2026, charge nothing per seat and pool up to 50 seats each. Standard runs $20 a month with $20 of usage included. Business runs $100 a month with $100 included. That usage covers model consumption, including large language model (LLM) inference and Context Engine usage, plus Cosmos compute time.

Augment bills usage at the provider's public API list price plus a flat 40% service fee on model inference, with no fee on compute. Teams purchase overage through top-ups, with top-up credits remaining valid for 12 months from purchase. A seat is any team member with access to the workspace, and teams above 50 seats move to Enterprise at custom pricing. In the model above, the allowance plus manager-approved top-ups forms the platform ceiling. This structure lets finance treat Cosmos consumption as pooled workspace exposure.

Other vendors price the same consumption in different units. A finance partner lining them up has to normalize first. Devin Enterprise bills in Agent Compute Units at a rate set in each customer's order form, and Cognition flags any session it sizes Large or Extra Large as unhealthy. On the standard session thresholds that starts above 10 ACUs, and above 100 on the Enterprise scale, where every band is ten times larger. GitHub prices its Copilot cloud agent in AI credits set by the model and the tokens consumed, at $0.01 per credit, plus the Actions minutes the session's workflow run burns. Subscribers still on legacy annual Pro and Pro+ request-based billing consume one premium request per session instead.

Governance Through Tiers and Nudges

Tiers with nudges protect the budget by making spend visible while an engineer can still change course. They preserve access for the engineers producing the most and require approval before the organization raises its exposure. Uber runs exactly this pattern across its interactive harnesses, nudging engineers in Slack at 50, 80, and 100% of expected spend. The signal reaches an engineer while the session is still running.

  • Live cost counter: A counter in each interactive tool's status line tracks spend per tool and across tools.
  • Shared tiers: One shared spend tier covers all interactive tools, with a separate tier for hosted fleet work.
  • Alerts before the ceiling: Teams configure alerts as usage approaches expected spend and again when it reaches the ceiling.
  • Manager sign-off: A manager must approve tier upgrades.

Uber's session-analysis dashboard flags 16 distinct anti-patterns and pairs each with a financial impact and a remediation, including multi-turn sessions run on a frontier model when a smaller one would do. Session traces are the raw material. Agent observability has to be in place before a spend dashboard can say anything.

Chargeback or showback by gate owner puts both lines on one statement. The owner is the review lead who holds a service's merge gate, a factory org design question before it is a finance one. A token-spend leaderboard is the anti-pattern, sometimes called “tokenmaxxing”: ranking engineers by consumption rewards agent overhead and says nothing about merged changes. Finance and the board see spend and supervision hours per merged change first, then the ceiling and its used share, with adoption reported separately on users and sessions. For Cosmos work, finance can use the pooled allowance as the shared ceiling rather than assigning a separate cap to every seat.

The platform team should set a cheaper default model for subagents that perform well-defined tasks and allow manual overrides when the task needs frontier reasoning. Cosmos makes the same trade automatically, routing each conversational turn to a model through Prism at 20 to 30% lower cost per task while preserving frontier quality. Interactive work and short-lived subagent tasks can justify different cache settings because their idle periods differ. Model Context Protocol (MCP) schema pre-loading can add request overhead before an engineer submits a prompt, especially when a session loads many tool definitions.

Cache time-to-live tuning, MCP payload overhead, and model-routing mechanics all belong to FinOps and platform economics, one level below the number the board sees.

Where the Fleet Budget Goes Wrong

Treating agent spend as a headcount substitute creates an unfunded review queue when agent-authored changes require human merge approval. Verification is the work that does not move. The workforce plan keeps the engineers who do it. Cutting junior roles to pay for the fleet adds a second cost: organizations that rely on AI to cut those roles will hollow out their pipeline of software engineering talent by 2028, Gartner predicted on July 7, 2026. Keep junior hiring in the workforce plan so the organization maintains the pipeline of engineers who will supervise future fleet work.

Open source
augmentcode/auggie281
Star on GitHub

The remaining failures come from optimizing at the wrong level or measuring the wrong unit.

  • Optimizing individual terminal sessions: Session-level tuning across thousands of engineers spends platform time on workloads that vary by person. Move the optimization hours to repeated fleet tasks with benchmarks, and leave interactive sessions to the tiers and nudges.
  • Using token budgets and hard caps: A token budget puts the middle terms of the cost equation, which belong to the platform team, in front of finance and turns the conversation into a debate about a vendor's price sheet. Keep token consumption in the platform team's dashboard. A hard monthly cap lands on whoever is shipping the most that month. Replace it with a tier, alerts before the ceiling, and manager sign-off, so the heavy user plans and the manager decides.

The failure modes above are cheap to fix before the fleet grows and expensive after. A tier set late has to be lowered against work already running. The platform engineering lead sets the model default, the review lead owns supervision hours, and a manager signs off on tier upgrades.

How to Build the Number

A defensible baseline ties the budget to delivered work and assigns every human or provider cost to the line that created it:

  • Delivery unit: Use spend per merged change for teams whose agents open pull requests, and spend per completed task for a fleet that closes tickets or triages incidents without a merge.
  • Human cost: Load the FTE line with salary, benefits, taxes, recruiting, and equipment.
  • Agent cost: Load the fleet line with inference at list price, the service fee, compute, and integration, and carry review hours, monitoring, and rework on the supervision line.
  • Supervision rule: Set the supervision ratio by task risk, so a schema change gets a named human approver and a dependency bump gets a spot check, and charge those hours to the supervision line.

Instrument the agent hourly rate on one team for one quarter before extrapolating. The trial gives finance a measured numerator against the chosen delivery unit before the model reaches other teams.

Wire the tiers before the first month of real volume, keeping Cosmos fleet work on its own. Then name the chargeback owner. The gate owner is the default, since the review lead is the one person who sees both lines. Read the spend and supervision figures next to cycle time and revert rate, the factory metrics that keep a falling cost per change from masking a rising revert rate.

The number changes as the factory matures. Where agents assist interactive sessions, supervision hours per change run high and fleet spend runs low. As tasks move to repeated fleet workflows with dedicated benchmarks, fleet spend per change falls and the supervision line concentrates on the gates that stay human. Rebase the budget as work moves up the software factory maturity model.

What to Do Next

Fleet capacity increases provider spend and reviewer demand on different schedules, so a seat-replacement promise cannot forecast both. Pick one team that already merges agent-authored pull requests. For one quarter, record fleet spend per merged change and the reviewer hours logged against each change, then calculate the share of the platform ceiling consumed. Present that number to finance at quarter end beside the team's cycle time, and rebase the fleet line for the next quarter from it.

Frequently Asked Questions About Budgeting AI Agents vs Headcount

Written by

Ani Galstian

Ani Galstian

Ani writes about enterprise-scale AI coding tool evaluation, agentic development security, and the operational patterns that make AI agents reliable in production. His guides cover topics like AGENTS.md context files, spec-as-source-of-truth workflows, and how engineering teams should assess AI coding tools across dimensions like auditability and security compliance

Related reading

Get Started

Give your codebase the agents it deserves

Install Augment to get started. Works with codebases of any size, from side projects to enterprise monorepos.