> ## Content Index
> Fetch the complete content index at: https://www.anubhavpateriya.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Your AI agents are waiting for delegation of authority
- URL: https://www.anubhavpateriya.com/ai-agents-delegation-of-authority/
- Published: 2026-08-11T23:41:46.000Z
- Updated: 2026-08-18T01:14:42.000Z
- Description: AI agents are set to walk because nobody with the authority to say otherwise has written down how far they may run. On a twenty billion dollar business that is a forty million dollar setting.
- Author: Anubhav Pateriya

## Somebody chose Balanced, and nobody wrote it down

Somewhere in your company, in the last year or so, a meeting happened that nobody minuted. A vendor walked a team through the setup screens for a pricing agent, and about forty minutes in, on a screen with three choices on it, somebody chose Balanced. The setting decides how much evidence the agent waits for before it moves a price, and how far it is allowed to move it. Not Conservative, which felt timid, and not Aggressive, which felt like something you would have to explain later. Balanced. The room moved on, because there were eleven more screens.

The person at the keyboard was probably a pricing analyst, and they were doing their job. Nothing about that afternoon was careless. In the same building, on the same day, that analyst could not have signed a fifteen thousand dollar contract without two names above theirs on the paperwork.

Put those two afternoons side by side. The contract needed two signatures because fifteen thousand dollars is real money. The setting needed none, and it governs how that agent behaves across a portfolio for a year. On a large consumer business the pot behind it runs into the billions, which makes a single point of movement worth tens of millions, and I will show that arithmetic in a moment.

So it reads as a story about weak controls. It is not. The controls are working as designed. What that afternoon settled is how much authority the machine has, and what it may do before it has to ask. Every decision it takes from then on happens inside that boundary.

## You bought agents to move faster, then set them to walk

Start with what the money was for. Nobody approved an agent program in order to have a machine ask permission more politely than the last one, in fluent and tireless English, around the clock. The case was decision throughput, and the gap is worth stating without a multiplier attached: an agent can evaluate every item in every account every day, while a commercial team reviews a category once a month and gets to the accounts that shout loudest before the ones that matter most. That gap is what you bought.

Almost nobody is capturing that gap, and the reason is not the technology. In the absence of a written mandate, nobody has the standing to widen anything, because a limit nobody approved is a limit you will personally own when it goes wrong. So whatever was chosen on the day stays chosen, the exception queue grows into a second full-time job for the person the agent was bought to free up, and the organization pays enterprise prices for a system left at walking pace.

> Autonomy you have paid for and never authorized is a subscription.

So let me put a number on it, and be clear about what kind of number it is. Nobody has studied what agent limits are worth, because the practice is barely eighteen months old. Nothing that follows is a finding. It is an estimate from the programs I have run, and you should discount it the way you would discount anyone else's, including the ones that arrive with a decimal point and a footnote pointing at nothing. The test is never who is saying it. It is whether they will tell you where it came from.

Here is where mine came from. Across twenty years of commercial programs in large consumer businesses, what I have seen when trade allocation genuinely improves, not the plan but the allocation, is two to four points of spend moving out of promotions that lose money and into promotions that pay. Below about five billion in revenue I have seen a good deal more, because the starting point is usually worse and less of the budget is locked into habit.

On a twenty billion dollar business, trade spend in the high teens of gross sales puts the pot near four billion, so a point of mix is worth about forty million dollars. Run it on your own numbers, and if your post-promotion analysis says something different, use yours. Agents do not raise that ceiling. They change how often you can act on it, which is worth nothing until somebody writes down how far they may go.

So the question is not whether your controls are strong enough. It is whether they are specific enough to let you go faster on purpose.

## Your approval rules are not broken. They are watching transactions, not settings.

Every company of any size owns a delegation of authority schedule, the grid that sets out which role may commit the company to what, up to which number, and who has to counter-sign. It was drafted by a general counsel and a finance director in a week neither of them enjoyed, and it gets opened about twice a year, once when a new director joins the board and once when something has already gone wrong.

You will hear it said that agents will bleed a company by a thousand cuts, each one too small for anyone to notice. It is a good line and, if your schedule is in decent condition, it is mostly wrong. Each role almost always carries two limits, one for a single transaction and one for the month or the year, for exactly the reason you would guess: somebody worked out long ago that four modest invoices attract less attention than one large one. The thousand-cuts problem was solved decades ago, by people who never met an agent and were guarding against something far older.

The agents are also better behaved than their reputation. In pricing this year the working pattern is that the agent proposes and a person disposes. On a Monday morning a category manager opens a queue of recommendations, works through the exceptions the system has flagged, accepts the rest, and anything touching a major customer goes up to somebody with a name and a bad week if it turns out badly. Every step carries an audit trail.

You might object that these systems are hardly unwatched, and you would be right. They get reviewed constantly, in performance decks, vendor QBRs and exception reports. Problem is who sits in those reviews. The people reading the numbers sit in technical and compliance functions, and nobody in that room has the commercial authority to change the setting. The review is technical but the decision is commercial, so whatever was chosen on the day survives every review of it.

So the money is guarded and the steps are logged. What is not guarded is the thing that decided what would be in that queue in the first place.

## The real decision was taken weeks earlier, on the setup screen

An analyst who clicks approve on eight hundred price recommendations has approved them. Nobody would say they decided them. A batch takes about as long as a single one and feels a good deal more useful, and that is not laziness, it is arithmetic, because nobody reads eight hundred of anything on a Tuesday.

The decision that mattered was taken back on the setup screen. What the agent may touch, which customers sit inside its scope, how much evidence it needs before it moves, what it protects when two goals pull against each other. Those four answers are the agent's operating limits, and once they are set, every approval afterwards happens inside the boundary they drew. The person clicking approve is choosing between options that somebody else already narrowed, at a meeting with no minutes.

And that earlier choice appears nowhere on the schedule. There is a row for capital spend, a row for contracts, a row for hiring, a row for writing off bad debt, and no row at all for who may set the operating limits of a system that will touch several thousand prices before Christmas.

![Timeline showing the setup screen in week zero marked as the only point a commitment was made, followed by the agent running all year, weekly batch approvals, and a year-end result where the budget lands on plan but the mix is wrong.](https://storage.ghost.io/c/5d/c9/5dc96621-b880-4c3b-8598-5623c532588f/content/images/2026/08/fig_authority-timeline_in-body.png)

Approvals continued all year. The decision was made once, in week zero.

Nobody decided to leave that row off the schedule. You could object that none of this is new, and you would be half right. The last generation of enterprise software encoded commercial policy too, in credit limits, pricing conditions and approval workflows, all set by implementation teams in rooms much like that one. What has changed is what sits downstream of the setting. In an ERP a person still executed each transaction, so a badly chosen rule met a human being before it met a customer, and somebody usually noticed. A pricing agent removes that step. The setting is the decision now, several thousand times over, with nobody in between. I have spent a good part of two decades on the consulting side of large consumer and retail businesses, and I read these schedules early in an engagement, usually in the first fortnight, because they tell you more about how a company works than the strategy deck does. I have never seen one with a line for this. I would be glad to be sent one.

That would matter less if the work sitting inside those limits were trivial.

## Pricing, promotion and markdown agents are already running inside those settings

The work is neither trivial nor hypothetical. It is unglamorous, which is exactly what makes it easy to miss, and it runs at a volume no team could match.

On trade promotion, an agent watches an event from the first day instead of the post-mortem, catches an underperforming mechanic while the money is still in the room, and adjusts inside the parameters the customer contract already allows. PepsiCo has published what this looks like at full scale, in a peer-reviewed account of two systems it runs called PromoAI and PricingAI. One searches millions of combinations of product, promotion and timing to build the promotional calendar5. The other sets base prices across the portfolio from estimated price elasticities. That is the trade calendar and the price list, computed rather than argued. On pricing, it weighs a competitor move, the stock position, the days left in the season and the margin target at the same moment, instead of in four spreadsheets owned by three people, one of whom is on holiday. In service it clears the small credit in ninety seconds instead of ninety hours.

On the retailer's side of the same desk the picture is a mirror image. This month Dunelm, the UK homewares chain, put a shopping agent into its iOS app, built on Google Cloud's Gemini, letting customers describe a room in their own words and get a tailored set of products back6. Ask what decides which products come back and you are asking a merchandising question, not a technology one. The same shift is running through markdown timing and personalized offers, which means the negotiation between a brand and its largest customer is picking up two automated participants, each pushing toward a different goal at a speed neither commercial team can follow, while the joint business plan that governs the relationship still gets renegotiated once a year over lunch.

Every one of those decisions is defensible on its own, and nothing in any single action would fail a review. The question is what all of them together do to the largest discretionary number on the P&L.

## Budgets check the total. Nothing checks the mix.

Take trade spend, because it is the clearest case. Let us be precise about the word mix first, since it does the work here: mix is not how much you spend, it is where the money lands, meaning which accounts, which packs, which mechanics, at which depth. Trade spend is the largest pot of discretionary money a consumer goods company controls, and it has an odd property. It never appears as its own line in the published accounts, because it comes off gross sales before any other cost is counted. Industry estimates put it in the high teens as a share of gross sales, and since nobody is required to disclose it, those estimates come from advisers rather than filings.

Now the part every revenue growth manager already knows. Nielsen has found for years that more than seventy percent of trade promotions do not break even4. Agents did not create that, and I am not going to pretend they did. It survives because every control around trade spend was built to answer one question: did we spend more than we said we would.

The agent makes that old problem faster and much harder to see forming, because it optimizes against what it can measure, and some things are far easier to measure than others. A deep price promotion in a high-velocity account produces a clean, attributable lift inside a week. Investment in a slower account, or in a pack that builds a category over three years, shows almost nothing in the same window. Point an agent at measurable return and it will drift toward the first and away from the second, every week, without ever breaching a parameter, and without once mentioning that it is doing so, because nobody asked it to.

So picture the year. The budget lands within a point of plan, every event sat inside approved parameters, every price move was inside the band, the audit trail is immaculate, and the business still ends up with more volume, thinner margin and two strategic accounts quietly underinvested. The review pack will be entirely green. Cumulative limits catch overspending and audit trails prove each step was authorized by somebody entitled to authorize it. Not one of those controls was built to ask whether nine hundred correct decisions add up to the mix you meant to hold, because until recently a human sat in the middle whose whole value was noticing when they did not.

> None of it was settled in a strategy session. It was settled on a screen with three options on it.

That noticing is what you are now buying back, and the only place to install it is in the settings themselves. Which is why the setup screen is not an IT decision. It is where your commercial strategy is now written, in a vendor's user interface, by whoever happened to be free that afternoon.

You might reasonably ask whether somebody is going to solve this on your behalf. This month the industry answered.

## The new agent standard leaves permissions to you, deliberately

On August 6, 2026, OpenAI, Amazon, Microsoft, Vercel and Cursor published Agent Plugins 1.0, a shared standard1 for how an agent's skills and tool connections get packaged so that any system can read them. Google joined as a core maintainer the same day. Five companies who agree on almost nothing agreed on this in about the time it takes a large enterprise to schedule a kickoff.

What the standard covers is the packaging. What it leaves out, by its own stated scope, is who may install one of these things, what it is allowed to reach once installed, and what stops it doing damage. It sets out how an agent carries its tools. It says nothing about what the agent may do with them.

That gap is deliberate, and the people who wrote it were right to leave it. Nobody outside your building can set your permissions, for the same reason nobody outside your building can write your trade terms, and anybody offering to would be handing you a template with somebody else's appetite for risk already loaded into it, usually somebody with no customers of their own to lose. What an agent may do to a price, a promotion or a customer relationship is a statement about what your company believes, and a standards body has no business having a view on that.

So take it as the notice it is. The plumbing arrived faster than anyone expected. The permission did not arrive, and it is unlikely to. I argued in [The Judgment Layer](https://www.anubhavpateriya.com/the-judgment-layer/) that this part has no vendor. Here is the specification agreeing, in writing.

If no standard is going to set your number, the next instinct is to find somebody who has set theirs and copy it. That turns out to be harder than it sounds, and the one company that has published a number is more interesting for what it did afterwards.

The clearest published number comes from an adjacent industry, and where it comes from turns out to be part of the point. Airbnb said this month that its AI assistant now resolves close to forty-five percent of inquiries that start with it without a human agent, up from forty percent the previous quarter, and that support cost per booking fell about sixteen percent year on year2. Most people read that as a cost story. Read the two numbers together instead. Forty, then forty-five, one quarter apart. Airbnb has not said it widened a limit, but the assistant runs in more than fifty languages now with voice next, so somebody decided it could handle more and the number followed.

The loop is still the part worth building, because it answers the objection that stops most of these conversations before they start. Nobody can tell you the right limit for your business on day one, and waiting until you are certain means waiting forever.

> You do not need the right number. You need a defensible starting number, a review date, and an agreed rule for what would let it widen.

Notice who did the disclosing, though. A platform business whose investors expect operating detail every quarter. It is hard to picture a manufacturer or a grocer telling the market what share of its pricing ran without a person, and harder still to picture an investor relations team volunteering it. So the benchmark you would like to copy probably does not exist, and there may be no industry average to point at when your board asks whether your number is sensible. Which makes the real question not what the number should be, but who is entitled to set it, and who is coming back to move it.

## Decide who is allowed to choose Balanced, then go about widening it

There is a reason to settle this before the next deployment rather than after it. Gartner expects more than forty percent of agentic AI projects to be cancelled by the end of 2027, and the causes it gives are escalating costs, unclear business value and inadequate risk controls3. Not one of those is a model problem. These programs are not dying because the agent could not do the work. They are dying because nobody would put a name to what it was allowed to do while doing it, so the work never widened, never produced a number worth defending, and drifted until a finance review put it out of its misery.

Which makes writing those limits the opposite of an administrative fix, and the wrong thing to hand to whoever manages the vendor. How much of your commercial position a machine may set is a delegation of authority, and it belongs where every other change to that document goes: approved by the board or the executive committee, recorded the same way, reviewed on the same cycle.

I have argued before that judgment has to be a governed artifact rather than a settings screen. This is where that starts: one workflow, one page, one name. A company does not get an enterprise judgment layer by declaring one. It gets there by writing enough of these that they have to agree with each other.

Three things make it real, and the third is the one companies skip. I will keep the examples in revenue growth management, where the money is most visible, though none of it is particular to that function.

**Write the limits, including how they widen.** Take the workflow that is live rather than in pilot and put its limits on one page. What it may touch, by category and by customer. How much of the portfolio it may move in a period. How certain it must be before it commits, which is a statement about your appetite for being wrong and not a technical setting. What it escalates, to whom by name, within how many hours. What the executive committee sees monthly, including how many decisions it made and what share a person overrode, because the override rate is the one number here a pilot cannot flatter. And the line most companies leave out: what would earn it more room, stated in advance. If overrides stay low for two quarters and the outcomes hold, the limit moves, by a stated amount, on a stated date. A limit with no widening rule never gets widened, because nobody wants to be the person who proposed loosening a control that has not yet failed.

![Four-step loop for a revenue growth management agent: set a limit, watch overrides, check outcomes, widen it, returning to the start on the review date. A marker between checking outcomes and widening reads: most decision makers stop here.](https://storage.ghost.io/c/5d/c9/5dc96621-b880-4c3b-8598-5623c532588f/content/images/2026/08/fig_widening-loop_in-body.png)

The widening loop, as a worked example for a revenue growth management (RGM) agent.

**Record them where authority lives.** Put that page on the schedule as its own row, with an approval level beside it and a review date. Not a function. A person, the way that document has handled every other commitment for a century.

**Then make the organization believe it.** A row changes nothing on its own, because the next implementation is already booked and the consultant running it will ask a room to pick a setting, exactly as last time, and somebody in that room will feel the pull of the middle option all over again. Nobody who configures these systems has ever needed sign-off for a configuration, and nothing in their training says this one is different. It has to be said out loud, more than once, to the commercial teams, to procurement and to the vendors, until the room knows that the person best placed to answer is not in it.

If you cannot fill one of those lines, you have not found a technology gap. You have found a commercial position your company has never taken, and the agents will take it for you this quarter, in whichever direction the vendor thought reasonable.

Writing that page is the cheapest thing you will ever do to make an expensive system earn out. No platform, no program, no new function. Your company has known how to put a name on a commitment for a hundred years. It has never once thought of a setting on a screen as one.

Somewhere in your building, somebody chose Balanced. They may well have been right for that quarter. The questions are whether anyone decided they were the one to choose, and whether anybody is coming back to widen it.

*Read next:* [*The Judgment Layer*](https://www.anubhavpateriya.com/the-judgment-layer/) *·* [*The Destination Problem*](https://www.anubhavpateriya.com/the-destination-problem/) *·* [*The line you won't cross*](https://www.anubhavpateriya.com/the-line-you-wont-cross/) *·* [*Decision Engineering*](https://www.anubhavpateriya.com/decision-engineering/)

## Sources

1. Agent Plugins 1.0, published 6 August 2026 by OpenAI, Amazon, Microsoft, Vercel and Cursor, with Google as a core maintainer. The specification covers packaging only and explicitly excludes distribution, permissions, provenance and sandboxing. [source](https://thenextweb.com/news/openai-agent-plugins-open-standard-skills-mcp?ref=anubhavpateriya.com)
2. Airbnb second quarter 2026 results and management commentary on AI-assisted customer service. [source](https://news.airbnb.com/airbnb-q2-2026-financial-results/?ref=anubhavpateriya.com)
3. Gartner, on agentic AI project cancellations by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. [source](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027?ref=anubhavpateriya.com)
4. Nielsen's analysis of more than one million UPCs and 39 million promotional events, covering $555 billion of US retail sales across 150 retail banners, found that almost three-quarters of promotions do not break even. Other industry findings put the figure nearer two-thirds. [source](https://www.foodnavigator.com/Article/2014/10/09/Three-quarters-of-CPG-promotions-don-t-break-even-Nielsen/?ref=anubhavpateriya.com)
5. PepsiCo's PromoAI and PricingAI, published in the INFORMS Journal on Applied Analytics, 11 June 2026\. PromoAI pairs machine-learning promotional forecasts with mixed-integer linear programming to build promotional calendars across trade channels; PricingAI sets base prices using Bayesian hierarchical models of own and cross-price elasticities. Both are optimization systems operating at scale rather than conversational agents. [source](https://pubsonline.informs.org/doi/abs/10.1287/inte.2025.0302?ref=anubhavpateriya.com)
6. Dunelm's Ask Dunelm shopping agent, live in its iOS app on Google Cloud's Gemini Enterprise. [source](https://www.retailgazette.co.uk/blog/2026/08/dunelm-launches-ai-shopping-assistant-with-google-cloud/?ref=anubhavpateriya.com)