Aaron Levie's list: the real work of deploying AI agents
Ricardo Argüello, October 11, 2026
CEO & Founder
General summary
On September 30, Box CEO Aaron Levie posted the list of work it takes to deploy AI agents inside a company: data, connections, redesigned workflows, human oversight, evals and continuous updates. The labs have already committed around $9.75 billion to teams that do this work. Three of the seven items can't be outsourced.
- Levie named seven workstreams: legacy systems to the cloud, data access, software connected to agents, redesigned workflows, human-in-the-loop design, evals, and updates with every new model.
- His core point is that with agents you no longer install a tool, you deliver work output inside a process, and the implementation job changes completely.
- Per Petar Petrov's analysis, OpenAI, Anthropic, Microsoft, Amazon and Google committed about $9.75 billion between April and early July to teams that embed in companies to deploy AI.
- Data access decisions, where a person approves before the agent acts, and your own evals stay with the company even if a partner handles everything else.
- IQ Source's AI Maestro program spends two months mapping the real process and ends in a Go/No-Go gate before anyone is paid to deploy.
Imagine buying a commercial oven for your restaurant. The vendor installs it, hooks up the gas and leaves. Now imagine hiring a cook instead. Nobody can install a cook and walk away, because someone has to say what to make, taste what comes out and decide what gets served. Deploying AI agents looks a lot more like the cook.
AI-generated summary
Most of the AI deployment conversation is about who gets paid to do it. Aaron Levie’s post on September 30 was more useful than that, because he wrote down what “it” actually is.
His list, lightly trimmed. Legacy systems moved to the cloud, data organization and access updated, software connected to agents in new ways, workflows reengineered for agents, human-in-the-loop figured out per process, evals generated and maintained, and the whole system updated as new models ship. Then the honest part. “And the full list may even be longer.”
I read it twice and did one thing with it. I sorted the seven items by who has to own them. Four you can buy. Three you can’t, and those three decide whether the other four are worth paying for.
Why the old implementation model breaks
Levie’s distinction is the one I’d underline. Software implementation used to mean installing a known category of tool, stepping back, and letting the customer run it. With agents, “you’re deploying work output in a process.”
An ERP never decided anything on its own. An agent that reconciles invoices does, every time it runs. Somebody has to own what happens when it’s wrong, and that somebody is rarely the firm that set it up and moved on to the next client.
The money is already moving
Petar Petrov tallied it in his September 30 newsletter. Between April and early July, OpenAI, Anthropic, Microsoft, Amazon and Google committed around $9.75 billion to teams that go inside a company and stay until the AI is used. I covered what that meant for Accenture back in July.
The number in his piece that matters more for a mid-size company comes from McKinsey, as he cites it. Among companies using AI, about 20% have redesigned workflows around it. Among the ones getting real financial results, it’s 55%.
Workflow redesign is item four on Levie’s list. It’s also the line between companies that see returns and companies that see invoices.
Seven items, two owners
| Workstream | Can you hire it out? | What breaks if you skip it |
|---|---|---|
| Legacy systems to cloud | Yes | The agent has nothing to plug into |
| Data organization and access | Only the plumbing | The agent sees too much, or too little |
| Software connected to agents | Yes | Brittle integrations that fail on every update |
| Workflows reengineered | Partly | AI bolted onto the old process |
| Human in the loop | No | Nobody knows who signs off before the agent acts |
| Evals | No | You can’t tell if it works, or if it got worse |
| Updates with each model | Yes, by contract | The system ages in months |
The three “no” rows share something. They’re decisions, and a vendor can only implement decisions somebody already made.
A partner can configure permissions on your payroll data. They can’t tell you who should have them. They can wire an approval step before the agent emails a client. They can’t tell you which emails need one. And evals, the set of real cases with known correct answers, can only be written by whoever knows the answers. That’s your operations lead who has done the work by hand for ten years, not an engineer who met your company last Tuesday.
Levie said it in seven words five days earlier, “You can’t automate what you can’t measure.” I’ve argued before that the eval is the part you keep when prompts and models change. Partners rotate off. Your evals stay.
What we do at IQ Source before anyone deploys
If you hire a deployment team before settling those three decisions, you’re paying them to make the decisions for you, using whatever they learned in a week of interviews.
That’s the gap AI Maestro is built for. It sits before deployment rather than replacing it, and it typically runs two months with deliverables every two weeks. Four come out of it. A Process Reality Map of how the company actually works (it rarely matches the org chart), a definitions audit that finds where the same word means two things in two departments, an AI Opportunity Score, and a ranked list of opportunities that ends in a Go/No-Go gate. It runs on whatever AI tools the company has already approved, so nothing new gets bought to start.
What comes out is the half of Levie’s list you can’t contract. Which data, who approves, how it’s measured. With that settled, the other half gets quoted for less, because the implementer stops guessing.
Levie closed by saying it’s a great time to be a forward deployed engineer. Agreed. It’s also a good time to walk into that first meeting with your list already sorted.
Sort your deployment list with AI Maestro firstFrequently Asked Questions
Box CEO Aaron Levie uses the AI deployment layer to mean the work of getting AI into a company's real processes: moving legacy systems to the cloud, updating data access, connecting software to agents, reengineering workflows, designing human-in-the-loop review, maintaining evals, and updating everything as new models ship. He calls it a huge opportunity.
According to Petar Petrov's analysis published September 30, 2026, OpenAI, Anthropic, Microsoft, Amazon and Google committed around $9.75 billion between April and early July 2026 to teams that embed inside companies to deploy AI, including the $4 billion OpenAI Deployment Company.
When deploying AI agents, three decisions stay with the company even if it hires an implementation partner: who gets access to which data, which process steps need a person to approve before the agent acts, and which evals define whether the agent's work is correct. Cloud migration and integrations can be contracted out.
With traditional software, an implementer installed a well-understood category of tool and stepped back. With AI agents, as Aaron Levie put it, you deliver work output inside a process, so the workflow has to be redesigned, human oversight defined, results measured with evals, and the system updated whenever the model changes.
Related Articles
Your marketing asks AI what to cut. Ask this instead.
Box created 13 new AI roles, one to market to industries it could not staff before. The question is not what AI lets you cut, but what it makes possible.
Evals Are the New Process Documentation
Aaron Levie and Garrett Lord agree: AI programs stall at pilot because the company cannot define what good looks like. Evals are that definition, written.
AI doesn't cheapen your product, it changes your margin
OpenAI launched Deployment Co. Anthropic hit $45B ARR. Stripe embeds 1 AI engineer per 20 employees. Prices aren't falling. The delivery stack changed.
Building costs zero. Distribution is the new investment.
Aaron Levie, Gergely Orosz, and Eric Siu published the same thesis in 36 hours: building got commoditized. Owned distribution loops are the 2026 moat.
AI isn't the problem. Your company isn't ready.
Aaron Levie and Daniel Miessler shipped the same diagnosis from opposite angles. Companies that can't describe themselves won't get value from AI.
Levie Is FDE-Pilled. Tanmai Gopal Named the Part That Costs.
Aaron Levie defended forward deployed engineering. Tanmai Gopal answered with what breaks the cheap version: shared context costs money every single day.