McDonald's vs Wendy's: AI Proposes, Business Rules Decide
Ricardo Argüello, October 9, 2026
CEO & Founder
General summary
An AI that makes decisions involving money cannot give two different answers to the same input. McDonald's, Wendy's and Hilbert show the working split: the model understands and proposes, while caps, validations and escalation to a person live in fixed rules around it.
- McDonald's tested IBM's automated voice ordering in about 100 drive-thrus from October 2021 and announced in June 2024 that the test was ending, after viral videos of orders nobody placed.
- Wendy's FreshAI reported a success rate near 99%, counting every order that reached the register even when a staff member had to step in and fix it.
- Hilbert, backed by a $28M a16z-led Series A, defines customers with models trained on years of history and lets its agents do no arithmetic beyond rounding.
- The pattern to copy: the model interprets and proposes, a deterministic rules layer checks amounts and quantities, and anything out of range goes to a named person.
Picture a cash register with a fast but distracted cashier. The cashier listens and keys in the order. The register refuses to ring up forty burgers without the manager's key. Nobody asks the register to understand accents, and nobody lets the cashier set prices unchecked. That is the split between the model and the rules.
AI-generated summary
Wendy’s said its drive-thru AI hit nearly 99%. Read the definition before you quote that number. According to Restaurant Dive’s December 2023 report, an order counted as a success if the chatbot started it and it reached the point of sale, even if a human had to join the conversation to fix an inaccuracy.
I’d call it the most useful design detail in the whole story. It’s also the first thing dropped when the case circulates on LinkedIn.
Two drive-thrus, one lesson
Andrew Amann’s newsletter “McDonald’s AI McTrouble” put the two chains side by side this week. His read is that McDonald’s asked the model to hear the order, make sense of it and ring it up, with no quantity caps and no escalation, while Wendy’s wrapped its model in business logic. “AI was the ears,” he writes, and the rules engine was part of the brain.
I checked what I could. McDonald’s and IBM started testing automated voice ordering in October 2021 across roughly 100 locations, and McDonald’s said in June 2024 it was ending the test, per Nation’s Restaurant News. The AP story describes a 2023 TikTok where the system kept stacking nuggets on one car’s tab while the customers, laughing, begged it to stop. Franchisees were told to switch it off by July 26.
Whether McDonald’s had zero caps, I can’t prove from public documents. The failure mode in those videos fits Amann’s diagnosis, though. A model that hears “more” keeps adding, because nothing outside the model knows what a normal order looks like.
On Wendy’s, the public numbers say something a bit different from “rules engine.” They say person in the loop. Twenty-two seconds faster than the local market average at a Columbus test site. A success rate near 99% that includes human fixes. And a stated target for fully hands-off orders of 30% or more, climbing to 70% by the end of 2024. I couldn’t find whether they reached 70%, so I won’t guess.
Either way, Wendy’s measured the system, model plus staff. McDonald’s story reads like the model was expected to carry it alone.
The startup that bans its agents from arithmetic
Then there’s Hilbert. It raised a $28M Series A led by a16z in April. In a Startupeable interview (in Spanish), CEO Nazli Tan, who ran growth at Getir, argues that piling more context on an agent makes it worse. Context usually means someone typed out their gut definition of “good customer” or “churn risk,” and the agent runs with it. Ask twice, get two lists.
Hilbert trains deep learning models on at least three or four years of real customer history and lets those models decide. Agents work inside the bounds. They can’t even do math beyond rounding. Tan’s reason: her customers route hundreds of millions in marketing spend off those answers, and nobody trusts a budget to an answer that changes when you ask again.
An agent company that won’t let its agents add. That’s the most mature thing I’ve read about agents this month.
Where the rules go
Here’s the split we draw before automating anything that charges, pays or discounts. Language goes to the model: a rushed order, an accent, engine noise in the background, a customer changing their mind halfway. Arithmetic and policy go to code.
- Caps on quantity and amount live in code, outside the prompt. You can talk a prompt out of a rule. You can’t talk a validation out of one.
- Prices, discounts and taxes come from your system of record. Never from what the model remembers.
- Escalation has an owner. Which cases go to a person, which person, and how long the customer waits.
- Measurement separates the two. Cases the model closed alone, cases that needed someone. Wendy’s did this, and it’s why their 99% means something once you read the footnote.
Skipping that count has burned this industry before. Amazon’s Just Walk Out got in trouble because nobody outside knew how much human review sat behind it. And that measurement compounds when you keep it as your own eval set, rerun every time you switch models.
A cheap test if you already have an agent near pricing, credit or orders. Pull ten real cases from last week. Run each one twice. Any pair that disagrees marks where your first rule belongs.
Let’s map the rules your agent should sit insideFrequently Asked Questions
McDonald's ran IBM's Automated Order Taker in about 100 restaurants starting October 2021 and announced in June 2024 that the test would end. Complaints and viral videos showed misread orders, including nuggets added over and over to one car's tab. McDonald's said it would explore voice ordering solutions more broadly.
Restaurant Dive reported that Wendy's FreshAI reached a success rate near 99%, defined as an order started by the chatbot and submitted to the point of sale even if a human had to join the conversation to fix an error. Wendy's goal for orders with no human intervention was 30% or more, rising to 70% by the end of 2024.
An AI agent that touches money should be split into two layers. The model interprets the request and proposes an action. A deterministic rules layer validates amounts, quantities and permissions, and routes anything out of range to a person. The same input then produces the same decision, and every exception is logged.
Hilbert raised a $28 million Series A led by Andreessen Horowitz in April 2026. CEO Nazli Tan says Hilbert trains deep learning models on three to four years of customer history to decide who is a good or at-risk customer, and its AI agents only act inside those bounds, without doing math beyond rounding.
Related Articles
Agent Autonomy Is a Liability, Not a Feature You Buy
Cognition raised $1B at a $26B valuation for an autonomous coding agent. In production, autonomy is the first thing that breaks. How much should you give it?
The pyramid before the agent: you almost never need one
The 'AI consultant' title expires. What stays is a discipline: start with a deterministic workflow and climb to an agent only when the problem demands it.
He Sells AI Agents. He Told His Team to Stop Using Them
Gumloop's founder told his team to stop automating everything with AI. The expensive failure of agentic AI isn't cost, it's the slop that loses customers.
Amazon Just Walk Out and the humans behind the cameras
A 2024 report about Amazon Just Walk Out is circulating again on LinkedIn this week. The lesson for anyone buying AI automation still holds.
SpaceX Automates Last. You Do It First.
The method SpaceX uses to redesign rockets puts automation dead last, after questioning, deleting and simplifying. Most AI projects start where it ends.
Anthropic Deleted 80% of Claude Code's System Prompt
Anthropic cut more than 80% of Claude Code's system prompt with no measurable loss on coding evals. The rules you added last year are now the ceiling.