Skip to main content

McDonald's vs Wendy's: AI Proposes, Business Rules Decide

McDonald's shut down its AI drive-thru. Wendy's reported 99%, counting orders a human fixed. Hilbert won't let its agents do math. Money needs rules.

McDonald's vs Wendy's: AI Proposes, Business Rules Decide

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

AI & Automation 4 min read

Wendy’s said its drive-thru AI hit nearly 99%. Read the definition before you quote that number. According to Restaurant Dive’s December 2023 report, an order counted as a success if the chatbot started it and it reached the point of sale, even if a human had to join the conversation to fix an inaccuracy.

I’d call it the most useful design detail in the whole story. It’s also the first thing dropped when the case circulates on LinkedIn.

Two drive-thrus, one lesson

Andrew Amann’s newsletter “McDonald’s AI McTrouble” put the two chains side by side this week. His read is that McDonald’s asked the model to hear the order, make sense of it and ring it up, with no quantity caps and no escalation, while Wendy’s wrapped its model in business logic. “AI was the ears,” he writes, and the rules engine was part of the brain.

I checked what I could. McDonald’s and IBM started testing automated voice ordering in October 2021 across roughly 100 locations, and McDonald’s said in June 2024 it was ending the test, per Nation’s Restaurant News. The AP story describes a 2023 TikTok where the system kept stacking nuggets on one car’s tab while the customers, laughing, begged it to stop. Franchisees were told to switch it off by July 26.

Whether McDonald’s had zero caps, I can’t prove from public documents. The failure mode in those videos fits Amann’s diagnosis, though. A model that hears “more” keeps adding, because nothing outside the model knows what a normal order looks like.

On Wendy’s, the public numbers say something a bit different from “rules engine.” They say person in the loop. Twenty-two seconds faster than the local market average at a Columbus test site. A success rate near 99% that includes human fixes. And a stated target for fully hands-off orders of 30% or more, climbing to 70% by the end of 2024. I couldn’t find whether they reached 70%, so I won’t guess.

Either way, Wendy’s measured the system, model plus staff. McDonald’s story reads like the model was expected to carry it alone.

The startup that bans its agents from arithmetic

Then there’s Hilbert. It raised a $28M Series A led by a16z in April. In a Startupeable interview (in Spanish), CEO Nazli Tan, who ran growth at Getir, argues that piling more context on an agent makes it worse. Context usually means someone typed out their gut definition of “good customer” or “churn risk,” and the agent runs with it. Ask twice, get two lists.

Hilbert trains deep learning models on at least three or four years of real customer history and lets those models decide. Agents work inside the bounds. They can’t even do math beyond rounding. Tan’s reason: her customers route hundreds of millions in marketing spend off those answers, and nobody trusts a budget to an answer that changes when you ask again.

An agent company that won’t let its agents add. That’s the most mature thing I’ve read about agents this month.

Where the rules go

Here’s the split we draw before automating anything that charges, pays or discounts. Language goes to the model: a rushed order, an accent, engine noise in the background, a customer changing their mind halfway. Arithmetic and policy go to code.

  1. Caps on quantity and amount live in code, outside the prompt. You can talk a prompt out of a rule. You can’t talk a validation out of one.
  2. Prices, discounts and taxes come from your system of record. Never from what the model remembers.
  3. Escalation has an owner. Which cases go to a person, which person, and how long the customer waits.
  4. Measurement separates the two. Cases the model closed alone, cases that needed someone. Wendy’s did this, and it’s why their 99% means something once you read the footnote.

Skipping that count has burned this industry before. Amazon’s Just Walk Out got in trouble because nobody outside knew how much human review sat behind it. And that measurement compounds when you keep it as your own eval set, rerun every time you switch models.

A cheap test if you already have an agent near pricing, credit or orders. Pull ten real cases from last week. Run each one twice. Any pair that disagrees marks where your first rule belongs.

Let’s map the rules your agent should sit inside

Frequently Asked Questions

McDonald's ran IBM's Automated Order Taker in about 100 restaurants starting October 2021 and announced in June 2024 that the test would end. Complaints and viral videos showed misread orders, including nuggets added over and over to one car's tab. McDonald's said it would explore voice ordering solutions more broadly.

Restaurant Dive reported that Wendy's FreshAI reached a success rate near 99%, defined as an order started by the chatbot and submitted to the point of sale even if a human had to join the conversation to fix an error. Wendy's goal for orders with no human intervention was 30% or more, rising to 70% by the end of 2024.

An AI agent that touches money should be split into two layers. The model interprets the request and proposes an action. A deterministic rules layer validates amounts, quantities and permissions, and routes anything out of range to a person. The same input then produces the same decision, and every exception is logged.

Hilbert raised a $28 million Series A led by Andreessen Horowitz in April 2026. CEO Nazli Tan says Hilbert trains deep learning models on three to four years of customer history to decide who is a good or at-risk customer, and its AI agents only act inside those bounds, without doing math beyond rounding.

McDonald's Wendy's Hilbert AI agents business rules AI automation AI operations

Related Articles

Agent Autonomy Is a Liability, Not a Feature You Buy
Business Strategy
· 7 min read

Agent Autonomy Is a Liability, Not a Feature You Buy

Cognition raised $1B at a $26B valuation for an autonomous coding agent. In production, autonomy is the first thing that breaks. How much should you give it?

AI agents autonomous agents software architecture
The pyramid before the agent: you almost never need one
Business Strategy
· 7 min read

The pyramid before the agent: you almost never need one

The 'AI consultant' title expires. What stays is a discipline: start with a deterministic workflow and climb to an agent only when the problem demands it.

automation pyramid AI agents deterministic workflows
He Sells AI Agents. He Told His Team to Stop Using Them
Business Strategy
· 6 min read

He Sells AI Agents. He Told His Team to Stop Using Them

Gumloop's founder told his team to stop automating everything with AI. The expensive failure of agentic AI isn't cost, it's the slop that loses customers.

AI slop AI automation AI agents
Amazon Just Walk Out and the humans behind the cameras
Business Strategy
· 4 min read

Amazon Just Walk Out and the humans behind the cameras

A 2024 report about Amazon Just Walk Out is circulating again on LinkedIn this week. The lesson for anyone buying AI automation still holds.

Amazon Just Walk Out AI automation
SpaceX Automates Last. You Do It First.
Business Strategy
· 5 min read

SpaceX Automates Last. You Do It First.

The method SpaceX uses to redesign rockets puts automation dead last, after questioning, deleting and simplifying. Most AI projects start where it ends.

AI automation SpaceX process redesign
Anthropic Deleted 80% of Claude Code's System Prompt
AI & Automation
· 4 min read

Anthropic Deleted 80% of Claude Code's System Prompt

Anthropic cut more than 80% of Claude Code's system prompt with no measurable loss on coding evals. The rules you added last year are now the ceiling.

Anthropic Claude Code context engineering