Skip to main content

Fish Audio is 70% cheaper. ElevenLabs grew anyway

Fish Audio opened its weights and charges $0.05 a minute against $0.18. ElevenLabs still went from $350M to over $500M ARR. Where the moat actually moved.

Fish Audio is 70% cheaper. ElevenLabs grew anyway

Ricardo Argüello

Ricardo Argüello
Ricardo Argüello

CEO & Founder

Business Strategy 6 min read

Fish Audio opened the weights on its voice model and charges roughly $0.05 a minute. ElevenLabs charges about $0.18 for the same thing.

And ElevenLabs grew from $350 million ARR to over $500 million while that was happening.

Those two sentences sitting next to each other are the whole story.

The model became a commodity and the business did not move

We already argued that the harness is the moat and the model is now a commodity. What the voice market shows is that argument with financials attached: commoditizing the model does not commoditize the vendor. It relocates the value.

When the price of a component drops 70% and the leader keeps growing, the conclusion is not that buyers are stupid or that the open alternative is bad. It is that model price was never most of what they were buying.

This matters because nearly every voice AI evaluation I see begins and ends at cost per minute. It is the easiest number to find and the worst predictor of whether the deployment will survive contact with production.

The numbers on both sides

Fish Audio is not a weekend project, and that is precisely the point.

It published Fish Speech S2 on March 10, 2026 as an open text-to-speech model. It has passed 26,000 GitHub stars, ships emotion control through natural language tags, multi-speaker generation in a single pass, and latency under 150 milliseconds. It raised a $52 million seed round with $21 million of ARR already running, per Open Source For You. The pricing is aggressive: a free tier with 7 minutes of S2, $11 a month for 200 minutes, $75 a month for 27 hours.

It is exactly the dual-track play Mistral and Qwen already ran: open weights for community adoption, a paid API for enterprise, each side funding the other. Remio frames it accurately as an explicit bet against voice AI’s closed giants.

On the other side, ElevenLabs closed 2025 at $350 million ARR and passed $500 million in the first four months of 2026. It did not grow despite the open competitor. It grew during.

What an enterprise voice buyer is actually paying for

Here is the part that shows up on no comparison table and decides whether the project survives.

Latency under real load, not in the demo. A model that answers in 150 milliseconds on a single request and 900 with three hundred concurrent ones is useless for a contact center, and that curve is published nowhere.

Interruption handling. A real voice conversation has someone cutting in mid-sentence. Solving that well is engineering around the model, not the model.

Voice licensing. If you are using a voice that sounds like a specific person, you need to know who signed what. That legal exposure does not download with the weights.

Logging and traceability. What was said, on which model version, and whether you can reconstruct it six months later when a customer disputes it.

And the connection to the systems where the work happens. A voice agent that cannot read real order status is an expensive answering machine.

None of those five improve because the model got cheaper. All of them cost engineering, and that engineering is what is being paid for.

There is also a risk that lands differently in voice than in text, and it is worth saying even if it makes those of us who like open weights uncomfortable. An open voice model cannot be recalled. If a consent problem surfaces tomorrow around a particular voice, a closed vendor has something it can switch off. Weights already downloaded by thousands of people have no such button. That is not an argument against using them. It is an argument for making sure the voice you run in production has an origin you can document, and for never cloning a real person’s voice without written permission that would survive an audit.

The same movie, one vertical at a time

What makes this case useful is that it is not specific to voice.

The day a vertical hits its commoditization moment, model price stops being a defensible advantage for everyone in it. It already happened with text. It is happening with voice. It will happen with video, with specialized transcription, with document vision. We wrote about the same shift from the frontier open-weights angle this week when Alibaba opened Qwen3.8-Max.

The practical lesson repeats identically across all of them. If your product is the model, your margin has an expiration date. If your product is what surrounds the model, a falling model price makes your input cheaper.

We put numbers on that before in the research measuring the seven components of the harness. The operating conclusion has not changed: the advantage sits in the layer almost nobody wants to build, because it does not show up in a demo.

How we evaluate a voice deployment

If you are about to put voice into customer service, scheduling, or collections, this is the order we work in, and the model comes last.

First we define the specific conversation and its acceptable success rate. Not “handle calls,” but “confirm an existing appointment and reschedule it on request, resolving 80% without a transfer.”

Then we measure the system that already exists. If order status takes four seconds to look up, no model is going to sound natural, because the silence will be coming from your database.

Then we test under concurrency that resembles reality, not a single demo call.

And only then do we compare vendors, with per-minute price as one criterion among several, alongside voice licensing, traceability, and what happens the day you want to switch.

That last point is worth double right now. An open model like Fish Speech S2 gives you a real exit if the closed vendor raises prices, but only if you built the integration layer so the model is swappable. Build it welded to one API and the cheap open alternative does you no good at all.

So when is self-hosting the open model actually the right call? When three things are true at once: you have enough volume that the per-minute gap covers the cost of running it, you have a data or latency reason to want the model close, and you already have somebody who answers when the service falls over on a Sunday. Miss any one of those and the closed API is cheaper than the table says, because you were leaving the rest of the cost out. It is the same arithmetic as open weights at frontier scale, and the answer changes company by company, not model by model.

A cheaper input is good news. You just have to be built to collect on it.

Let’s build the layer that makes the model swappable

Frequently Asked Questions

voice AI ElevenLabs Fish Audio open models moat enterprise AI integration

Related Articles

Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Business Strategy
· 8 min read

Qwen3.8-Max redesigned a chip from 8,298 gates to 678

Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.

Qwen Alibaba open weights
The model is a commodity. Governance is the moat.
AI & Automation
· 5 min read

The model is a commodity. Governance is the moat.

Enterprise AI does not fail because the model cannot reason. It fails because nobody owns the control tower: who approves what, under which policy.

AI governance AI agents enterprise AI
Building costs zero. Distribution is the new investment.
Business Strategy
· 11 min read

Building costs zero. Distribution is the new investment.

Aaron Levie, Gergely Orosz, and Eric Siu published the same thesis in 36 hours: building got commoditized. Owned distribution loops are the 2026 moat.

distribution strategy moat Aaron Levie