Fish Audio is 70% cheaper. ElevenLabs grew anyway
Ricardo Argüello — August 16, 2026
CEO & Founder
General summary
Fish Audio released Fish Speech S2 in March 2026, an open text-to-speech model with over 26,000 GitHub stars and sub-150ms latency, charging roughly $0.05 per minute against ElevenLabs' $0.18. It raised $52 million on $21 million of ARR. Over the same stretch, ElevenLabs went from $350 million ARR at the end of 2025 to more than $500 million in the first four months of 2026. The model got commoditized. The business did not.
- Fish Audio released Fish Speech S2 on March 10, 2026 as an open text-to-speech model, with over 26,000 GitHub stars, emotion control through natural language tags, and latency under 150 milliseconds
- Estimated cost per minute is roughly $0.05 on Fish Audio against $0.18 on ElevenLabs, about 70% cheaper
- Fish Audio raised a $52 million seed round after reaching $21 million ARR, betting that open weights and self-hosting can compete against the closed leaders
- ElevenLabs ended 2025 at $350 million ARR and passed $500 million in the first four months of 2026, growing while the open substitute got cheaper
- The pattern repeats vertical by vertical: once the model becomes a commodity, value shifts to the integration, compliance, and operational-guarantee layer around it
Imagine that for years the coffee business depended on access to a bean almost nobody could source. One day that bean gets cheap and anyone can buy it. The sellers who only had the bean lose the business. The ones who also had the calibrated machine, the location, the delivery window, and the contract with the office next door keep charging the same. The bean stopped being the business without the business disappearing.
AI-generated summary
Fish Audio opened the weights on its voice model and charges roughly $0.05 a minute. ElevenLabs charges about $0.18 for the same thing.
And ElevenLabs grew from $350 million ARR to over $500 million while that was happening.
Those two sentences sitting next to each other are the whole story.
The model became a commodity and the business did not move
We already argued that the harness is the moat and the model is now a commodity. What the voice market shows is that argument with financials attached: commoditizing the model does not commoditize the vendor. It relocates the value.
When the price of a component drops 70% and the leader keeps growing, the conclusion is not that buyers are stupid or that the open alternative is bad. It is that model price was never most of what they were buying.
This matters because nearly every voice AI evaluation I see begins and ends at cost per minute. It is the easiest number to find and the worst predictor of whether the deployment will survive contact with production.
The numbers on both sides
Fish Audio is not a weekend project, and that is precisely the point.
It published Fish Speech S2 on March 10, 2026 as an open text-to-speech model. It has passed 26,000 GitHub stars, ships emotion control through natural language tags, multi-speaker generation in a single pass, and latency under 150 milliseconds. It raised a $52 million seed round with $21 million of ARR already running, per Open Source For You. The pricing is aggressive: a free tier with 7 minutes of S2, $11 a month for 200 minutes, $75 a month for 27 hours.
It is exactly the dual-track play Mistral and Qwen already ran: open weights for community adoption, a paid API for enterprise, each side funding the other. Remio frames it accurately as an explicit bet against voice AI’s closed giants.
On the other side, ElevenLabs closed 2025 at $350 million ARR and passed $500 million in the first four months of 2026. It did not grow despite the open competitor. It grew during.
What an enterprise voice buyer is actually paying for
Here is the part that shows up on no comparison table and decides whether the project survives.
Latency under real load, not in the demo. A model that answers in 150 milliseconds on a single request and 900 with three hundred concurrent ones is useless for a contact center, and that curve is published nowhere.
Interruption handling. A real voice conversation has someone cutting in mid-sentence. Solving that well is engineering around the model, not the model.
Voice licensing. If you are using a voice that sounds like a specific person, you need to know who signed what. That legal exposure does not download with the weights.
Logging and traceability. What was said, on which model version, and whether you can reconstruct it six months later when a customer disputes it.
And the connection to the systems where the work happens. A voice agent that cannot read real order status is an expensive answering machine.
None of those five improve because the model got cheaper. All of them cost engineering, and that engineering is what is being paid for.
There is also a risk that lands differently in voice than in text, and it is worth saying even if it makes those of us who like open weights uncomfortable. An open voice model cannot be recalled. If a consent problem surfaces tomorrow around a particular voice, a closed vendor has something it can switch off. Weights already downloaded by thousands of people have no such button. That is not an argument against using them. It is an argument for making sure the voice you run in production has an origin you can document, and for never cloning a real person’s voice without written permission that would survive an audit.
The same movie, one vertical at a time
What makes this case useful is that it is not specific to voice.
The day a vertical hits its commoditization moment, model price stops being a defensible advantage for everyone in it. It already happened with text. It is happening with voice. It will happen with video, with specialized transcription, with document vision. We wrote about the same shift from the frontier open-weights angle this week when Alibaba opened Qwen3.8-Max.
The practical lesson repeats identically across all of them. If your product is the model, your margin has an expiration date. If your product is what surrounds the model, a falling model price makes your input cheaper.
We put numbers on that before in the research measuring the seven components of the harness. The operating conclusion has not changed: the advantage sits in the layer almost nobody wants to build, because it does not show up in a demo.
How we evaluate a voice deployment
If you are about to put voice into customer service, scheduling, or collections, this is the order we work in, and the model comes last.
First we define the specific conversation and its acceptable success rate. Not “handle calls,” but “confirm an existing appointment and reschedule it on request, resolving 80% without a transfer.”
Then we measure the system that already exists. If order status takes four seconds to look up, no model is going to sound natural, because the silence will be coming from your database.
Then we test under concurrency that resembles reality, not a single demo call.
And only then do we compare vendors, with per-minute price as one criterion among several, alongside voice licensing, traceability, and what happens the day you want to switch.
That last point is worth double right now. An open model like Fish Speech S2 gives you a real exit if the closed vendor raises prices, but only if you built the integration layer so the model is swappable. Build it welded to one API and the cheap open alternative does you no good at all.
So when is self-hosting the open model actually the right call? When three things are true at once: you have enough volume that the per-minute gap covers the cost of running it, you have a data or latency reason to want the model close, and you already have somebody who answers when the service falls over on a Sunday. Miss any one of those and the closed API is cheaper than the table says, because you were leaving the rest of the cost out. It is the same arithmetic as open weights at frontier scale, and the answer changes company by company, not model by model.
A cheaper input is good news. You just have to be built to collect on it.
Let’s build the layer that makes the model swappableFrequently Asked Questions
Fish Audio is a speech synthesis company that releases its models openly, including Fish Speech S2, published March 10, 2026 with over 26,000 GitHub stars. Unlike ElevenLabs, which runs closed models behind an API, Fish Audio lets you download the weights and self-host.
Estimated cost per minute is roughly $0.05 on Fish Audio against $0.18 on ElevenLabs, about 70% cheaper. Fish Audio also offers a free tier with 7 minutes of S2 generation, an $11 per month plan with 200 minutes, and a $75 per month plan with 27 hours.
Because per-minute model price is not what an enterprise buyer is actually purchasing. ElevenLabs went from $350 million ARR at the end of 2025 to more than $500 million in the first four months of 2026, selling guaranteed latency, voice licensing, compliance, and integrations that already work.
It means picking a vendor on model price matters less and less. What separates a voice deployment that works is the layer around it: real latency under load, interruption handling, licensing of the voice used, audit logging, and connection to the systems where the work actually happens.
Related Articles
Qwen3.8-Max redesigned a chip from 8,298 gates to 678
Alibaba opened the weights of a 2.4T-parameter model that ran ~500 closed-loop chip design iterations alone. Your evaluation criteria just went stale.
The model is a commodity. Governance is the moat.
Enterprise AI does not fail because the model cannot reason. It fails because nobody owns the control tower: who approves what, under which policy.
Building costs zero. Distribution is the new investment.
Aaron Levie, Gergely Orosz, and Eric Siu published the same thesis in 36 hours: building got commoditized. Owned distribution loops are the 2026 moat.