Luke Oliff.

Fish Audio's $52M Seed: Open Weights Got Them Here

·TTS·9 min read·Luke Oliff

Fish Audio announced a $52 million seed round on July 28, 2026, led by Coreline Ventures and Capital Today, exactly one year after the company launched. The Palo Alto startup reports $21 million in annual recurring revenue and more than 8 million users, and shipped a new flagship model, S2.1 Pro, alongside the round. The most telling detail sits in the release strategy: after open-sourcing three of its four speech-generation models, Fish Audio is keeping S2.1 Pro closed, available only through its paid API.

That last part is the story I want to write about, because I called Fish’s S2 Pro the default self-hosting recommendation in my open TTS roundup last week, and the successor to that model will never appear on a list like it.

From a single gaming GPU to a $52M seed in one year: Fish Speech open-sourced (July 2025, 31k+ GitHub stars), $21M ARR and 8M users built on open distribution, then S2.1 Pro launches closed on July 28, 2026.

How did Fish Audio grow this fast?

The origin story is genuinely good, which is presumably why the PR leans on it so hard. Shijia Liao, a former NVIDIA video researcher and self-described VTuber and anime fan, got annoyed enough at flat synthetic voices to train his own model on a single gaming GPU, then open-sourced it. That project, Fish Speech, now has more than 31,000 GitHub stars and became the on-ramp for indie developers, game designers, and creators who eventually turned into paying customers of the hosted platform.

One year in, the numbers Fish reports are $21 million ARR, 8 million users across open-source and hosted versions, and a production customer list that includes HeyGen, LiveKit, Retell, Sanas, OpenArt, and Telnyx. Those are company-stated figures, not audited ones, but the customer names are real and they sit at usefully different points in the voice stack. HeyGen wants realism for avatars, Retell wants cheap per-minute economics for phone agents, LiveKit wants latency. If you can serve all three, you have a platform rather than a demo.

The round itself is a seed in name only. $52 million at $21 million ARR, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphalist Partners, is a category bet on Fish becoming the default voice layer, and the stated use of funds (voice-native LLMs, speech-to-speech, enterprise sales) reads like a company planning to fight well above the TTS SKU it launched with.

What does S2.1 Pro actually claim?

The launch claims are aggressive and worth listing precisely, because every one of them is company-reported. Voice cloning from a 5-second clip in roughly 15 seconds. Support for 83+ languages. Word-level control over emotion and pacing through more than 15,000 natural language tags. Twice as fast as Cartesia. One sixth the cost of ElevenLabs, with third-party writeups putting S2.1 Pro at $15 per million characters against ElevenLabs’ $60 to $165. And a blind listening test where 67% of listeners preferred it over leading competitors.

I have no reason to assume bad faith on any of these, and no reason to repeat them as fact either. Vendor blind tests are marketing until you run your own scripts through them, latency multipliers depend entirely on region and concurrency, and “1/6th the cost” comparisons always pick the competitor tier that flatters the ratio. What I will say is that the free API window (the s2.1-pro-free model string, running through August 31, 2026, fair use, no SLA) exists precisely so you can check, and checking costs you an afternoon. Fish also published an inference engineering post claiming they cut serving cost roughly 4x, which at least suggests the free tier is structural rather than pure investor subsidy.

Where does it sit against the top of the board? Here’s the part the launch material doesn’t mention: S2.1 Pro is already on the blind-vote Speech Arena, listed since June, and as of July 29 it sits 16th at 1,138 Elo from 1,582 votes. That’s a real improvement over the open S2 Pro at 1,123, and 91 points behind Simba 3.2, which holds first place at 1,229. A 91-point Elo gap is one listeners can hear, and it sits awkwardly next to a 67% blind-preference claim over “leading competitors”. Both can be true if the comparison set was priced like ElevenLabs (1,173 Elo at $100 per million) rather than ranked like the leaders. Which is exactly why you run your own scripts.

Why is S2.1 Pro closed when Fish Speech was open?

Because open weights were the customer acquisition strategy, and the strategy worked. Fish shipped five models in its first year, four speech-generation and one speech-to-text, and open-sourced three of the speech models. Each open release compounded the GitHub gravity, pulled in self-hosters, and built exactly the developer trust that a one-year-old company cannot buy with ads. Then the model good enough to anchor a $52 million round arrived, and it is API-only.

Mistral ran this play in text. Meta arguably ran a version of it before pulling back. The pattern is consistent enough now to treat as the default lifecycle of an open-weights AI company: open until the frontier is yours, then closed where the margin lives. I noted in last week’s open TTS piece that open speech models trail the closed leaders by 100+ Elo, and Fish’s decision fits that gap. The odd part is how small the defended delta currently is: 15 arena points between the closed S2.1 Pro and the open S2 Pro. Either Fish is protecting the model because the roadmap built on it matters more than this checkpoint, or the arena hasn’t caught up with what the model does off-leaderboard (emotion control and cloning don’t show up in blind voice votes). Probably some of both.

None of this is a criticism, to be clear. Fish’s open models remain available, self-hosting remains a real path at roughly $1 to $5 per million characters in GPU time, and a company that gave the community 31,000 stars’ worth of genuinely usable speech tech owes nobody its frontier weights. But if your architecture assumes the best Fish model will always be self-hostable, yesterday ended that assumption, and the roadmap (voice-native LLMs, speech-to-speech) will live behind the API too.

What should TTS buyers do with this news?

Test it, during the free window, on your own scripts and your own languages. That’s the whole recommendation. The eval that matters is 20 production scripts in your top two languages, cloned from both 5-second and 30-second references, measured for time-to-first-audio at your real concurrency, priced at your actual monthly character volume. Fish’s enterprise pitch includes a hook that if they can’t cut your voice AI costs by 50% they’ll give you a year free, which is a sales wager you can only call with a documented baseline invoice.

The wider signal is the one I keep writing about: the price floor for good speech keeps falling. Google cut decent speech by 70% in June, open weights put self-hosting in the $1 to $5 range, and now a well-funded lab is pushing the hosted market at $15 per million characters with a free-until-September API. At Speechify we publish Simba 3.2 at $10 per million characters list and $6 at volume, first on the arena as of this week, so I’m comfortable with where that fight goes on quality per dollar. But every vendor in this market, mine included, now operates on the assumption that a credible competitor can appear from a bedroom GPU inside twelve months. That’s the actual lesson of Fish Audio’s first year, and $52 million says the market believes it.

FAQ

What is Fish Audio’s S2.1 Pro?

S2.1 Pro is Fish Audio’s flagship text-to-speech model, publicly launched July 28, 2026 alongside the company’s $52 million seed round. It claims voice cloning from 5 seconds of audio, 83+ language support, and word-level emotion control via 15,000+ natural language tags, and scores 1,138 Elo on the independent Speech Arena. Unlike Fish’s earlier S2-family models, it is not open weights: it is available only through the paid API, with a free evaluation tier (s2.1-pro-free) running until August 31, 2026.

How much does Fish Audio S2.1 Pro cost compared to other TTS APIs?

S2.1 Pro is listed at $15 per million characters on the Speech Arena’s pricing column, against ElevenLabs’ Eleven v3 at $100 per million. Speechify’s Simba 3.2, first on the blind-vote Speech Arena at 1,229 Elo, lists at $10 per million characters and $6 at volume on the Scale tier. Self-hosting Fish’s older open S2 Pro costs roughly $1 to $5 per million in GPU time, excluding engineering overhead.

Is Fish Audio open source?

Partly, and decreasingly. The original Fish Speech project is open source with more than 31,000 GitHub stars, and Fish has open-sourced three of its four speech-generation models. Its newest and best model, S2.1 Pro, is closed and API-only. The open S2-family checkpoints remain available for self-hosting where the licence allows, but the frontier of Fish’s catalogue now sits behind the paid API.

Is S2.1 Pro better than Simba 3.2?

By independent blind votes, no. As of July 29, 2026, the Speech Arena has S2.1 Pro 16th at 1,138 Elo from 1,582 votes, while Simba 3.2 holds first at 1,229, a 91-point gap listeners can hear. Fish’s own 67% blind-test preference claim was run against an undisclosed competitor set. On price, Simba 3.2 also lists lower: $10 per million characters ($6 at volume) against S2.1 Pro’s $15, though S2.1 Pro is free to test until August 31, 2026.

Who uses Fish Audio in production?

Fish Audio’s named customers include HeyGen (AI avatars), LiveKit (real-time infrastructure), Retell (phone agents), Sanas (accent technology), OpenArt, and Telnyx. The company reports 8 million total users across its open-source and hosted products, and offers on-premises deployment, zero-data-retention configurations, and HIPAA-compliant setups for regulated industries.