One year.

$52M seed.

8M+ builders.

Read Our Story
Jul 27, 2026Company

5 Models, 22 People, 1 Year

5 Models, 22 People, 1 Year

Fish Audio raised $52M in seed funding. How a hobby project on a single 4090 became this, and what we're building next. Read our story.

Today marks another incredible milestone for Fish Audio. We’re announcing $52 million in seed funding on our first anniversary. Thank you to Coreline Ventures and Capital Today, who led the round, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphaist Partners, and the angels who put money in when there wasn't much to look at.

Text-to-speech has been good enough for about a decade, as long as you only need one sentence. The cracks show by the tenth. The whole category fails in the same place, for the same reason: getting the words right and getting the delivery right are different problems, and almost everyone has only ever worked on the first.

Years in voice AI at Amazon Alexa and Meta taught me where legacy TTS ends: a voice built to read never learns to perform. So we started from the other end. Fish Audio isn't a better voice assistant. It's the voice layer for everything after it.

Two company founders sitting side by side in wooden armchairs, facing the camera, with a floor-to-ceiling window and outdoor trees in the background.

The 4090

Shijia Liao, our chief scientist, reached the same conclusion from a different direction. As a research engineer at NVIDIA working on video, he spent his free time watching VTuber streams and couldn't get past how lifeless the voices sounded. He started training models on a single 4090 in his bedroom, making an early bet on a dual autoregressive architecture that the rest of the industry wouldn't catch up to for another two years. That work became the foundation of S2.

What We Did in 1 Year

  • Five state-of-the-art audio models shipped. Four text-to-speech, one speech-to-text.
  • Scaled the team from 3 to 22 people, and still hiring.
  • Zero to $21M in annual recurring revenue.
  • 8M+ creators, developers, and enterprises on the platform.
  • 2M+ voice models in the community library.
  • S2.1 Pro, preferred by 66% of listeners over leading competitors in blind listening tests.
  • 83+ languages with native cadence and word-level emotion control across 15,000+ natural language controls.

The Community Built the Model

We run post-training on real user preferences, which means the model gets better when people use it more. That's why we hold up on accented English, Mandarin, Japanese, Korean, and Spanish, where models trained mostly on standard American English fall apart. This ability came from people using the tool in languages, accents, and edge cases we'd never have thought to test.

Fish Speech went from a repo to a company because of game developers, anime fans, and people dubbing characters for fun. From the start, we’ve believed in transparency and openness, and making expressive, high-quality voice AI accessible to everyone.

And eight million of you trusted us to do so.

The Model Enterprises Now Trust

That trust is why enterprises are turning to Fish Audio too. Together with developers, enterprise is now two-thirds of our revenue.

We give them uptime, latency, and a security posture that survives procurement: on-premises deployment, zero-data-retention policies, and HIPAA-compliant configurations, built in from the start. S2.1 Pro already wins blind listening tests against the leading models, runs at 8,000+ tokens per second on a single H200, and holds up in languages where models tuned for one dialect quietly break.

Enterprise customers come to us because getting the delivery right, not just the words, is exactly the problem we set out to solve. It’s the same mission we’re now able to go further on.

What's Next?

Text-to-speech was never the whole goal. It was the part we could build first with a single GPU and prove the theory: delivery matters as much as words, and nobody was solving for it.

Now we go after the rest of the stack. Audio Understanding Language Model (Audio LM) and speech-to-speech. We're also doubling down on developer tooling and the API, expanding our enterprise team, and going deeper with partners like LiveKit and Retell to get expressive voice into more places, faster.

Free Through August

To celebrate this major milestone, S2.1 Pro is free for every developer via API through the end of August. This is the same model our enterprise customers pay for. We’re also offering 50% off creator plans and 3 months free if you're migrating from another provider.

The reason we can do this is technical, and our research team gets the credit. They rebuilt our entire inference stack this year: custom FP8 kernels, 8,000+ tokens per second on a single H200. That dropped our cost per request far enough that giving the model away became a successful strategy.

Thank You

One year ago, Fish Audio was a hobby project started from our living room. Everyone who trained a voice, dubbed a character, built on the API, or bet on us before there was much to bet on — this is yours as much as ours.

We got a lot of things wrong along the way. We're about to have the resources to get a lot more things right.

Here's to year two.

Rissa Cao

Rissa CaoX

Rissa is the CEO and co-founder of Fish Audio, pushing breakthroughs in AI voice technology. Find her latest work at @rissa_cao.

Read more from Rissa Cao

Create voices that feel real

Start generating the highest quality audio today.

Already have an account? Log in