
Fish Audio raises $50M seed for AI voice models

Fish Audio, a Palo Alto-based startup building AI voice models for creators and enterprises, announced a $50 million seed round led by Coreline Ventures and Capital Today. The company, founded by former NVIDIA researcher Shijia Liao, offers a library of more than 15,000 natural language controls for AI-generated voices. Since launching last year, it has attracted over 8 million users across its open-source and hosted models, generating $21 million in annual recurring revenue. Fish Audio‘s open-source repository, Fish Speech, has over 31,000 GitHub stars and is used by indie developers, video game designers, and creators.
The startup has released five models in the past year: four speech generation models and one speech-to-text model. Three of its speech generation models are open-sourced, while its latest S2.1 Pro model is available only via a paid API. The company offers monthly plans for creators, including voice cloning features, and an enterprise version used by organizations like HeyGen, Sanas, and Plaud. CEO and co-founder Rissa Cao explained that enterprises have diverse needs: HeyGen requires realism for AI avatars, gaming studios need expressive voices for characters, and voice agent companies like LiveKit want natural, low-latency voices for calls.
To build its voice library, Fish Audio encouraged users to submit their own voices for training and compensated them if used. This led to controversy when some creators alleged their voices were uploaded without consent. Fish Audio had a DMCA takedown process, but it was slow. Cao said the company has now automated takedowns, allowing creators to submit a voice sample or contract to prove ownership and have their voice removed in under three minutes. However, Oskue Honda, a partner at Coreline Ventures, noted that trust requires consent, transparency, and attribution to be built into the product, not treated as afterthoughts. He suggested the industry should move toward verified voice ownership, clear licensing, easy reporting, and revenue-sharing models.
Cao said the startup operated efficiently without outside funding when focused on open-source and creator plans, but sought capital to develop more advanced models and accommodate enterprises as investor interest grew. Future plans include releasing an audio understanding model and a speech-to-speech model this year. The speech generation market is crowded, with competitors like ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp. Rico Mallozzi of 359 Capital highlighted Fish Audio‘s fine-grained controls and cost-efficient training as advantages that help it compete with larger AI labs, praising the team’s ability to close the gap between artificial and human-like voices.


