Fish Audio
Hanabi AI Inc.
Expressive real-time speech synthesis and voice cloning from 15 seconds of audio, in 30+ languages.

Platforms
Overview
Fish Audio is a voice AI platform operated by Hanabi AI Inc. Its latest model, S2.1 Pro, is described as "the most expressive, emotionally controllable real-time voice model," with tags such as [laughing] and [whispering] to steer emotion and delivery. It covers voice cloning from 15 seconds of audio, speech recognition and APIs for voice agents, in more than 30 languages, with a public library of over 2 million voices. The free plan is for personal use only; commercial use requires a paid plan. The interface is also available in Japanese.
Main features
- Speech synthesis with emotion tags (S2.1 Pro)
- Voice cloning from 15 seconds of audio
- Speech to text
- APIs for voice agents
- More than 30 languages
- A library of over 2 million voices
- Japanese interface
Recommended use cases
Voice production
- Video voiceovers
- Audiobooks
- Character voices
- Voice chatbots
Pricing
Pricing not yet verified — check the official website.
Screenshots

Related AI tools
ElevenLabs
ElevenLabs is an AI voice generation service that can be used for narration and dubbing.
Cartesia
Real-time text-to-speech and speech-to-text APIs, aimed at building voice agents.
Murf AI
Murf AI is an AI voice generation service useful for narration and video production.
Hume AI
Voice AI APIs built around emotion: speech synthesis, a conversational voice interface and expression measurement.