Fish Audio, run by Hanabi AI, differs from most closed voice-cloning platforms by open-sourcing its underlying model, Fish Speech, on GitHub, letting developers inspect, self-host, or fine-tune the same technology the hosted product runs on rather than treating it as a black box. That openness has helped build an unusually large community voice library, over two million shared voices at last count.
Voice cloning needs only about ten seconds of reference audio to produce a usable clone, and emotion tags, marking a line as angry, whispering, or laughing, give more direct control over delivery than adjusting abstract sliders. Support for more than thirty languages and ultra-low-latency streaming round out a feature set aimed at both casual creators and developers building real-time voice applications.
A free tier covers limited generation, and paid plans start around $5 a month for more volume. For a developer who wants the option to self-host or inspect the underlying model rather than depend entirely on a closed API, Fish Audio's open-source foundation is a genuine point of difference in a category dominated by proprietary systems.








