Is FREE AI Text to Speech and Voice Cloning Finally Good? (Fish Audio)

Fish Audio’s S2 Pro is a versatile, open-source text-to-speech model offering advanced emotional controls, multilingual support, and efficient voice cloning with just seconds of reference audio, available for both local use and faster cloud API access. While local setup is straightforward and functional, the cloud API delivers superior speed and quality, making it ideal for real-time applications and developers seeking quick integration.

The video reviews Fish Audio’s S2 Pro, a leading text-to-speech (TTS) model available on Hugging Face and GitHub. This model stands out due to its fine-grained natural language control tags, allowing users to add emotions and speech effects like astonishment, laughter, or whispering. Fish Audio offers both a slow mode with 4 billion parameters and a super-fast mode with 400 million parameters, enabling a balance between quality and speed. The model supports 15,000 different control options and multiple languages, making it versatile for various applications. The inference code is open source, and users can run it locally or access it via free APIs currently available for a limited time.

Setting up the model locally is straightforward, with detailed installation instructions provided, including options for command line, web UI, and Docker environments. The model runs efficiently on Macs using the GPU and can also be run on Nvidia GPUs or CPU-only setups. The presenter demonstrates running the model locally on a MacBook Pro, showing how to generate speech with different emotional tags and multi-turn speaker support. The local generation, while functional, is slower compared to the cloud API, taking several minutes for complex prompts.

The cloud API, powered by H200 GPUs, offers significantly faster inference times, producing speech almost instantaneously. The presenter tests the API using curl commands and highlights the ease of obtaining a free API key. The API supports dynamic emotional speech and multi-speaker dialogues, with the cloud version handling some tags like laughter more accurately than the local version. This speed and quality make the cloud API suitable for real-time applications and developers looking for quick integration.

A notable feature of Fish Audio’s S2 Pro is its voice cloning capability, which requires only 5 to 10 seconds of reference audio. The presenter experiments with cloning his own voice, comparing local and cloud generation times. The cloud service produces a professional-grade voice clone rapidly, while local generation is much slower. The voice cloning results are impressive, capturing the nuances of the reference voice, although better samples can improve quality further. This feature is particularly useful for content creators needing to dub or recreate voices efficiently.

In conclusion, Fish Audio’s S2 Pro offers a powerful, flexible, and accessible TTS solution with advanced emotional controls, multilingual support, and voice cloning. While local usage is possible and useful for experimentation, the cloud API provides superior speed and quality, making it ideal for developers and real-time applications. The model’s open-source nature and extensive features position it as a strong contender in the AI TTS space, especially given the current free API access. The presenter encourages viewers to explore the model and share their experiences.