Voice & music
Cartesia
Real-time, expressive text-to-speech API built for voice agents.

Voice & music
Real-time, expressive text-to-speech API built for voice agents.
Product insights
Low-latency streaming voices that sound human in conversation.
Slow or robotic TTS breaks the feel of a live voice conversation.
Developers and companies building voice products.
Usage-based API pricing with a free tier; see the official pricing page.
Cartesia AI voice generator is a text-to-speech API. Its Sonic model streams audio fast enough for live conversations, which is why it is used inside customer-support bots, AI companions, and other voice agents.
Sign up and create a key in the dashboard.
Choose a voice and language in the playground.
Call the streaming API from your agent.
Sonic is Cartesia’s text-to-speech model; Sonic-3.6 is the current version.
Cartesia lists 44 languages for Sonic.
Official website snapshot
cartesia.ai
Cartesia provides Sonic, a streaming text-to-speech model now at Sonic-3.6. It generates natural, expressive voices, including laughter, in 44 languages with low latency, and is designed for AI agents and interactive apps.