Cartesia
Cartesia: Sonic AI voices for narration and live applications
Cartesia turns text into speech with Sonic 3.6, offering 44 languages, streaming audio and custom voices. Its browser Playground and speech APIs support narration, interactive apps and voice agents.
Free plan availableAccount required

- Access
- Free and paid plans
- Example plan
- USD 5 / monthPro
- Available on
- Web, API
A speech platform for interactive products
Cartesia provides a voice layer for applications, narration and conversational agents. Sonic 3.6 is its current text-to-speech model, with 44 languages and a stable snapshot released in August 2026. The browser Playground provides voice selection and delivery controls, while developers can integrate speech through the API. A dated model snapshot gives production teams a fixed version; the general Sonic 3.6 identifier follows stable updates.
Custom voices come in two forms. Instant cloning begins with a short reference recording and is available on Pro and above. Sonic 3.6 can use up to 60 seconds to preserve more of the speaker’s accent. Professional cloning requires at least 30 minutes of audio and a Startup or higher plan. It learns the recording’s pacing and loudness, so request-time speed and volume changes do not apply. Emotion guidance is an English-only beta feature rather than a universal control for all languages.
The API offers different delivery formats for different products: the Bytes endpoint can return MP3 or WAV files, while streaming endpoints provide raw audio. Credits are shared with Cartesia’s transcription models, and standard TTS is metered at approximately one credit per character. Managed voice agents use separate dollar-based billing. Commercial use begins on paid plans. Enterprise zero retention covers speech generation and transcription inference; cloning still requires retained source material. Standard terms permit model training unless an applicable opt-out or different agreement applies.
Best for
- Developers adding speech to interactive applications
- Teams building multilingual voice agents
- Creators generating narration with custom voices
Limitations
Speech generation and transcription share credits. Managed agent calls have separate charges. Professional clones do not support request-time speed or volume changes; emotion tags are English-only beta. Free output is noncommercial. Enterprise zero retention excludes cloning and voice creation.
Cartesia speech and voice features
Sonic 3.6 multilingual speech
Generate speech in 44 languages with stable model snapshots available for fixed production versions.
Streaming and file output
Use raw audio for SSE/WebSocket streaming or save MP3/WAV through the Bytes endpoint.
Instant custom voices
Pro and higher support short-reference cloning, with up to 60 seconds of source speech on Sonic 3.6.
Professional voice training
Startup and above support clones trained on 30+ minutes, with organisation-level slot limits.
Delivery controls
Guide speed, loudness and English emotional delivery, subject to the selected voice type.
Enterprise retention options
Enterprise zero retention can cover TTS/STT content; operational metadata and cloning material are outside that scope.
Cartesia pricing
Free plan available. Subscription plans. Visit Cartesia for full plan details and current offers.
Example plans.
Pro
USD 5 / month
Cartesia speech platform
Plan details
100,000 shared speech credits/month; commercial use and instant cloning. TTS is approximately 1 credit/character.
Startup
USD 49 / month
Cartesia speech platform
Plan details
1.25 million shared speech credits/month; two professional-clone slots per organisation.
Scale
USD 299 / month
Cartesia speech platform
Plan details
8 million shared speech credits/month; four professional-clone slots and 15 concurrent TTS requests.
Technical specifications
Voice generation
| Feature | Cartesia |
|---|---|
| Speech languages | Sonic 3.6 supports 44 languages, including Odia and Urdu. A cloned voice starts in its recorded language; additional native accents use the voice-localisation capability.Language coverage is model-specific |
| Voice and pronunciation controls | Speed and volume guide delivery; English emotion controls are beta. Instant clones support speed adjustment, while professional clones take pacing and loudness from their training recordings.Controls vary by voice type; emotion tags English only |
| Voice cloning | Instant cloning starts on Pro with a short reference clip. Professional cloning needs 30+ minutes and Startup or above, with two Startup or four Scale slots per organisation.Paid tiers; use your own voice or obtain explicit consent |
| Audio exports and streaming | The Bytes API returns MP3, WAV or raw audio; SSE/WebSocket streaming returns raw audio. Output sample-rate options span 8–48 kHz.Container support differs by endpoint |
| Commercial-use terms | Pro and higher include commercial use subject to content rights and acceptable-use rules. Standard terms permit model training unless an applicable opt-out or separate agreement applies.Paid commercial rights; Enterprise retention scope |
| Output sample rate (Hz) | 48000Maximum selectable output rate; not a voice-quality score |
| API access | SupportedBytes, SSE and WebSocket endpoints; account/API key required |
Cartesia alternatives
Other speech products offer browser production tools or a self-hosted model family.
ElevenLabs
Hosted speech and custom voices alongside a studio, dubbing and other audio tools.
Compare with CartesiaResemble AI (Chatterbox)
Open-source voice models for local deployment and self-managed infrastructure.
Compare with CartesiaMurf
A browser narration editor with pronunciation controls and separate developer API plans.
Compare with CartesiaCartesia FAQs
Is Cartesia free?
The Free plan includes 20,000 monthly speech credits for noncommercial use. Pro starts at USD 5/month and adds commercial rights and instant cloning. Standard TTS uses approximately one credit per character, with transcription sharing the credit balance.
What is the difference between instant and professional cloning?
Instant cloning uses a short reference recording and starts on Pro. Professional cloning needs at least 30 minutes and starts on Startup, with limited slots per organisation. Professional clones learn pacing and loudness from the recordings rather than request-time controls.
Can Cartesia generate MP3 or WAV audio?
Yes, through the TTS Bytes endpoint. SSE and WebSocket endpoints stream raw audio instead. The API offers output sample rates from 8 to 48 kHz.
Do Cartesia credits roll over?
Paid plans roll unused credits forward up to twice the monthly allocation. Optional overages are billed separately; with overages disabled, requests stop when the available allocation is exhausted. Managed agent usage is metered in dollars rather than speech credits.
Is this your product? Claim this page to update your listing.