Amazon Polly
Amazon Polly: multilingual text-to-speech on AWS
Amazon Polly converts text into speech for applications, narration and telephone systems. Its Standard, Neural, Generative and Long-form engines combine multilingual voices with pronunciation controls, streaming audio and reusable speech files.
Paid plansAccount required

- Access
- Paid access
- Usage pricing
- USD 4 / 1,000,000 charactersStandard
- Available on
- Web, API
Speech generation for applications and reusable narration
Amazon Polly provides the speech layer for software, spoken notifications, accessible content and narration. A developer supplies text, selects a compatible voice and engine, and receives audio. The AWS console offers a browser interface for previewing speech, while SDKs and the command-line tools support application integration. Generated files can be stored and replayed, making Polly useful for recurring prompts and published narration as well as speech created on demand.
Four engines serve different production needs and budgets: Standard, Neural, Generative and Long-form. Their voice catalogues and regional availability differ. Speech Synthesis Markup Language controls delivery, and pronunciation lexicons help with names, abbreviations and specialist terms. Standard voices support pitch changes; the other engines support rate and volume but not the same pitch attribute. Speech Marks supply word, sentence and mouth-shape timing for highlighting text or coordinating animation. These metadata requests have their own character charges.
The API supports compressed audio, PCM and telephony encodings, with sample rates determined by the format. Longer scripts can use asynchronous synthesis rather than the smaller synchronous request limit. Custom Brand Voice projects require a separate engagement with AWS and a voice actor. AWS states that Polly output belongs to the customer as between customer and AWS, while permissions for supplied text remain necessary. AWS may retain inputs for service and model improvement; organisations can use the documented AI services opt-out policy.
Best for
- Developers adding speech to AWS applications
- Publishers producing reusable audio narration
- Teams building telephone prompts and spoken interfaces
Limitations
Voice availability and SSML controls vary by engine and region. Standard pitch controls do not apply to Neural, Long-form or Generative voices. Speech Marks add request charges, and other AWS infrastructure costs are separate. Brand Voice requires a custom engagement. Free Tier eligibility depends on the account; input-data improvement use has an opt-out policy.
Amazon Polly speech features
Four speech engines
Standard, Neural, Generative and Long-form provide different voice catalogues and per-character rates.
Multilingual voices and accents
Choose supported languages and regional voices for narration, spoken interfaces and telephone prompts.
SSML and pronunciation lexicons
Control delivery and specialist pronunciation, with supported tags and attributes defined for each engine.
Speech timing metadata
Word, sentence and viseme Speech Marks support synchronized text highlighting and animation.
Files, streams and telephony audio
Generate MP3, Ogg, PCM or A-law/μ-law output and reuse stored audio without another synthesis request.
Custom Brand Voice projects
Work with AWS and a voice actor on a dedicated neural voice through a separately negotiated engagement.
Amazon Polly pricing
Paid plans. Usage-based pricing. Visit Amazon Polly for full plan details and current offers.
Example plans and usage rates; additional charges may apply.
Standard
USD 4 / 1,000,000 characters
Standard engine
Plan details
Per million billed input characters; SSML tags excluded. GovCloud uses different rates.
Neural
USD 16 / 1,000,000 characters
Neural engine
Plan details
Per million billed input characters; compatible voice required. GovCloud uses different rates.
Generative
USD 30 / 1,000,000 characters
Generative engine
Plan details
Per million billed input characters; available voices and regions vary.
Technical specifications
Voice generation
| Feature | Amazon Polly |
|---|---|
| Speech languages | Multilingual voices and regional accents, including bilingual English/Hindi voices. Available voices depend on the engine and AWS region; Polly speaks the supplied language rather than translating text.Engine and region availability vary |
| Voice and pronunciation controls | SSML and pronunciation lexicons guide speech. Standard supports pitch, rate and volume; Neural, Long-form and Generative support rate and volume, not pitch. Generative prosody wraps full sentences.SSML tags differ by engine |
| Voice cloning | Brand Voice is a custom engagement with AWS and a voice actor to create an exclusive neural voice. Pricing and production timeline are negotiated; it is separate from ordinary stock-voice synthesis.Contact AWS; custom engagement |
| Audio exports and streaming | MP3, Ogg Vorbis, Ogg Opus, mono PCM and A-law/μ-law audio; JSON Speech Marks provide timing metadata. MP3/Vorbis support up to 48 kHz, PCM up to 16 kHz and telephony formats 8 kHz.Output format determines sample-rate options |
| Commercial-use terms | Polly output belongs to the customer as between customer and AWS. Third-party input text requires appropriate rights. Replaying cached audio does not incur another synthesis charge.AWS customer agreement and input rights |
| Output sample rate (Hz) | 48000MP3/Ogg Vorbis maximum; Ogg Opus 48 kHz; not PCM or telephony |
| API access | SupportedAWS account and permissions; SDK, CLI and console access |
Amazon Polly alternatives
Other speech platforms offer managed voice APIs or browser narration studios.
Google Cloud Text-to-Speech
Cloud speech with Chirp voices, Gemini delivery prompts and model-specific billing.
Compare with Amazon PollyCartesia
Sonic speech APIs with streaming output and short-reference cloning on paid plans.
Compare with Amazon PollyMurf
A browser voiceover editor for scripts, pronunciation edits and production exports.
Compare with Amazon PollyAmazon Polly FAQs
How much does Amazon Polly cost?
Published USD rates are 4 per million characters for Standard, 16 for Neural, 30 for Generative and 100 for Long-form. SSML tags are excluded from billed characters. Speech Marks requests are billed separately, and GovCloud pricing differs.
Is Amazon Polly free?
Eligible new AWS accounts can apply Free Tier credits. Accounts created under the scheme introduced in July 2025 have time-limited credit eligibility; the free account plan ends after six months or when credits run out. Older account allowances differ. Polly is otherwise billed by usage.
Can I use Polly audio commercially?
AWS says Polly output belongs to the customer as between customer and AWS. You need rights to any third-party text supplied. Generated audio can be stored and replayed without another Polly synthesis charge, subject to the AWS agreement.
Does Amazon Polly clone voices?
Its stock voice API does not include a short-recording cloning tool. Brand Voice is a separate custom engagement in which AWS works with a voice actor to create an exclusive neural voice; pricing and timelines are negotiated.
Is this your product? Claim this page to update your listing.