For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending `.md` to the URL, or by setting the `Accept` header to `text/markdown`.
Skip to main content
Speech to TextAvailability

Feature availability

Availability depends on the combination of interaction pattern, model, and deployment, so check the combination rather than the product name. This page covers SaaS on Cloud. For container and virtual appliance deployments, see On-prem availability.

A check mark marks an available combination. An em dash means the combination is either unavailable or not established, and carries no roadmap position either way.

Readiness by combination

Readiness is a separate question from availability: availability says whether something exists for a combination, readiness says whether that combination is usable and documentable. Readiness is a property of the combination, so it is stated here once and not repeated on individual feature pages.

Interaction patternModelSaaS on CloudOn-prem
pre-recordedStandardGAReleased
pre-recordedEnhancedGAReleased
pre-recordedMelia 1GAReleased
streamingStandardGAReleased
streamingEnhancedGAReleased
streamingMelia 1Preview
agent STTLinden 1Preview

Preview applies to SaaS on Cloud only, because on-prem ships as versioned containers: the question there is which release you need.

Streaming with Melia 1, and agent STT with Linden 1, are available for evaluation and feedback. They are not production-ready and not ready to scale.

Pre-recorded transcription features

Pre-recorded transcription runs on the Standard, Enhanced, and Melia 1 models, in the EU, US, and AUS regions.

Language coverage and selection

ItemStandardEnhancedMelia 1
56 languages
Mixed-language transcription
Language hints
Language labeling
Transcription language packs (including bilingual)
Automatic language identification

Output and formatting

ItemStandardEnhancedMelia 1
Output locale
Smart formatting
Punctuation and casing
Word-level timings
Segment-level timings
Confidence scores

Transcript content and tagging

ItemStandardEnhancedMelia 1
Custom dictionary
Medical domain
Entity detection (basic, legacy)
Disfluency tagging
Profanity tagging
Text replacement (find and replace)

Speakers and channels

ItemStandardEnhancedMelia 1
Speaker diarization
Channel diarization
Speaker identification

Audio input

ItemStandardEnhancedMelia 1
Audio events
Audio filtering (volume filtering)
Fetch URL

Operational

ItemStandardEnhancedMelia 1
App usage tracking
Notifications

Add-ons

Add-ons are separate products that produce an output derived from a completed transcript. They are selected in addition to transcription.

ItemStandardEnhancedMelia 1
Translation
Chapters
Topics
Summaries
Sentiment
Audio alignment

Audio alignment is available to Enterprise customers only.

Pre-recorded regions

RegionStandardEnhancedMelia 1
EU
US
AUS

Streaming transcription features

Streaming transcription runs on the Standard, Enhanced, and Melia 1 models. Streaming is available in the EU and US regions, and is not available in the AUS region. Melia 1 for streaming is in Preview.

Streaming language coverage and selection

ItemStandardEnhancedMelia 1
56 languages
Mixed-language transcription
Language labeling
Transcription language packs (including bilingual)

Streaming output and formatting

ItemStandardEnhancedMelia 1
Output locale
Smart formatting
Punctuation and casing
Word-level timings
Confidence scores

Streaming transcript content and tagging

ItemStandardEnhancedMelia 1
Custom dictionary
Medical domain
Entity detection (basic, legacy)
Disfluency tagging
Profanity tagging
Text replacement (find and replace)

Streaming speakers and channels

ItemStandardEnhancedMelia 1
Speaker diarization
Channel diarization
Speaker identification

Streaming audio input

ItemStandardEnhancedMelia 1
Audio events
Audio filtering (volume filtering)

Responding

These features control when the Realtime API returns a result, and how much of the transcript each result contains.

ItemStandardEnhancedMelia 1
Force end of utterance
Turn detection
Word-level partials

Streaming operational and add-ons

ItemStandardEnhancedMelia 1
App usage tracking
Translation

Agent STT features

Agent STT runs on the Linden 1 model in the EU and US regions, and is in Preview. Because agent STT offers a single model, the following lists name what Linden 1 supports rather than comparing columns.

Language and output

  • 56 languages
  • Transcription language packs (including bilingual)
  • Output locale
  • Smart formatting
  • Punctuation and casing
  • Segment-level timings

Transcript content and tagging

  • Custom dictionary
  • Medical domain
  • Entity detection (basic, legacy)
  • Text replacement (find and replace)

Speakers

  • Speaker diarization
  • Speaker identification

Responding

  • Turn detection
  • Voice activity detection (VAD)
  • Force end of utterance
  • Segment-level partials

Audio input

  • Audio filtering (volume filtering)

Agent STT regions

Agent STT processes audio in the EU and US regions.