For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending `.md` to the URL, or by setting the `Accept` header to `text/markdown`.
Skip to main content
Speech to TextAvailability

Feature availability

Availability depends on the combination of interaction pattern, model, and deployment, so check the combination rather than the product name. This page covers SaaS on Cloud. For container and virtual appliance deployments, see On-prem availability.

A check mark marks an available combination. An em dash means the combination is either unavailable or not established, and carries no roadmap position either way.

Readiness by combination​

Readiness is a separate question from availability: availability says whether something exists for a combination, readiness says whether that combination is usable and documentable. Readiness is a property of the combination, so it is stated here once and not repeated on individual feature pages.

Interaction patternModelSaaS on CloudOn-prem
pre-recordedStandardGAReleased
pre-recordedEnhancedGAReleased
pre-recordedMelia 1GAReleased
streamingStandardGAReleased
streamingEnhancedGAReleased
streamingMelia 1Preview—
agent STTLinden 1Preview—

Preview applies to SaaS on Cloud only, because on-prem ships as versioned containers: the question there is which release you need.

Streaming with Melia 1, and agent STT with Linden 1, are available for evaluation and feedback. They are not production-ready and not ready to scale.

Pre-recorded transcription features​

Pre-recorded transcription runs on the Standard, Enhanced, and Melia 1 models, in the EU, US, and AUS regions.

Language coverage and selection​

ItemStandardEnhancedMelia 1
56 languages✓✓✓
Mixed-language transcription——✓
Language hints——✓
Language labeling——✓
Transcription language packs (including bilingual)✓✓—
Automatic language identification✓✓—

Output and formatting​

ItemStandardEnhancedMelia 1
Output locale✓✓✓
Smart formatting✓✓✓
Punctuation and casing✓✓✓
Word-level timings✓✓✓
Segment-level timings✓✓✓
Confidence scores✓✓—

Transcript content and tagging​

ItemStandardEnhancedMelia 1
Custom dictionary✓✓—
Medical domain—✓—
Entity detection (basic, legacy)✓✓—
Disfluency tagging✓✓—
Profanity tagging✓✓—
Text replacement (find and replace)✓✓—

Speakers and channels​

ItemStandardEnhancedMelia 1
Speaker diarization✓✓✓
Channel diarization✓✓✓
Speaker identification✓✓—

Audio input​

ItemStandardEnhancedMelia 1
Audio events✓✓—
Audio filtering (volume filtering)✓✓—
Fetch URL✓✓✓

Operational​

ItemStandardEnhancedMelia 1
App usage tracking✓✓✓
Notifications✓✓✓

Add-ons​

Add-ons are separate products that produce an output derived from a completed transcript. They are selected in addition to transcription.

ItemStandardEnhancedMelia 1
Translation✓✓—
Chapters✓✓—
Topics✓✓—
Summaries✓✓—
Sentiment✓✓—
Audio alignment✓✓—

Audio alignment is available to Enterprise customers only.

Pre-recorded regions​

RegionStandardEnhancedMelia 1
EU✓✓✓
US✓✓✓
AUS✓✓—

Streaming transcription features​

Streaming transcription runs on the Standard, Enhanced, and Melia 1 models. Streaming is available in the EU and US regions, and is not available in the AUS region. Melia 1 for streaming is in Preview.

Streaming language coverage and selection​

ItemStandardEnhancedMelia 1
56 languages✓✓✓
Mixed-language transcription——✓
Language labeling——✓
Transcription language packs (including bilingual)✓✓—

Streaming output and formatting​

ItemStandardEnhancedMelia 1
Output locale✓✓✓
Smart formatting✓✓✓
Punctuation and casing✓✓✓
Word-level timings✓✓✓
Confidence scores✓✓—

Streaming transcript content and tagging​

ItemStandardEnhancedMelia 1
Custom dictionary✓✓—
Medical domain—✓—
Entity detection (basic, legacy)✓✓—
Disfluency tagging✓✓—
Profanity tagging✓✓—
Text replacement (find and replace)✓✓—

Streaming speakers and channels​

ItemStandardEnhancedMelia 1
Speaker diarization✓✓—
Channel diarization✓✓✓
Speaker identification✓✓—

Streaming audio input​

ItemStandardEnhancedMelia 1
Audio events✓✓—
Audio filtering (volume filtering)✓✓—

Responding​

These features control when the Realtime API returns a result, and how much of the transcript each result contains.

ItemStandardEnhancedMelia 1
Force end of utterance✓✓—
Turn detection✓✓—
Word-level partials✓✓✓

Streaming operational and add-ons​

ItemStandardEnhancedMelia 1
App usage tracking✓✓✓
Translation✓✓—

Agent STT features​

Agent STT runs on the Linden 1 model in the EU and US regions, and is in Preview. Because agent STT offers a single model, the following lists name what Linden 1 supports rather than comparing columns.

Language and output

  • 56 languages
  • Transcription language packs (including bilingual)
  • Output locale
  • Smart formatting
  • Punctuation and casing
  • Segment-level timings

Transcript content and tagging

  • Custom dictionary
  • Medical domain
  • Entity detection (basic, legacy)
  • Text replacement (find and replace)

Speakers

  • Speaker diarization
  • Speaker identification

Responding

  • Turn detection
  • Voice activity detection (VAD)
  • Force end of utterance
  • Segment-level partials

Audio input

  • Audio filtering (volume filtering)

Agent STT regions​

Agent STT processes audio in the EU and US regions.