For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending `.md` to the URL, or by setting the `Accept` header to `text/markdown`.
Skip to main content
Speech to TextAvailability

On-prem availability

Pre-recorded and streaming transcription run on-prem, as a container deployment or a virtual appliance. This page covers on-prem availability. For Speechmatics-hosted deployments, see Feature availability.

A check mark marks an available combination. An em dash means the combination is either unavailable or not established, and carries no roadmap position either way.

Region does not apply on-prem. Regions are an attribute of SaaS on Cloud, because on-prem deployments are customer-hosted.

Readiness on-prem​

On-prem ships as versioned containers, so the question is which release you need rather than whether a combination is production-ready.

Interaction patternModelOn-prem
pre-recordedStandardReleased
pre-recordedEnhancedReleased
pre-recordedMelia 1Released
streamingStandardReleased
streamingEnhancedReleased

Feature availability on-prem depends on the container release you are running. Check Accessing images for the current releases, and the release notes for the change that introduced a given feature.

Pre-recorded transcription features on-prem​

Pre-recorded transcription runs on the Standard, Enhanced, and Melia 1 models on-prem.

Language coverage and selection​

ItemStandardEnhancedMelia 1
56 languages✓✓✓
Mixed-language transcription——✓
Language hints——✓
Language labeling——✓
Transcription language packs (including bilingual)✓✓—
Automatic language identification✓✓—

Output and formatting​

ItemStandardEnhancedMelia 1
Output locale✓✓✓
Smart formatting✓✓✓
Punctuation and casing✓✓✓
Word-level timings✓✓✓
Segment-level timings✓✓✓
Confidence scores✓✓—

Transcript content and tagging​

ItemStandardEnhancedMelia 1
Custom dictionary✓✓—
Medical domain—✓—
Entity detection (basic, legacy)✓✓—
Disfluency tagging✓✓—
Profanity tagging✓✓—
Text replacement (find and replace)✓✓—

Speakers and channels​

ItemStandardEnhancedMelia 1
Speaker diarization✓✓✓
Channel diarization✓✓✓
Speaker identification✓✓—

Audio input​

ItemStandardEnhancedMelia 1
Audio events✓✓—
Audio filtering (volume filtering)✓✓—
Fetch URL✓✓✓

Audio events on-prem requires a GPU container.

Operational​

ItemStandardEnhancedMelia 1
Notifications✓✓✓

Add-ons​

ItemStandardEnhancedMelia 1
Translation✓✓—
Sentiment✓✓—

Translation on-prem runs in a separate inference container. See Translation GPU inference container.

Streaming transcription features on-prem​

Streaming transcription runs on the Standard and Enhanced models on-prem.

Streaming language coverage and selection​

ItemStandardEnhanced
56 languages✓✓
Transcription language packs (including bilingual)✓✓

Streaming output and formatting​

ItemStandardEnhanced
Output locale✓✓
Smart formatting✓✓
Punctuation and casing✓✓
Word-level timings✓✓
Confidence scores✓✓

Streaming transcript content and tagging​

ItemStandardEnhanced
Custom dictionary✓✓
Medical domain—✓
Entity detection (basic, legacy)✓✓
Disfluency tagging✓✓
Profanity tagging✓✓
Text replacement (find and replace)✓✓

Streaming speakers and channels​

ItemStandardEnhanced
Speaker diarization✓✓
Channel diarization✓✓
Speaker identification✓✓

Streaming audio input​

ItemStandardEnhanced
Audio events✓✓
Audio filtering (volume filtering)✓✓

Responding​

ItemStandardEnhanced
Force end of utterance✓✓
Turn detection✓✓
Word-level partials✓✓

Streaming add-ons​

ItemStandardEnhanced
Translation✓✓