For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending `.md` to the URL, or by setting the `Accept` header to `text/markdown`.
Skip to main content
Speech to TextAvailability

On-prem availability

Pre-recorded and streaming transcription run on-prem, as a container deployment or a virtual appliance. This page covers on-prem availability. For Speechmatics-hosted deployments, see Feature availability.

A check mark marks an available combination. An em dash means the combination is either unavailable or not established, and carries no roadmap position either way.

Region does not apply on-prem. Regions are an attribute of SaaS on Cloud, because on-prem deployments are customer-hosted.

Readiness on-prem

On-prem ships as versioned containers, so the question is which release you need rather than whether a combination is production-ready.

Interaction patternModelOn-prem
pre-recordedStandardReleased
pre-recordedEnhancedReleased
pre-recordedMelia 1Released
streamingStandardReleased
streamingEnhancedReleased

Feature availability on-prem depends on the container release you are running. Check Accessing images for the current releases, and the release notes for the change that introduced a given feature.

Pre-recorded transcription features on-prem

Pre-recorded transcription runs on the Standard, Enhanced, and Melia 1 models on-prem.

Language coverage and selection

ItemStandardEnhancedMelia 1
56 languages
Mixed-language transcription
Language hints
Language labeling
Transcription language packs (including bilingual)
Automatic language identification

Output and formatting

ItemStandardEnhancedMelia 1
Output locale
Smart formatting
Punctuation and casing
Word-level timings
Segment-level timings
Confidence scores

Transcript content and tagging

ItemStandardEnhancedMelia 1
Custom dictionary
Medical domain
Entity detection (basic, legacy)
Disfluency tagging
Profanity tagging
Text replacement (find and replace)

Speakers and channels

ItemStandardEnhancedMelia 1
Speaker diarization
Channel diarization
Speaker identification

Audio input

ItemStandardEnhancedMelia 1
Audio events
Audio filtering (volume filtering)
Fetch URL

Audio events on-prem requires a GPU container.

Operational

ItemStandardEnhancedMelia 1
Notifications

Add-ons

ItemStandardEnhancedMelia 1
Translation
Sentiment

Translation on-prem runs in a separate inference container. See Translation GPU inference container.

Streaming transcription features on-prem

Streaming transcription runs on the Standard and Enhanced models on-prem.

Streaming language coverage and selection

ItemStandardEnhanced
56 languages
Transcription language packs (including bilingual)

Streaming output and formatting

ItemStandardEnhanced
Output locale
Smart formatting
Punctuation and casing
Word-level timings
Confidence scores

Streaming transcript content and tagging

ItemStandardEnhanced
Custom dictionary
Medical domain
Entity detection (basic, legacy)
Disfluency tagging
Profanity tagging
Text replacement (find and replace)

Streaming speakers and channels

ItemStandardEnhanced
Speaker diarization
Channel diarization
Speaker identification

Streaming audio input

ItemStandardEnhanced
Audio events
Audio filtering (volume filtering)

Responding

ItemStandardEnhanced
Force end of utterance
Turn detection
Word-level partials

Streaming add-ons

ItemStandardEnhanced
Translation