Skip to main content
A variety of microphones can be used for speech to text workflows. Optimal configuration of such devices depends on environmental factors and use case. This page provides guidance of configuration best practice for common scenarios. Be sure to adapt for your local needs.
See audio format requirements for full details on allowable MIME types and supported audioFormat configuration.

Microphone Configuration

Dictation

Ambient Conversation

Maintain average loudness around –12 dBFS RMS with peaks near –3 dBFS for optimal speech-to-text normalization.

Channel Configuration

Choosing the right channel configuration ensures accurate transcription, speaker separation, and diarization across different use cases.
Mono input supports transcription with diarization; however, speaker separation may be unreliable when there is not clear turn-taking in the dialogue.Multichannel input (two audio channels, one per participant, in telehealth workflow) provides opportunity for improved speaker separation and labeling.

Streams Endpoint

Transcripts Endpoint


Audio streaming recommendations

Additional Notes

  • Enabling diarization is typically only required on mono audio.
  • Mono audio with diarization disabled will produce transcripts with one channel (-1), whereas diarized-mono transcripts will have two channels (0, 1).
  • For multichannel audio, each channel should capture only one speaker’s microphone feed in order to avoid cross-talk or echo between channels.
  • A maximum of 8 audio channels are supported in a given audio stream or recording.
  • Keep all channels aligned in time; do not trim or delay audio streams independently.
  • Ensure each channel contains only one participant’s feed to avoid duplicated transcript content.
  • Recommended capture format is 16-bit / 16 kHz PCM

Please contact us if you need more information about supported audio formats or are having issues processing an audio file.Additional references and resources: