> ## Documentation Index
> Fetch the complete documentation index at: https://docs.corti.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Audio Configuration

> Learn about best practices for microphone and channel configuration

A [variety of microphones](/stt/microphones) can be used for speech to text workflows. Optimal configuration of such devices depends on environmental factors and use case. This page provides guidance of configuration best practice for common scenarios. Be sure to adapt for your local needs.

<Tip>
  See [audio format requirements](/stt/audio) for full details on **allowable MIME types**  and supported `audioFormat` configuration.
</Tip>

***

## Microphone Configuration

### Dictation

| Setting          | Recommendation | Rationale                                                                                                                                                                                                                                                                                                                                                                         |
| :--------------- | :------------- | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| echoCancellation | Off            | Ensure clear, unfiltered audio from near-field recording.                                                                                                                                                                                                                                                                                                                         |
| autoGainControl  | Off            | Manual calibration of microphone gain level provides optimal support for consistent dictation patterns (i.e., microphone placement and speaking pattern). Recalibrate when dictation environments change (e.g., moving from a quiet to noisy environment). Recommend setting input gain with average loudness around –12 dBFS RMS (peaks near –3 dBFS) to prevent audio clipping. |
| noiseSuppression | Mild (-15dB)   | Removes background noise (e.g., HVAC); adjust as needed to optimize for your environment.                                                                                                                                                                                                                                                                                         |

### Ambient Conversation

| Setting          | Recommendation | Rationale                                                                                                                        |
| :--------------- | :------------- | :------------------------------------------------------------------------------------------------------------------------------- |
| echoCancellation | On             | Suppresses "echo" audio that is being played by your device speaker, e.g. remote call participant's voice + system alert sounds. |
| autoGainControl  | On             | Adaptive correction of input gain to support varying loudness and speaking patterns of conversational audio.                     |
| noiseSuppression | Mild (-15dB)   | Removes background noise (e.g., HVAC); adjust as needed to optimize for your environment.                                        |

<Check>Maintain average loudness around –12 dBFS RMS with peaks near –3 dBFS for optimal speech-to-text normalization.</Check>

***

## Channel Configuration

Choosing the right channel configuration ensures accurate transcription, speaker separation, and diarization across different use cases.

| Audio type               | Workflow                                          | Rationale                                                                                                                                                                                                                            |
| :----------------------- | :------------------------------------------------ | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Mono                     | Dictation or in-room doctor/patient conversation  | Speech to text models expect a single coherent input source. Mono also reduces bandwidth and file size without affecting accuracy.                                                                                                   |
| Multichannel (dual mono) | Telehealth or remote doctor/patient conversations | Assigns each participant to a dedicated channel, allowing the speech to text system to perform accurate speaker attribution. Provides better control over noise suppression and improves transcription accuracy when voices overlap. |

<Tip>
  Mono input supports transcription with diarization; however, speaker separation may be unreliable when there is not clear turn-taking in the dialogue.

  Multichannel input (two audio channels, one per participant, in telehealth workflow) provides opportunity for improved speaker separation and labeling.
</Tip>

<br />

### Streams Endpoint

<AccordionGroup>
  <Accordion title="Mono Audio Stream with Diarization">
    ```json highlight={6-10} theme={null}
    {
    "type": "config",
    "configuration": {
        "transcription": {
        "primaryLanguage": "en",
        "diarize": true,
        "isMultichannel": false,
        "participants": [
            {"channel": 0, "role": "multiple"}
          ]
        },
        "mode": {
        "type": "facts",
        "outputLocale": "en"
        }
      }
    }
    ```
  </Accordion>

  <Accordion title="Multichannel Audio Stream with Participants">
    ```json highlight={6-11} theme={null}
    {
    "type": "config",
    "configuration": {
        "transcription": {
        "primaryLanguage": "en",
        "diarize": false,
        "isMultichannel": true,
        "participants": [
            {"channel": 0, "role": "Doctor"},
            {"channel": 1, "role": "Patient"}
          ]
        },
        "mode": {
        "type": "facts",
        "outputLocale": "en"
        }
      }
    }
    ```
  </Accordion>

  <Accordion title="Diarization Disabled">
    ```json highlight={6-8} theme={null}
    {
    "type": "config",
    "configuration": {
        "transcription": {
        "primaryLanguage": "en",
        "diarize": false,
        "isMultichannel": false,
        "participants": []
        },
        "mode": {
        "type": "facts",
        "outputLocale": "en"
        }
      }
    }
    ```
  </Accordion>
</AccordionGroup>

### Transcripts Endpoint

<AccordionGroup>
  <Accordion title="Mono Audio File with Diarization">
    ```json highlight={5-9} theme={null}
    {
    "recordingId": "uuid",
    "primaryLanguage": "en",
    "spokenPunctuation": true,
    "isMultichannel": false,
    "diarize": true,
    "participants": [
        {"channel": 0, "role": "multiple"}
      ]
    }
    ```
  </Accordion>

  <Accordion title="Multichannel Audio File with Participants">
    ```json highlight={5-10} theme={null}
    {
    "recordingId": "uuid",
    "primaryLanguage": "en",
    "spokenPunctuation": true,
    "isMultichannel": true,
    "diarize": false,
    "participants": [
        {"channel": 0, "role": "doctor"},
        {"channel": 1, "role": "patient"}
      ]
    }
    ```
  </Accordion>

  <Accordion title="Audio File Diarization Disabled">
    ```json highlight={5-7} theme={null}
    {
    "recordingId": "uuid",
    "primaryLanguage": "en",
    "spokenPunctuation": true,
    "isMultichannel": false,
    "diarize": false,
    "participants": []
    }
    ```
  </Accordion>
</AccordionGroup>

***

## Audio streaming recommendations

|                                          |                                                                                                                                                                                                                                                        |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Sample rate 16 kHz**                   | Captures the full range of human speech frequencies, with **higher rates offering negligible recognition benefit** but increasing computational cost                                                                                                   |
| **Audio chunk size of 250 milliseconds** | Optimal speed to support both dictation and AI scribing workflows, with **sending much smaller chunks more frequently can degrade recognition accuracy** without improving latency                                                                     |
| **Stream at real-time speed**            | Audio should be streamed at or near real-time speed. **Streaming audio faster than real time is not recommended** and may cause buffering issues, degraded results, or stream termination. Pace audio chunks according to their actual audio duration. |

## Additional Notes

* Enabling diarization is typically only required on mono audio.
* Mono audio with diarization disabled will produce transcripts with one channel (-1), whereas diarized-mono transcripts will have two channels (0, 1).
* For multichannel audio, each channel should capture only one speaker’s microphone feed in order to avoid cross-talk or echo between channels.
* A maximum of 8 audio channels are supported in a given audio stream or recording.
* Keep all channels aligned in time; do not trim or delay audio streams independently.
* Ensure each channel contains only one participant’s feed to avoid duplicated transcript content.
* Recommended capture format is 16-bit / 16 kHz PCM

<br />

<Note>
  Please [contact us](mailto:help@corti.ai) if you need more information about supported audio formats or are having issues processing an audio file.

  Additional references and resources:

  * [Wikipedia - Audio file format](https://en.wikipedia.org/wiki/Audio_file_format)
  * [Wikipedia - Audio bit depth](https://en.wikipedia.org/wiki/Audio_bit_depth)
  * [Hugging Face - Introduction to audio data](https://huggingface.co/learn/audio-course/chapter1/audio_data)
</Note>
