This guide assumes a working Corti agent integration is already in place, for example the On-Demand Agent, Wake-Command Agent, or Conversational Agent pattern. It covers only the last step: turning that agent’s text response into audio.
messageSend. To close the loop into a voice interface, that text needs to become speech the user can hear. This guide walks through pairing your agent with an external text to speech provider: choosing one, sending it the agent’s response, and playing back the result.
Why this is a pairing, not a Corti endpoint
Text to speech models and hosting are widely available and improve quickly, including strong open-source options. Rather than adding a Corti-hosted proxy in front of a fast-moving space, we recommend connecting your agent integration directly to a text to speech provider. That gives you direct control over latency, voice quality, language coverage, and cost, and lets you change providers without waiting on a Corti release.Choose a text to speech provider
The right provider depends on your priorities. For example, providers with a streaming API (WebSocket or server-sent events) return the first audio chunk well before the full response finishes synthesizing. If low time-to-first-audio matters for your use case, prefer a provider’s streaming endpoint over its plain REST endpoint. A few vendors to consider include the following:Prompt the agent for speakable text
An agent tuned for a text-based chat interface will often return markdown formatting, bullet lists, or headers, none of which sound right read aloud. Add explicit instructions to the agent’ssystemPrompt so its responses are written to be heard, not read:
Speakable-response system prompt
Send the agent’s response to your provider
OncemessageSend resolves, extract the text from the response and hand it to your chosen provider’s text to speech endpoint. The exact request shape is provider-specific: consult that provider’s docs for authentication, voice selection, and output format. The pattern is the same regardless of provider:
JavaScript
speak is your own function, not part of the Corti API, that calls the chosen provider’s text to speech endpoint with replyText and plays back the resulting audio, for example through the Web Audio API or an <audio> element fed from a MediaSource. Match the playback method to the provider’s output format: compressed formats like mp3 or opus play fine through an <audio> element, but raw PCM output needs the Web Audio API instead.
Account for end-to-end latency
Time to first audio in this pattern is the sum of three stages: speech to text on the user’s utterance, the agent’smessageSend response time, and the provider’s text to speech synthesis time. Each stage adds up, so:
- Prefer providers with a streaming synthesis API to shorten the third stage.
- Keep agent responses concise; shorter text reaches the provider’s synthesis step sooner and produces less audio to wait for.
- Measure the full pipeline, not just the text to speech call in isolation. A fast provider can still feel slow if it sits behind a slow agent response.
Handle interruptions
Some text to speech providers offer their own barge-in control (for example, a streamingInterrupt or flush/cancel message that stops synthesis server-side). That’s worth using when it’s available, but not every provider offers it, and it only stops the provider’s audio, not your local playback. Detecting the interruption itself doesn’t have to depend on the provider at all: your existing Corti speech to text session can do it, so the same approach works regardless of which provider you paired with above.
The core idea: don’t stop listening while the agent’s response plays back. Keep the dictation session running through playback, and treat a new incoming transcript as the user talking over the agent.
1
Keep the microphone hot during playback
Don’t pause or close the dictation session while
speak is running. Track whether playback is active with a simple flag so you can distinguish “the agent is talking” from “the user is talking.”2
Turn on echo cancellation
With the speaker and microphone active at the same time, the agent’s own audio can loop back into the mic and be misread as user speech. Set
echoCancellation: true on the capture track, the same setting Corti recommends for ambient conversation capture whenever a device’s own speaker output could reach its microphone.3
Treat an interim transcript during playback as a barge-in
Listen for the dictation session’s
transcript event as usual. If an interim result arrives while isSpeaking is true, the user has started talking over the agent: stop audio playback immediately and cancel the in-flight text to speech request if your provider call is still streaming or generating.JavaScript
4
Feed the new turn through your normal flow
Once the barge-in is handled, the transcript that triggered it continues through the same debounce and turn-finalization logic from the Conversational Agent guide, so it’s sent to the agent as the next turn like any other utterance.
Fall back gracefully on provider failures
The text to speech provider is now a dependency your integration doesn’t control. Treat it as one that can fail independently of the agent call: wrap it with a timeout and a fallback so a slow or failed provider request doesn’t stall the conversation or leave the user with nothing.JavaScript
What this pattern doesn’t cover
This is a starting point, not a complete voice interface. It does not handle:- Long-running tasks. If the agent needs to do substantial work before responding, this pattern has nothing to say back to the user in the meantime.
- SSML or fine-grained prosody control. Handled entirely by your chosen provider, if it supports it.
Next steps
On-Demand Agent
The agent lifecycle and message API, for a single isolated pass.
Wake-Command Agent
Gate the agent behind a spoken wake phrase and hold a multi-turn thread.
Conversational Agent
Always-on voice agent, no wake phrase, every utterance forwarded with speculative prefetch.
Context & Memory
How conversation threads and memory work in the Agentic Framework.
Please contact us for help or questions.