Skip to main content

AI Realtime Voice

Realtime voice lets people talk to a WinnerWare AI assistant and hear it talk back, like a phone call. There is no typing and no waiting for a text reply. The assistant listens, answers aloud, and stops when the person talks over it.

This page explains the two ways WinnerWare delivers realtime voice, how to choose between them, how ElevenLabs fits in, how to set up each one, and what to do when something does not work.

Overview​

A realtime assistant needs to do three jobs on every turn:

  1. Hear the person and turn their speech into words.
  2. Think of a reply, using the profile's instructions, tools, and knowledge.
  3. Speak the reply.

WinnerWare can do these three jobs in two ways:

WayWhat it isIn one sentence
Native realtimeOne speech-to-speech model does all three jobs.The model listens and speaks by itself.
Cascaded realtimeThree separate deployments do one job each: speech-to-text (STT), a normal chat model, and text-to-speech (TTS).WinnerWare chains three services into one voice.

STT (speech-to-text) is a service that writes down what a person says. TTS (text-to-speech) is a service that reads text aloud in a spoken voice.

Realtime is used through the Conversation chat mode. On a profile set to Conversation, the Conversation deployment is the model that speaks, and the Chat deployment stays the text model that answers typed messages. Both native and cascaded realtime deployments appear in the same Conversation deployment list, and the people who use the assistant see the same Start speaking button either way. People can type and talk in the same chat.

Key Features​

  • Holds a live spoken conversation in the front-end chat, in external chat widgets, and in Chat Interactions.
  • Keeps typed and spoken turns in one thread. People can end the voice session and continue to type.
  • Lets people interrupt the assistant by talking over it.
  • Keeps the profile's instructions, tools, and knowledge in voice conversations.
  • Saves spoken turns to the session history, like typed turns.
  • Supports native speech-to-speech models from OpenAI and Azure OpenAI.
  • Supports Cascaded Realtime deployments, so you can pair any chat model with the listening and speaking services you prefer, for example ElevenLabs.
  • Lets each listener set their own microphone, speaker, volume, language, interruptions, and push-to-talk from a gear button.

How It Works​

Native realtime​

A native realtime deployment uses a speech-to-speech model, such as OpenAI's gpt-realtime models on OpenAI or Azure OpenAI. The model receives the person's voice directly and answers with its own voice.

  • One service, less delay. The audio goes to one model and comes back from the same model, so the reply starts quickly.
  • The model decides when you have finished. It listens for the end of a thought, so a short pause in the middle of a question does not make it answer too soon.
  • The model's own voices. You choose from the provider's fixed voice list. For OpenAI and Azure OpenAI these are alloy, ash, ballad, cedar, coral, echo, marin, sage, shimmer, and verse.
  • Knowledge is looked up on every spoken turn. When the profile has a knowledge source, WinnerWare searches it with what the person said, then gives the results to the model before it answers. If the search is slow, the assistant says a short line such as "let me look that up" while it waits.

Cascaded realtime​

"Cascaded" means the work flows down a chain, one step after another. A Cascaded Realtime deployment has no model of its own. It names three other deployments and chains them together:

microphone -> speech-to-text -> chat model -> text-to-speech -> speaker

On each turn:

  1. The speech-to-text deployment listens all the time and writes down what the person says. When it decides the person has finished speaking, it sends the finished sentence on.
  2. The chat deployment writes the reply. This is a normal text chat model. The profile's instructions, tools, and knowledge are applied here, so a voice assistant can do everything a text assistant can.
  3. The text-to-speech deployment speaks the reply, one sentence at a time. The first sentence starts to play while the chat model is still writing the rest, so people do not wait for the whole answer. Very short pieces are joined to the next piece so the voice does not sound choppy.

Some details to know:

  • Interruptions. As soon as the speech-to-text service hears the person talking, the reply that is playing stops and the unplayed audio is dropped. The part the person already heard stays in the conversation history.
  • One turn at a time. A new question waits until the previous reply has finished or been stopped.
  • Voices come from the speaking deployment. The Voice list on the profile shows the voices of the text-to-speech deployment in the cascade.
  • More delay than native. WinnerWare shows this on the deployment editor: "A reply passes through three services, so expect more delay before the assistant starts speaking than with a provider's own speech-to-speech model."

Native realtime compared with cascaded realtime​

Native realtimeCascaded realtime
What you set upOne deployment with Realtime (speech-to-speech) checked.Three deployments (speech-to-text, chat, text-to-speech) plus one Cascaded Realtime deployment that names them.
Voice quality and choiceThe provider's fixed voice list.Any voice from the speaking provider. With ElevenLabs, that is the voice library in your ElevenLabs account, including custom and cloned voices.
Delay before the assistant speaksLowest. One service hears and speaks.Higher. Each reply passes through three services.
Which modelsSpeech-to-speech models only, such as gpt-realtime on OpenAI or Azure OpenAI.Any chat model WinnerWare can use, including models that have no realtime version.
Tools and knowledgeRun on the realtime model. The realtime deployment needs Tool calling checked. Knowledge is looked up on every turn before the model answers.Run on the chat deployment in the cascade. That chat deployment needs Tool calling checked.
InterruptionsThe model hears speech over its own audio and stops.The first partial words from the speech-to-text service stop the reply.
End of a turnDecided by the model, which listens for the end of a thought.Decided by the speech-to-text service, based on a pause in speech.
What you pay forThe realtime model's audio usage, from one provider.Three bills: audio minutes for speech-to-text, text tokens for the chat model, and characters for text-to-speech. Check each provider's price list.
Choose it whenSpeed and natural turn-taking matter most, and the provider's voices are good enough.You want a specific voice or brand voice, a chat model that has no realtime version, or separate control over each step.

How ElevenLabs fits​

The Speech Services (Eleven Labs) module adds ElevenLabs as an AI provider. ElevenLabs can fill two of the three steps of a cascade:

  • Listening (speech-to-text). ElevenLabs realtime speech-to-text writes down the person's speech while they talk. It only transcribes. It never writes or speaks a reply, so it cannot be a realtime assistant by itself. Use it only as the speech-to-text step of a cascade.
  • Speaking (text-to-speech). ElevenLabs text-to-speech reads the reply aloud with a voice from your ElevenLabs account. WinnerWare asks ElevenLabs for raw audio at the rate the voice session uses, which is the format a cascade needs.

ElevenLabs does not provide a chat model. The middle step of the cascade must be a chat deployment from another provider, such as OpenAI or Azure OpenAI.

Which ElevenLabs models to use​

WinnerWare does not show a list of ElevenLabs models. You type the ElevenLabs model ID in the deployment's Model name box, and WinnerWare sends it to ElevenLabs unchanged. ElevenLabs adds and retires models, so confirm the current IDs in your ElevenLabs account.

StepModel IDNotes
Speech-to-text (realtime)scribe_v2_realtimeElevenLabs' realtime transcription model (Scribe v2 Realtime) at the time of writing. The realtime step needs a realtime model. The standard scribe_v1 and scribe_v2 models are for recorded audio, not live speech.
Text-to-speecheleven_flash_v2_5Flash. The fastest ElevenLabs voice model and the best choice for live conversation. Supports many languages.
Text-to-speecheleven_turbo_v2_5Turbo. Low delay, with a balance of speed and sound quality.
Text-to-speecheleven_multilingual_v2Multilingual. Very natural sound, but slower to start. Better for recorded prompts than for live conversation.
Text-to-speecheleven_v3Eleven v3. The most expressive voice model, and the slowest. Not recommended for live conversation.
StepChoose
Speech-to-textElevenLabs, Model name scribe_v2_realtime
ChatA fast chat model, for example a "mini" model on OpenAI or Azure OpenAI. Do not use a reasoning model: it thinks before it answers, which adds a long pause in a spoken conversation.
Text-to-speechElevenLabs, Model name eleven_flash_v2_5

Keep the profile's tool list short and its instructions focused on short, spoken answers. Every tool and every long answer adds delay.

ElevenLabs is not the only choice for the listening and speaking steps. Any speech-to-text deployment with Realtime (speech-to-speech) checked can listen, and any text-to-speech deployment that can produce raw audio can speak. Azure AI Services speech deployments can also be the speaking step.

How the Conversation mode picks its voice​

A profile in the Conversation chat mode uses the first of these that is available:

  1. The Conversation deployment on the profile.
  2. The site's default realtime deployment. The admin has no field for this value today, so it is set only when your hosting team sets it outside the admin.
  3. The first deployment that has Realtime (speech-to-speech) checked.

If the Conversation deployment on the profile is not a realtime deployment, WinnerWare does not use a different realtime model in its place.

When no realtime deployment is available, the Conversation mode falls back, in this order:

FallbackWhat it needsWhat it does
Speech-to-text plus text-to-speechA speech-to-text and a text-to-speech deploymentRecords what the person says, transcribes it, sends it to the Chat deployment as a normal text message, and reads the reply aloud. The button reads Start Conversation instead of Start speaking. It is slower than realtime, and it has no Voice settings gear.
Audio inputA speech-to-text deployment onlyThe person dictates with the microphone button, and the reply is text.
Text onlyNothingTyping only.

This fallback is not the same as a Cascaded Realtime deployment. A cascade is a realtime deployment, so it gets Start speaking, interruptions, and the Voice settings gear.

Configuration​

Before you start​

Make sure these are in place:

  • The AI Chat features are enabled, and you have an AI Profile or can create a Chat Interaction. See AI Chat.
  • A connection for your chat provider, such as OpenAI or Azure OpenAI, on Artificial Intelligence -> Provider Connections. See AI Deployments.
  • For ElevenLabs: the Speech Services (Eleven Labs) feature is enabled. The ElevenLabs deployment editor has no API key box. Your hosting team adds the ElevenLabs API key (and, if you want one, a default voice) to the site configuration, in the CloudSolutions_ElevenLabs_Speech section. See Speech ElevenLabs.
  • The site uses HTTPS, and the people who will use the assistant have a microphone and speakers or a headset.

The Default Deployments tab under Settings -> Artificial Intelligence has no default realtime deployment. Pick the realtime deployment as the Conversation deployment on each profile or Chat Interaction. If you leave it empty, WinnerWare uses the first deployment that has Realtime (speech-to-speech) checked.

Set up native realtime​

  1. Open Artificial Intelligence -> Deployments and select Add Deployment.
  2. In Available Providers, choose OpenAI or Azure OpenAI.
  3. In Model name, enter the realtime model or Azure deployment name, for example gpt-realtime.
  4. In Connection name, choose the connection.
  5. On the Model capabilities card, check Realtime (speech-to-speech). Keep Tool calling checked if the assistant needs tools or knowledge search.
  6. Select Save.
  7. Turn on realtime on a profile or Chat Interaction (see below).

Set up cascaded realtime with ElevenLabs​

You create four deployments: one for each step, and one Cascaded Realtime deployment that chains them. New deployments start with Text conversation, Tool calling, and Streaming checked, so clear the boxes that do not apply as you go. If you leave Text conversation checked on an ElevenLabs deployment, it appears in chat deployment lists where it cannot work.

1. The listening deployment (ElevenLabs speech-to-text)

  1. Open Artificial Intelligence -> Deployments and select Add Deployment.
  2. In Available Providers, choose ElevenLabs.
  3. In Model name, enter the ElevenLabs realtime transcription model, for example scribe_v2_realtime.
  4. On the Model capabilities card, check Realtime (speech-to-speech). The cascade only offers deployments with this box checked as the listening step, because only those can transcribe live speech. Clear Text conversation, Tool calling, and Streaming.
  5. Select Save.

2. The speaking deployment (ElevenLabs text-to-speech)

  1. Add another ElevenLabs deployment.
  2. In Model name, enter the ElevenLabs voice model, for example eleven_flash_v2_5.
  3. On the Model capabilities card, check Text to speech (synthesis). Clear Text conversation, Tool calling, and Streaming.
  4. Select Save.

3. The chat deployment

Use an existing chat deployment, or add one from OpenAI or Azure OpenAI. It must have Text conversation checked and Realtime (speech-to-speech) cleared. Check Tool calling if the assistant needs tools or knowledge. Check Streaming.

4. The Cascaded Realtime deployment

  1. Select Add Deployment. In Available Providers, choose Cascaded Realtime. The page title changes to New 'Cascaded Realtime' deployment.

  2. In Model name, enter a short label such as voice-cascade. WinnerWare does not send this value to any provider.

  3. On the Cascaded realtime card, choose:

    OptionWhat to chooseWhich deployments the list shows
    Speech-to-text deploymentThe ElevenLabs listening deployment.Deployments with Realtime (speech-to-speech) checked. Other cascaded deployments are not shown.
    Chat deploymentThe chat deployment that writes replies.Deployments that can hold a text conversation and are not realtime.
    Text-to-speech deploymentThe ElevenLabs speaking deployment.Deployments with Text to speech (synthesis) checked.
  4. Select Save. WinnerWare checks Realtime (speech-to-speech) on this deployment for you, because a cascade is realtime by design. All three steps are required.

Turn on realtime on an AI Profile​

  1. Open Artificial Intelligence -> Profiles and open the profile.
  2. On the Deployments & Interactions tab, keep a text model as the Chat deployment. Its hint reads "The text model this profile talks to. It answers typed messages, including those typed during a voice conversation." Realtime deployments are not in this list.
  3. Set Chat mode to Conversation. This option appears when a speech-to-text deployment or a realtime deployment exists.
  4. In Conversation deployment, choose your native realtime deployment or your Cascaded Realtime deployment. Use the site's default realtime deployment uses the order in How the Conversation mode picks its voice.
  5. In Voice, choose a voice:
    • For native realtime, the list shows the model's voices.
    • For a cascade, the list shows the voices of its text-to-speech deployment. With ElevenLabs, these are the voices in your ElevenLabs account.
    • Default voice uses the provider's default. For an ElevenLabs cascade, this is the default voice in the site configuration. If none is configured, the assistant cannot speak, so choose a voice.
  6. Select Save.

Every place that uses this profile now offers voice: the front-end chat, external chat widgets, and the admin chat.

A profile that was set up before the Conversation mode, with a realtime model as its Chat deployment, still speaks. The editor shows that model as the Conversation deployment and the chat mode as Conversation. Choose a text model as its Chat deployment so that typed messages get an answer from the model you want.

Turn on realtime in a Chat Interaction​

  1. Open Settings -> Artificial Intelligence, select the Chat Interactions tab, and set Chat mode to Conversation. This option needs a speech-to-text and a text-to-speech deployment.
  2. Open the Chat Interaction.
  3. In the settings panel, set Conversation deployment to a realtime deployment. This field appears only when the site Chat mode is Conversation.
  4. In Voice, choose a voice, or keep Default voice.

Usage​

What people see​

In the Conversation chat mode, a soundwave button is next to the message box. Its name is Start speaking when a realtime deployment is used, and Start Conversation when the speech-to-text plus text-to-speech fallback is used.

  1. Select the soundwave button. The browser asks for microphone access the first time. The status line shows Waiting for microphone access…, then Connecting…, then Listening.
  2. Talk normally. The assistant answers aloud, and the spoken words appear in the conversation as text.
  3. Talk over the assistant to interrupt it, if Allow interruptions is on.
  4. Select End Conversation to stop. The message box comes back, and the person can continue to type in the same thread.

While a voice session is live:

  • The message box and Send are hidden, and End Conversation takes their place.
  • A line below says "Speaking. Use headphones to avoid echo — end the conversation to go back to typing."
  • With a realtime deployment, the dictation microphone is hidden, because the voice session already uses the microphone.
  • If a typed message is sent, the voice session ends first. The typed message then goes to the Chat deployment as a normal text turn.

A voice session can also stop by itself. The chat then shows one of these messages, with a Resume button:

  • "Voice paused after a period of silence."
  • "Voice paused — this session reached its time limit. You can start another."
  • "Voice session ended — the connection dropped."
  • "Voice session ended — the microphone was disconnected."
  • "Voice session ended unexpectedly."

This works the same in the front-end chat and in external chat widgets. See AI Chat for how to embed a voice widget on another site.

Voice settings (gear button)​

With a realtime deployment, a gear button next to End Conversation opens Voice settings. The gear shows only while the voice session is live. Each person's choices are saved in their own browser, so they do not affect anyone else. These labels are always in English.

SettingWhat it does
MicrophoneChooses the microphone. Default microphone uses the system default.
SpeakerChooses where the assistant's voice plays. Automatic uses the device the system uses for calls. Select Choose speaker… to pick another device. If you cannot hear the assistant, choose the speakers you are actually listening to.
Assistant volumeSets how loud the assistant is, from 0% to 100%.
LanguageAutomatic lets the service detect the language. Choose a language (English, Spanish, French, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic, Hindi, or Russian) to fix it. In a cascade, this language goes to both the listening and the speaking steps.
Allow interruptionsLets people talk over the assistant to stop it. Turn it off if the assistant keeps hearing itself through the speakers.
Push-to-talkThe person holds Space, or the button, to talk. Use it in very noisy places.

Tips for a good voice assistant​

  • Keep the tool list short. A voice assistant chooses tools well from a small, clear list. Turn on only the few capabilities it needs. See AI Capabilities.
  • Ask for short, spoken answers in the instructions. Lists, tables, and links do not work well when read aloud.
  • Confirm before changes. Speech can be misheard, especially names and numbers. Tell the assistant to repeat details back before it saves anything.

Troubleshooting​

SymptomLikely causeFix
There is no Realtime option in the Chat mode list.This is expected. Realtime is not a chat mode.Set Chat mode to Conversation, and set the Conversation deployment to a realtime deployment.
There is no Conversation option in the Chat mode list.No speech-to-text deployment and no realtime deployment exists. For Chat Interactions, the site Chat mode needs a speech-to-text and a text-to-speech deployment.Add a realtime deployment, or the speech deployments. See AI Deployments.
No realtime deployment appears in the Conversation deployment list.No deployment has Realtime (speech-to-speech) checked.Check the box on a native realtime deployment, or create a Cascaded Realtime deployment. See AI Deployments.
Cascaded Realtime is missing from Available Providers.The AI deployment features are not enabled, or you do not have the Manage AI deployments permission.Ask your administrator to check the features and your role.
A step list on the Cascaded realtime card is empty.No deployment has the capability that step needs.Listening step: check Realtime (speech-to-speech) on the ElevenLabs speech-to-text deployment. Speaking step: check Text to speech (synthesis) on the text-to-speech deployment. Chat step: use a text chat deployment that is not realtime.
The session fails as soon as it starts, with a message that ElevenLabs realtime sessions "cannot generate replies".The ElevenLabs speech-to-text deployment was picked directly as the Conversation deployment. It only transcribes.Pick the Cascaded Realtime deployment as the Conversation deployment instead.
The assistant shows Listening but never answers.The listening step has the wrong Model name (for example a recorded-audio model instead of a realtime model), the microphone is muted or on the wrong device, or the chat step failed.Check the ElevenLabs speech-to-text Model name. Check Microphone in Voice settings. Test the chat deployment in a normal text chat.
The assistant hears and answers in text, but there is no voice.No voice is chosen and no default ElevenLabs voice is configured, the speaking step has the wrong Model name, the browser blocked audio, or Speaker points at the wrong device.Choose a Voice on the profile. Check the text-to-speech Model name (for example eleven_flash_v2_5). If the status says "Audio is blocked by the browser — click the page to enable it", click the page. Check Speaker and Assistant volume in Voice settings.
An error says the text-to-speech deployment "answered with" an audio format the session cannot play.The speaking provider cannot produce the raw audio a cascade needs.Use a text-to-speech deployment that can, such as ElevenLabs or Azure AI Services.
Long pauses before the assistant speaks.A cascade passes each reply through three services. A slow chat model, a reasoning model, many tools, a knowledge search, or long answers add more delay.Use native realtime if speed matters most. For a cascade, use eleven_flash_v2_5, a fast chat model without Reasoning, a short tool list, and instructions that ask for short answers.
The assistant answers itself or keeps stopping mid-sentence.Its own voice from the speakers reaches the microphone.Use a headset, lower the volume, or turn off Allow interruptions. In noisy places, turn on Push-to-talk.
The microphone is blocked in an external widget, and the browser never asks.The host site does not allow the microphone for the chat frame, or its own permissions policy is stricter than WinnerWare's.Keep microphone in the frame's allow value, make sure the host site's permissions policy allows the microphone, and register the host domain on the widget. See AI Chat.
The browser asks for the microphone, but the person selected Block.The browser remembers the choice.The person must allow the microphone for the site in the browser's site settings, then reload the page.
The assistant answers in the wrong language, or writes down the wrong words.Language in Voice settings is set to a different language, or Automatic guessed wrong.Set Language to the language the person speaks. Tell the assistant in the profile instructions which language to answer in. For ElevenLabs, use a multilingual voice model such as eleven_flash_v2_5.
The voice assistant talks but does not use its capabilities or knowledge.Tool calling is not checked on the deployment that writes the reply.Check Tool calling on the native realtime deployment, or on the chat step of the cascade.
The button reads Start Conversation, not Start speaking, and there is no gear.No realtime deployment was found, so the chat uses the speech-to-text plus text-to-speech fallback.Set the Conversation deployment to a realtime deployment. See How the Conversation mode picks its voice.
Typed messages get no answer, or an answer from an unexpected model.Typed messages go to the Chat deployment, not to the Conversation deployment.Choose a text model as the Chat deployment.

Operational Notes​

  • Audio travels over a direct browser connection (WebRTC) when the network allows it. When it does not, for example on a network that blocks this traffic, WinnerWare changes to a standard web connection by itself. People do not need to do anything.
  • A cascade depends on three providers. If one of them is slow or down, the whole voice assistant is affected. Check each provider's status when a cascade stops working.
  • The ElevenLabs voice list is cached for about 15 minutes. A new voice in your ElevenLabs account can take that long to appear in the Voice list.
  • The Agent Trainer simulator does not use realtime voice. Its Audio mode uses push-to-talk with the default speech-to-text and text-to-speech deployments. See AI Agent Trainer.
  • Plan for people who decline the microphone prompt or work in shared spaces. They can still type in the same chat.