AI Deployments
Every AI feature in WinnerWare talks to an AI model through a deployment. A deployment tells WinnerWare which model to use, which account to reach it through, and what that model can do. This page explains connections and deployments in plain terms, walks through every option on the deployment editor, and shows how the Model capabilities checkboxes decide what your assistants can and cannot do.
Overview
Two records work together:
| Record | What it is | Example |
|---|---|---|
| Connection | The account details WinnerWare uses to reach an AI provider: the address (endpoint) and how to sign in (key or managed identity). One connection can serve many deployments. | Your company's Azure OpenAI resource. |
| Deployment | One model, reached through a connection, plus a list of what that model can do. Profiles, chat interactions, and settings pick deployments, never connections. | gpt-4.1-mini on your Azure OpenAI resource, with tool calling and streaming checked. |
Where to find them in the admin menu:
| Screen | Menu path | Who can open it |
|---|---|---|
| AI Deployments | Artificial Intelligence -> Deployments | Roles with the Manage AI deployments permission. |
| AI Provider Connections | Artificial Intelligence -> Provider Connections | Roles with the Manage AI Provider Connections permission. The menu item appears only when the AI Connection Management feature is enabled. |
| Default deployments | Settings -> Artificial Intelligence, Default Deployments tab | Roles with the Manage AI profiles permission. |
Key Features
- Keeps provider account details in one place (the connection), so many deployments can share them.
- Describes each model by its capabilities, so WinnerWare only offers a deployment where it can actually do the job.
- Removes anything a model has not declared from every request, which avoids provider errors but also means an unchecked box silently turns a feature off.
- Lets you set a site-wide default deployment for each job (chat, utility, embedding, image, vision, speech-to-text, text-to-speech), and lets each profile or chat interaction override the chat and utility choice.
- Supports reasoning models, with a Reasoning effort setting you can narrow on the deployment and choose per profile or chat interaction.
- Supports realtime voice deployments, including a Cascaded Realtime deployment that chains three other deployments together.
How It Works
Connections hold the account details
A connection is the "how do we sign in" part. You create it once per provider account and then point deployments at it.
Connections come from one of two places:
- Created in the admin on Artificial Intelligence -> Provider Connections. You can edit or delete these.
- Configuration-based connections, set up by your hosting team outside the admin. They appear in a separate, locked list at the bottom of the screen and cannot be edited there.
Deployments describe one model
A deployment names the model, picks the connection, and declares the model's capabilities on the Model capabilities card. The capabilities are what WinnerWare uses to decide:
- Where the deployment is offered. For example, only deployments with Text embedding checked are offered as the embedding deployment.
- What is sent to the model. If a capability is not checked, WinnerWare strips the related options from every request before it reaches the provider. The most important example is Tool calling: with it unchecked, the model never receives any tools. See Tool calling.
The deployment list shows each deployment's capabilities as badges, next to the provider name and the connection (or Contained connection for deployments that carry their own sign-in details). This is the quickest place to check what a deployment declares.
Like connections, deployments can also be Configuration-based deployments. These appear in a locked list and have no Edit button.
Deployment roles
Each job in WinnerWare needs a deployment with a specific capability. A deployment can fill a role only when it declares that capability.
| Role | Used for | Capability the deployment must have |
|---|---|---|
| Chat | The main assistant reply in AI Profiles and Chat Interactions. | Text conversation, and not Realtime (speech-to-speech). |
| Utility | Lighter background work, such as generating session titles, planning, and intent detection. | Text conversation, and not Realtime (speech-to-speech). |
| Embedding | Turning text into search vectors for knowledge, documents, and memory indexes. | Text embedding |
| Image generation | Creating images. | Image output |
| Vision | Understanding images, for example image-aware Chat Interactions and describing pictures inside uploaded documents. | Image input (vision) |
| Speech to text | Turning a person's speech into text for the Audio input and Conversation chat modes. | Speech to text (transcription) |
| Text to speech | Reading the assistant's reply aloud. | Text to speech (synthesis) |
| Realtime | Spoken, back-and-forth voice conversations in the Conversation chat mode. It is picked as the Conversation deployment. | Realtime (speech-to-speech) |
A realtime deployment never fills the chat or utility role, even when Text conversation is also checked, because a speech-to-speech model cannot answer a normal text request.
How WinnerWare chooses a deployment
When a feature needs a deployment, WinnerWare works down this list and uses the first deployment that exists and has the required capability:
- The deployment picked on the profile or chat interaction. This applies to the Chat deployment and Utility deployment selectors.
- The site default for that role on Settings -> Artificial Intelligence -> Default Deployments.
- Utility only: if no utility deployment is found, WinnerWare uses the chat deployment instead: first the one picked on the profile or interaction, then the Default chat deployment.
- Last resort: the first deployment in the list that can fill the role.
A deployment that is picked but does not have the required capability is skipped without a message, and WinnerWare moves on to the next step. If a profile seems to use the "wrong" model, check the picked deployment's capabilities first.
The default deployments
Settings -> Artificial Intelligence, Default Deployments tab, holds one default per role. Each list only shows deployments that can fill that role.
| Setting | What it does |
|---|---|
| Default chat deployment | Used for chat sessions and chat interactions whose Chat deployment is set to Default. |
| Default utility deployment | Used for background tasks such as planning and intent detection when a profile or interaction does not pick its own. |
| Default embedding deployment | Used to build and search knowledge, document, and memory indexes. |
| Default image deployment | Used for image generation. |
| Default vision deployment | Used for image-aware chat when no specific vision-capable deployment is configured. |
| Default speech-to-text deployment | Used to transcribe speech. On a profile, Audio input appears only when a speech-to-text deployment is available, and Conversation appears when a speech-to-text or a realtime deployment is available. |
| Default text-to-speech deployment | Used to read replies aloud. |
| Default text-to-speech voice | The voice used when a profile does not pick one. The list loads the voices of the selected text-to-speech deployment. Leave it on Default voice to use the provider's default. |
The tab has no default realtime deployment. A profile or chat interaction in the Conversation chat mode uses its own Conversation deployment, or the first deployment that has Realtime (speech-to-speech) checked. See Realtime and voice deployments.
If the Default chat deployment or Default utility deployment is empty, the profile and chat interaction editors show a warning above their deployment selectors.
Configuration
Add a connection
- Open Artificial Intelligence -> Provider Connections and select Add Connection.
- In Available Providers, choose the provider.
- Fill in the fields below and select Save.
| Option | What it does |
|---|---|
| Title | The display name people see in lists and in deployment editors. |
| Technical name | The unique name used in settings and recipes. You cannot change it after you save. |
| Endpoint | The provider's web address, for example your Azure OpenAI resource URL. |
| Authentication type | Azure providers only. Default authentication uses the sign-in method set up by your hosting team, Managed identity uses the server's Azure identity, and API key uses a key you enter. |
| API key | The provider's secret key. After you save, the key is stored securely and the box stays blank. Leave it blank to keep the stored key, or type a new one to replace it. |
| Identity client ID | Shown for Managed identity. The client ID of a user-assigned managed identity. Leave it blank to use the system-assigned identity. |
Add a deployment
- Open Artificial Intelligence -> Deployments and select Add Deployment.
- In Available Providers, choose the provider that hosts the model (for example Azure OpenAI). The page title changes to New 'Azure OpenAI' deployment.
- Fill in the fields and the Model capabilities card, then select Save.
Deployment editor fields
| Option | What it does | Notes |
|---|---|---|
| Model name | The model or deployment name WinnerWare sends to the provider. For Azure OpenAI, this is the deployment name you created in Azure. | Required. It does not need to be unique, so two deployments can point at the same model with different capabilities. |
| Technical name | The unique name that settings, profiles, and recipes use to refer to this deployment. | Filled in from the model name while you type. You cannot change it after you save. |
| Connection name | The connection this deployment uses. | Required when the provider uses connections. If only one connection exists, it is selected for you. If none exist, the editor shows There are no configured connections. |
| Endpoint, Authentication type, API key, Identity client ID | Shown instead of Connection name for providers that keep their sign-in details on the deployment itself, such as Azure AI Services speech deployments. | Same meaning as on the connection editor. |
| Cascaded realtime card | Shown only for Cascaded Realtime deployments. It picks a Speech-to-text deployment, a Chat deployment, and a Text-to-speech deployment to chain into one voice deployment. | See AI Realtime Voice. |
| Model capabilities card | Declares what the model can do. | See the next section. |
The deployment editor has no temperature or response-length settings. You set those on the AI Profile or Chat Interaction.
The Model capabilities card
The card asks you to declare what the model supports. It has two parts:
- Trained features — checkboxes for things the model either can or cannot do.
- Model parameters — options that people can then tune on a profile or chat interaction. Today the only one is Reasoning effort, and it sits under the Reasoning checkbox.
Rules that apply to the whole card:
- At least one capability is required. The deployment will not save with every box cleared.
- New deployments start with Text conversation, Tool calling, and Streaming checked. Every other box starts cleared.
- Only what you check is used. Anything not declared is never shown in the profile and interaction editors and is never sent to the provider. Check a box only when the model really supports it, and do not clear a box the model supports, because WinnerWare treats the card as the full truth about the model.
| Checkbox | What it means | When it is on | When it is off |
|---|---|---|---|
| Text conversation | The model can hold a normal text chat. | The deployment can be a chat or utility deployment. | The deployment is not offered for chat or utility. Clear it only for a speech-to-speech model that cannot handle text. |
| Tool calling | The model can call tools, which is how an assistant takes action. | The assistant receives the tools you turn on in the Capabilities tab, plus built-in tools such as knowledge search. | Every tool is removed from every request. Capabilities such as Send Emails silently do nothing. See Tool calling. |
| Structured outputs | The model can reply in an exact JSON format. | Background features that ask the model for a fixed format, such as post-session analysis and data extraction, can require that format. | WinnerWare removes the format requirement from the request. Those features then depend on the model following the written instructions alone, which is less reliable. |
| Streaming | The model can send its reply word by word as it writes. | Replies appear in the chat as they are written. | WinnerWare waits for the full reply and then shows it all at once. Nothing breaks, but the chat feels slower. |
| Reasoning | The model thinks through a problem before it answers. | The Reasoning effort parameter becomes available. | WinnerWare removes all reasoning settings from the request. See Reasoning. |
| Image input (vision) | The model can look at images. | The deployment can be the vision deployment, used for image-aware Chat Interactions and for describing pictures in uploaded documents. | The deployment is not used for anything that needs to see an image. |
| Image output | The model can create images. | The deployment can be the image generation deployment. | The deployment is not used for image generation. |
| Audio input | The chat model itself accepts audio. | No WinnerWare feature reads this box today. | No effect today. Check it only if it is true, so the record stays accurate. |
| Audio output | The chat model itself produces audio. | No WinnerWare feature reads this box today. | No effect today. |
| Video input | The model can understand video. | No WinnerWare feature reads this box today. | No effect today. |
| Video output | The model can create video. | No WinnerWare feature reads this box today. | No effect today. |
| Realtime (speech-to-speech) | The model supports live, two-way voice conversations. | The deployment appears in the Conversation deployment list of profiles and chat interactions, and in the Speech-to-text deployment list of a Cascaded Realtime deployment. It is never offered as a Chat deployment. | The deployment cannot be used for realtime voice. If no deployment has this box checked, the Conversation chat mode uses speech-to-text plus text-to-speech instead. |
| Text embedding | This is a dedicated embedding model that turns text into search vectors. | The deployment can be the embedding deployment for knowledge, documents, and memory. | The deployment is not offered for embedding. Check this only on a real embedding model, never on a chat model. |
| Speech to text (transcription) | This is a dedicated transcription model, such as Whisper. | The deployment can be the speech-to-text deployment. | The deployment is not offered for transcription. This is different from Audio input. |
| Text to speech (synthesis) | This is a dedicated voice model that reads text aloud. | The deployment can be the text-to-speech deployment, and its voices appear in voice pickers. | The deployment is not offered for text-to-speech. This is different from Audio output. |
Tool calling
Tool calling is the checkbox that most often explains "the assistant ignores its capabilities."
Everything an assistant does, as opposed to says, runs through a tool: Send Emails, searching content, creating a content item, building a workflow, saving a memory, generating an image, and searching a knowledge data source. You turn those tools on in the Capabilities tab of an AI Profile or Chat Interaction (see AI Capabilities).
When Tool calling is not checked on the deployment that answers the chat, WinnerWare removes every tool from the request just before it is sent to the provider. This happens on every request and for every tool, no matter what the Capabilities tab says. The model never learns the tools exist, so:
- It cannot send an email, create content, or run any other capability.
- It may still say it did something, or describe the steps you should take, because it is only writing text.
- There is no message in the chat. The only trace is a warning in the application log that the deployment does not declare tool calling and that its tools were removed.
Knowledge answers are affected too. With Enable preemptive retrieval-augmented generation (RAG) turned on (the default, on the Default Orchestrator tab of Settings -> Artificial Intelligence), WinnerWare looks up knowledge before it calls the model, so grounded answers keep working. With that setting turned off, the model is told to search with a tool it no longer has, and it answers from its own general training instead of your knowledge base.
The check is made against the deployment that actually answers:
- For a text chat, that is the profile's or interaction's Chat deployment (or the Default chat deployment when it is set to Default).
- For a spoken turn in the Conversation chat mode, that is the realtime deployment, or the Chat deployment chosen inside a Cascaded Realtime deployment. When the Conversation mode falls back to speech-to-text plus text-to-speech, it is the profile's Chat deployment.
- A typed message during a voice conversation always goes to the profile's Chat deployment.
To fix it, open the deployment, check Tool calling on the Model capabilities card, and save. Only do this for models that really support tool calling. Nearly all current chat models do.
Reasoning
Reasoning models think through a request before they answer. Two settings control this, in two places.
On the deployment (Model capabilities card):
- Check Reasoning. A Reasoning effort box appears under it.
- Check Reasoning effort to let people tune it. Two more fields appear:
- Supported values — the effort levels this model accepts: Minimal, Low, Medium, High, and Extra high. Leave nothing selected to allow every level.
- Default value — the level used when a profile or interaction does not choose one. It starts at Medium. Choose No default to send no effort at all and let the provider decide.
On the AI Profile or Chat Interaction:
- On an AI Profile's Deployments & Interactions tab, a Reasoning effort list appears under Chat deployment (and under Utility deployment for that deployment's own setting).
- In a Chat Interaction, the same list appears in the settings panel under the deployment selectors.
- The list shows Use deployment default (with the default level in brackets) plus only the levels the deployment supports.
- The list only appears when the selected deployment has Reasoning and Reasoning effort checked. It stays hidden while the selector is set to Default, because the editor only reads the deployment you pick by name. The deployment's Default value still applies at run time.
What happens at run time:
| Setup | Result |
|---|---|
| Reasoning checked, Reasoning effort checked | The chosen level is sent. If a profile has a level the deployment no longer supports, WinnerWare uses the deployment's Default value instead, or sends no level if there is no default. |
| Reasoning checked, Reasoning effort not checked | No effort level is sent. The provider uses its own default. |
| Reasoning not checked | WinnerWare removes all reasoning settings from the request, so the provider never rejects the request for them. |
What people see in the chat: WinnerWare does not show the model's reasoning ("thinking") text, and there is no option to show a reasoning stream. People only see the final answer. The difference they notice is time: a higher Reasoning effort usually gives a more careful answer but a longer wait before the first words appear, and it costs more per request.
Realtime and voice deployments
Three kinds of deployment support voice:
- A speech-to-text deployment (Speech to text (transcription) checked) turns speech into text for the Audio input and Conversation chat modes.
- A text-to-speech deployment (Text to speech (synthesis) checked) reads replies aloud and supplies the voice list.
- A realtime deployment (Realtime (speech-to-speech) checked) holds a live spoken conversation. This can be a provider's own speech-to-speech model, or a Cascaded Realtime deployment that chains a speech-to-text, a chat, and a text-to-speech deployment.
Realtime is used through the Conversation chat mode. Set Chat mode to Conversation, and pick a realtime deployment as the Conversation deployment. The Chat deployment stays a text model: it answers typed messages, and realtime deployments are not in its list. If no realtime deployment is available, the Conversation mode uses the speech-to-text and text-to-speech deployments instead.
For the full voice setup, including native versus cascaded realtime and which models to choose, see AI Realtime Voice.
Usage
Pick deployments on a profile or chat interaction
- AI Profile: open the profile's Deployments & Interactions tab and choose a Chat deployment and a Utility deployment. Leave either on Default to follow the site default.
- Chat Interaction: choose the Chat deployment and Utility deployment in the interaction's settings panel.
Deployment lists are grouped by connection. A deployment whose name differs from its model name shows the model name in brackets, for example support-assistant (gpt-4.1-mini).
Checklist for a new chat deployment
- Model name matches the provider's deployment name exactly.
- Connection name points at the right account.
- Text conversation, Tool calling, and Streaming are checked (the defaults), unless the model truly lacks one of them.
- Reasoning is checked only for a reasoning model, and Reasoning effort is set up if you want people to tune it.
- Image input (vision) is checked only if the model can read images.
- Realtime (speech-to-speech), Text embedding, Speech to text (transcription), and Text to speech (synthesis) stay cleared on a normal chat model.
- The deployment is chosen as the Default chat deployment, or on the profiles that should use it.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The assistant says it sent an email, but nothing arrived. | Tool calling is not checked on the deployment that answered, so the Send Emails tool never reached the model. | Check Tool calling on that deployment. If it is already checked, see the next rows. |
| Every capability is ignored, and the assistant only describes what it would do. | Tool calling is not checked on the chat deployment. | Check Tool calling. Remember that a profile set to Default uses the Default chat deployment. |
| A capability works for an administrator but not for other people. | The role does not have permission to use that tool. | Grant the tool's permission under Security -> Roles. See AI Capabilities. |
| Send Emails runs but reports a failure. | Email delivery is not set up. | Configure Settings -> Email. |
| Answers ignore the knowledge base. | Tool calling is off and Enable preemptive retrieval-augmented generation (RAG) is also off. | Check Tool calling, or turn preemptive RAG back on under Settings -> Artificial Intelligence -> Default Orchestrator. |
| No realtime deployment appears in the Conversation deployment list. | No deployment has Realtime (speech-to-speech) checked. | Add a realtime deployment, or check the box on one that really supports it. See AI Realtime Voice. |
| A realtime deployment is missing from the defaults or from the Chat deployment list. | This is expected. Realtime deployments never fill the chat or utility role, and there is no default realtime deployment. | Set Chat mode to Conversation, and pick the realtime deployment as the Conversation deployment. |
| The voice assistant talks but cannot use capabilities. | The deployment that generates the voice reply does not have Tool calling checked. | Check Tool calling on the realtime deployment, or on the chat leg of a Cascaded Realtime deployment. |
| The Reasoning effort list does not appear on a profile. | The selected deployment does not have Reasoning and Reasoning effort checked, or the selector is set to Default. | Set up both boxes on the deployment, and pick the deployment by name to see the list. |
| No reasoning is shown in the chat. | This is expected. WinnerWare shows only the final answer. | No fix needed. Use Reasoning effort to trade speed for care. |
| The provider rejects requests with an error about a setting. | A box is checked that the model does not support, such as Reasoning or Structured outputs on a model without them. | Clear the box the model does not support. |
| Replies appear all at once instead of word by word. | Streaming is not checked. | Check Streaming if the model supports it. |
| A deployment is missing from a settings or profile list. | It does not have the capability that list needs, for example Text embedding for the embedding list. | Check the matching capability, but only if the model really has it. |
| The Audio input chat mode is missing. | No speech-to-text deployment is available. | Add a deployment with Speech to text (transcription) checked, and set it as the Default speech-to-text deployment. |
| The Conversation chat mode is missing. | No speech-to-text deployment and no realtime deployment is available. | Add a realtime deployment, or a speech-to-text deployment. |
| The deployment will not save and says a capability is required. | Every box on the Model capabilities card is cleared. | Check at least one capability. |
| You cannot edit a deployment or connection. | It is configuration-based and managed by your hosting team. | Ask your hosting team to change it. |
Operational Notes
- Deployments created before the Model capabilities card existed were given capabilities from their old purpose. Chat and utility deployments received Text conversation, Tool calling, and Streaming. Review older deployments and adjust the boxes to match each model.
- A deployment imported from a recipe that declares no capabilities is saved with only Text conversation checked, which means no tool calling. Open imported deployments and check their capabilities.
- Configuration-based chat deployments declare Text conversation, Tool calling, and Streaming. Configuration-based utility deployments declare only Text conversation and Streaming, so do not use one as the Chat deployment of an assistant that needs capabilities.
- A deployment's Technical name cannot change after you save, so choose one that will still make sense if you later change the model behind it.
- Capabilities belong to the deployment, not to the profile. A change applies to every profile, chat interaction, and default that uses that deployment.