Speech Insights
Speech Insights is the WinnerWare module for turning call recordings into searchable, governable, and reportable insight. It collects recordings from the places where they are delivered, transcribes them, removes sensitive data, applies your tags and indicators, and gives reviewers one place to search, listen, and export.
Overview
Speech Insights covers the full post-call workflow in six steps:
| Step | What happens | Where you manage it |
|---|---|---|
| 1. Ingestion | WinnerWare collects recordings and their metadata files from an ingestion source (file system, FTP/FTPS, or SFTP). | Speech Insights -> Ingestion Sources |
| 2. Transcription | Each recording is sent to Azure AI Speech, which returns the text of the call with speaker and timing information. | Speech Azure |
| 3. Redaction | AI finds sensitive values (card numbers, PINs, passwords, and so on) and masks them in the text, the search index, and the audio. | Speech Insights -> Redaction Rules |
| 4. Tags and indicators | Tags classify the call (for example "Escalation"). Indicators extract values from the call (for example "Cancellation reason"). | Speech Insights -> Transcript Tags, Indicators |
| 5. Search and review | Reviewers search completed calls, open a call to read and listen to it, and export results. | Speech Insights -> Search: Index Name |
| 6. Retro processing | New or changed tags and indicators are applied to calls that were already processed. | Speech Insights -> Retro Processing |
The Processing Queue shows where each call is in steps 2 to 4, and lets administrators retry calls that could not finish.
Key Features
- Collects recordings on a schedule (every hour) and on demand from one or more Ingestion Sources.
- Reports the last run of each source on the source list, and offers Test connection, Get statistics, and Clone actions.
- Supports Candidate locales so a source can transcribe calls in more than one language.
- Masks sensitive data with AI Redaction Rules before tags and indicators run.
- Supports Keyword, Semantics, and Expression transcript tags, plus AI Indicators that extract values from each call.
- Organizes tags and indicators into Transcript Tag Groups and Indicator Groups.
- Lets authors test an inactive tag or indicator against existing calls before turning it on (Rule Evaluations).
- Provides Keyword, Semantic, and Expression search, sorted by most recent call by default.
- Shows each call with collapsible sections, redacted audio playback, per-phrase silence colors, and an overall Transcript confidence.
- Supports Saved Lookups, CSV export, and an OData endpoint for tools such as Tableau, Power BI, or Excel.
- Provides a Processing Queue to monitor and retry transcript processing.
- Re-applies tags and indicators to older calls with Retro Processing.
How It Works
1. Bring recordings in through ingestion sources
An ingestion source tells WinnerWare where recordings are delivered and how to read their metadata. Go to Speech Insights -> Ingestion Sources, click Add Source, and pick a source type in Available Sources:
| Source type | Feature needed | Details |
|---|---|---|
| File System | Filesystem Recording Ingestion Source | Reads a folder in the tenant's own storage (under the tenant's Recordings folder). |
| FTP/FTPS | Speech Insights FTP | See Speech Insights FTP. |
| SFTP | Speech Insights SFTP | See Speech Insights SFTP. |
Source fields
Every source type has these fields:
| Field | What it does |
|---|---|
| Name | A business name for the source. Required. |
| Client | The client the calls belong to. Required. Tags, indicators, redaction rules, and search can be filtered by client. |
| Locale | The language used to transcribe the calls, for example en-US. Required. |
| Candidate locales | Optional. Select two or more locales (at most 10) to let the transcription service detect which of them is spoken and transcribe each part in that language. The language can change during the call. Select none to always use the Locale. |
| Active? | When on, the source is included in the hourly and manual ingestion runs. When off, the source is skipped. |
| Path | File System sources only. The folder to read, relative to the tenant's Recordings folder. Two sources cannot use the same path. |
Metadata File Map
Recordings are ingested only when a metadata file describes them. WinnerWare reads .json files (an array of objects) and .csv files (with a header row). Each row describes one recording.
The Metadata File Map tells WinnerWare which column in your metadata file holds each standard value. Enter your column name next to each field. When you leave a field empty, WinnerWare looks for a column with the same name as the field.
| Field | Meaning |
|---|---|
AudioFileName | The name of the audio file for this row. Required. |
CallDate | The date and time of the call. Required. |
CallId | Your unique call ID. WinnerWare stores it as the External ID and uses it to skip a recording that was already ingested for the same source. |
ANI, DNIS | Calling and called numbers. |
AgentName, AgentId | The agent on the call. |
Disposition, QueueName, SiteName, Direction, Language, CustomerId | Call details used by filters and search. |
TalkTime, HoldTime, AfterCallWorkTime | Call times. Enter a number of seconds (for example 480) or a clock value (for example 00:08:00). |
Columns that are not in the map are kept as custom metadata. They appear in the Metadata section of the call and can be used with Add Filter in search and in exports.
Supported audio files are .wav, .mp3, .opus, .flac, .alaw, and .ulaw.
How an ingestion run works
- Every hour, and when you click Ingest recordings, WinnerWare checks each active source.
- Files changed in the last 2 minutes are left for the next run, so files that are still uploading are not read. Empty audio files are skipped.
- For each metadata row, WinnerWare finds the audio file, creates a transcript record, and sends the audio for transcription.
- After a recording is accepted, WinnerWare removes it from the source. When every row in a metadata file is done, the metadata file is removed too.
- A recording that cannot be decoded (corrupt, cut off, or an unsupported format) is moved to a
failedfolder next to it at the source, and a failed transcript record is kept with the reason in its IngestionError metadata. It is not tried again. - Any other recording that fails is left at the source and tried again on the next run.
Only one run can work on a source at a time. If a run is still working on a source when the next run starts, the new run skips that source. This stops two runs from transcribing (and paying for) the same recording twice.
What the source list shows
Each source on Speech Insights -> Ingestion Sources shows:
| Item | Meaning |
|---|---|
| Active / Inactive | Whether the source is included in ingestion runs. |
| Processed time | When the last run finished. Hover over it to see how many recordings were ingested and how many failed ("Ingested N, failed N."). A warning icon shows that the last run stopped with an error. |
| Never processed | The source has not run yet. |
| Last error | Shown under the source name when an error stopped the last run. It is cleared when a run completes. |
A single recording or metadata file that fails does not create a Last error. It is counted in the failed number and tried again on the next run.
When an FTP or SFTP server cannot be reached during a run, WinnerWare logs the error, but the server looks empty to the run. The source then shows a normal Processed time with 0 recordings ingested. Use Test connection to confirm that the server can be reached.
Source actions
The Actions menu on each source has:
| Action | What it does |
|---|---|
| Test connection | Connects with the saved settings and lists the root folder. The result, the Endpoint, and the Duration show on the source row. |
| Get statistics | Reports what is waiting at the source without ingesting it: Files, Audio files, Metadata files, Other files, Ready to ingest, Waiting to settle, Empty audio files, Subdirectories, and Total size. |
| Ingest recordings | Starts an ingestion run for this source now, in the background. The source must be active. |
| Clone | Creates an inactive copy named "Name (Copy)" with the same type, client, locale, field map, and connection settings, including credentials. WinnerWare opens the copy so you can review it before you activate it. Candidate locales are not copied, so set them again if you need them. For a File System source, change the Path, because two sources cannot use the same path. |
Use Edit to change a source and Delete to remove it. You can also select several sources and delete them together.
2. Transcribe the recording
Each accepted recording is sent to Azure AI Speech for batch transcription. When the transcript is ready, WinnerWare downloads the text, identifies which speaker is the Agent and which is the Customer, and then runs the processing stages below. See Speech Azure for the transcription pipeline.
When a source has two or more Candidate locales, Azure identifies the spoken language continuously, and each phrase stores the language that was detected for it.
3. Redact sensitive data
Before tags and indicators run, each transcript is sent to the AI deployment for sensitive-data redaction. The AI replaces sensitive values (for example card numbers, social security numbers, PINs, passwords, and CVVs) with ***** and keeps labels such as "account number" visible.
- Redaction Rules add your own instructions to the base redaction prompt. WinnerWare creates a Default PII Redaction rule during setup.
- Redaction is applied to the phrase text and to each word. When the redacted text cannot be matched word by word (for example when a multi-word card number becomes one token), the whole phrase is masked so that no raw value is saved or indexed.
- WinnerWare then creates a redacted audio file, which is the audio reviewers hear.
- A transcript is never marked complete until redaction has finished.
4. Apply tags and indicators
After redaction, WinnerWare applies every active tag and indicator whose filters match the call.
Transcript tags
A tag marks a call with a label. Go to Speech Insights -> Transcript Tags and click Add Tag. The Type decides how the tag is matched:
| Type | How it matches | Uses AI |
|---|---|---|
| Keyword | Looks for a word or phrase in each phrase of the call. | No |
| Semantics | The AI reads the whole call and decides whether the tag's Description applies. | Yes |
| Expression | Uses the same syntax as Expression search, for example speaker:customer after:60s (refund OR cancel). See Speech Insights Expression Search. | No |
Tag options:
| Option | Applies to | What it does |
|---|---|---|
| Name | All | The tag label shown on calls and in filters. Required. |
| Description | All | Explains the tag. For Semantics tags, the AI uses it to decide whether the tag applies. |
| Good examples | Semantics | Optional. Conversations or phrases where the tag should be applied. |
| Bad examples | Semantics | Optional. Conversations or phrases where the tag should not be applied. |
| Search scope | Keyword | Entire transcript, Customer phrases only, or Agent phrases only. |
| Match type | Keyword | Contains, Starts with, Ends with (compared with each phrase, not case-sensitive), or Lucene query for operators such as +, -, AND, OR, NOT, wildcards, fuzzy, and proximity. The editor shows a syntax reference for Lucene query. |
| Keyword / Expression | Keyword, Expression | The text or expression to match. An invalid expression cannot be saved. |
| After offset (seconds) | Keyword | Only matches phrases that occur after this many seconds from the start of the recording. |
| In last (seconds) | Keyword | Only matches phrases in the last this-many seconds of the recording. |
| Clients, Ingestion sources | All | Only applies the tag to calls from these clients or sources. |
| Groups | All | Adds the tag to one or more Transcript Tag Groups. |
| Recording date from, Recording date to | All | Only applies the tag to recordings in this date range. |
| Dispositions, Queues, Agents | All | Only applies the tag to calls with these values. Enter names separated by commas. |
| Active? | All | When on, the tag is applied to all new recordings that match the filters. |
In an Expression tag, tag:"Name" can refer to tags already on the call and to Keyword or Expression tags matched earlier in the same pass. Semantics tags are classified after expression tags, so on a new call a tag: reference to a Semantics tag does not match yet.
Indicators
An indicator uses AI to extract a value from each call, for example the reason for a cancellation. Go to Speech Insights -> Indicators and click Add Indicator.
| Option | What it does |
|---|---|
| Name | The indicator name. Required. |
| Description | Explains the indicator. On the indicator list, hover over the info icon to read it. |
| Extraction Prompt | The AI instructions that say what to extract and how to format it. Required. |
| Clients, Ingestion sources | Only applies the indicator to calls from these clients or sources. |
| Groups | Adds the indicator to one or more Indicator Groups. |
| Recording date from, Recording date to | Only applies the indicator to recordings in this date range. |
| Index | Optional. Pick a search index to load the values for the pick lists below. The index is used only to fill the lists. It is not used when matching calls. |
| Dispositions, Queues, Agents, Sites, Directions | Only applies the indicator to calls with the selected values. |
| Talk time (seconds), Hold time (seconds), After call work time (seconds), Total silence (seconds), Longest silence (seconds) | Only applies the indicator to calls whose value is in the From / To range. To must be greater than or equal to From. |
| Active? | When on, the indicator is evaluated on all new recordings that match the filters. |
Each indicator is extracted with its own AI request, so its prompt is evaluated alone. If the AI provider refuses a call's text because of its content policy, that indicator is skipped for that call, the other indicators still run, and the call still completes.
Tag groups and indicator groups
Transcript Tag Groups and Indicator Groups have a Name, a Description, and an Active? setting. Use them to organize long lists. The tag and indicator lists can be filtered by Group, Client, and Ingestion source, and the tag list also by Type.
Test a tag or indicator before you turn it on
An inactive tag or indicator shows Evaluate against existing transcripts on its edit screen. The evaluation first applies the rule's saved filters to find candidate calls, then evaluates a sample:
| Option | What it does |
|---|---|
| Sample size | How many candidate calls to evaluate. It is capped by a hard limit (by default 5,000 for keyword tags and indicators, and 250 for semantics and expression tags). |
| Evaluation mode | Summary only stores only the counts. Full Evaluation stores each sampled call so you can page through the results. |
| Created from, Created to | Optional. Only include transcripts created in this date range. |
Click Run evaluation. The run is saved immediately and processed in the background. Follow it in Speech Insights -> Rule Evaluations, where each run shows In Progress, Succeeded, or Failed, with Candidates, Sampled, and Matched counts. For Full Evaluation, the stored rows show Matched or No match, the Matched text or Extracted value, and the Reason.
When the results look right, turn on Active?. To apply the rule to calls that already exist, use Retro Processing.
5. Search and review
Search a transcript index
Each Speech Insights search index adds a Speech Insights -> Search: Index Name page. The search form has collapsible sections. WinnerWare remembers which sections you opened or closed in your browser.
| Section | What it contains |
|---|---|
| Search Form | The search mode (Keyword, Semantic, or Expression), the query, the speaker (Any Role, Agent, or Customer), a Date range, and the Tags, Sites, Clients, Dispositions, Queues, Agents, and Indicators filters. |
| Silence Filters | Total Silence From (seconds), Total Silence To (seconds), Longest Silence (seconds), and Show silence time, which adds a silence column to the results. |
| Call Time Filters | Talk Time (seconds), Hold Time (seconds), and After Call Work Time (seconds), each with Equals, Not Equals, Greater Than, Less Than, or Between. |
| Advanced Filters | Add Filter for custom metadata fields and Add Indicator Filter for extracted indicator values, with operators such as Equals, Contains, Starts With, Greater than, Advanced Search, or Regex, and match options such as Case Insensitive and Fuzzy (Auto). |
| Result Sorting | Sort by and Direction. The default is Most recent (default): newest call first. Choose Relevance to rank by search match, or pick a field such as call date, silence, talk time, hold time, or confidence. |
The search modes work like this:
- Keyword finds words and phrases, together with the filters.
- Semantic finds calls by meaning. WinnerWare reads the question, picks out filter details such as agents or dates, and compares meaning with embeddings created when the call was processed.
- Expression uses the speech-analytics syntax described in Speech Insights Expression Search.
The result list shows the fields selected in Settings -> Speech Insights -> Search result fields, the tags on each call, and a View link.
Review a call
Click View to open a call. The page has collapsible sections, in this order:
| Section | What it shows |
|---|---|
| Recording Information | General call details in two columns (file name, ANI, DNIS, call date, agent, disposition, site, customer ID, direction, queue, language), then the timing values: Talk time, Hold time, After call work time, Total silence, and Transcript confidence (the average confidence of the recognized phrases). |
| Metadata | Custom metadata columns from the metadata file. |
| Indicators | The value extracted for each indicator, or No value extracted. |
| Tags | The tags applied to the call. |
| Recording | The redacted audio player and the phrase-by-phrase transcript. Click a time to jump to it, copy a link to a phrase, or play from a phrase. |
Each phrase shows the silence before it. Silences are gray when short, orange from 3 seconds, and red from 10 seconds. You can change these thresholds in Settings -> Speech Insights -> Silence highlighting. The audio player appears only for users with the Download redacted speech audio file permission.
Saved lookups
A saved lookup stores a set of search filters under a name, so reviewers can repeat a search in one click.
- Create one in Speech Insights -> Saved Lookups with Add Lookup. Pick the index, enter the Display text, and fill in the search form under Lookup Values.
- On the search page, the Saved Lookups panel lists saved lookups. Click Apply to load one into the search form.
Export results
When the Speech Insights Export feature is on, the Query & Export panel offers:
- Export File: choose the standard fields, Tags, Indicators, and Additional Metadata to include, and download a CSV file. Up to 10,000 records download directly. Larger exports run in the background, and WinnerWare sends a notification when the file is ready.
- Elasticsearch DSL Query JSON: the query behind the current search.
- The OData Service URL and Metadata URL for tools such as Tableau, Power BI, or Excel. The endpoint uses basic authentication and requires the Access indexed data using API permission.
Exported files are kept for the number of days in Report retention (days). A weekly task removes older files.
6. Monitor processing and retry failures
Speech Insights -> Processing Queue shows each call's processing task. Filter by All, Pending, Processing, Completed, or Abandoned, search by file name or External ID, and sort by Created or Last attempt. A banner shows Queue is healthy or Queue has abandoned items.
| Status | Meaning |
|---|---|
| Pending | Waiting to be processed, or waiting for its next retry (see Next retry). |
| Processing | Being processed now. |
| Completed | Every stage finished: download, speaker identification, redaction, tags, indicators, redacted audio, and clean-up. |
| Abandoned | The task failed on every attempt. Open Error details to see why. |
A failed task is retried automatically after 1, 5, 15, 30, and 60 minutes, up to 6 attempts (Attempt X/Y). Each retry continues from the stage that failed. After the last failed attempt, the task is Abandoned. Fix the cause, then click Retry on the task, or use Retry selected or Retry all abandoned.
7. Apply rules to older calls
New and changed tags and indicators apply only to new calls. To apply them to calls that were already processed, use Speech Insights Retro Processing.
Configuration
Enable the features your rollout needs
| Admin display name | Feature ID | Use when |
|---|---|---|
| Speech Insights | CloudSolutions.Speech.Insights | You need the main Speech Insights admin, search, tagging, redaction, saved lookup, retro processing, and queue features. |
| Filesystem Recording Ingestion Source | CloudSolutions.Speech.Insights.Filesystem | You want recordings to be ingested from a folder in the tenant's storage. |
| Speech Insights Export | CloudSolutions.Speech.Insights.Export | You want CSV exports and the Settings -> Speech Insights page (result fields, silence highlighting, and export retention). |
| Speech Insights Export - Azure | CloudSolutions.Speech.Insights.Export.Azure | You want exported files stored in Azure instead of the local tenant file system. |
Transcription also needs Speech Service (Azure). See Speech Azure.
Create a search index
Search pages come from search indexes. Go to Search -> Indexes, add an Elasticsearch index, and choose the Speech Insights source. Each index adds its own Search: Index Name menu item. Users need permission to query search indexes to open it.
Set the AI deployments
Speech Insights uses the tenant defaults in Settings -> Artificial Intelligence:
| Deployment | Used for |
|---|---|
| Utility deployment (or the Chat deployment when no utility deployment is set) | Redaction, Semantics tags, indicators, rule evaluations, retro processing, and reading Semantic search questions. |
| Embedding deployment | Phrase embeddings created during processing and Semantic search. |
See AI Deployments for how to create deployments and set the defaults.
Speech Insights settings
With the Speech Insights Export feature on, Settings -> Speech Insights has these settings. They require the Manage Speech Insights Settings permission.
| Setting | What it does |
|---|---|
| Report retention (days) | How long exported reports are kept and can be downloaded. Default 7. Must be at least 1. |
| Search result fields | Which standard fields appear on the search results. Available fields: agent name, agent ID, client name, disposition, queue name, site name, direction, language, call date, talk time, hold time, after call work time, total silence, file name, customer ID, ANI, and DNIS. When nothing is selected, WinnerWare shows agent name, client name, disposition, queue name, talk time, total silence, and call date. |
| Orange silence starts at (seconds) | Silences at or above this value are orange. Default 3. |
| Red silence starts at (seconds) | Silences at or above this value are red. Default 10. Must be greater than or equal to the orange value. |
Main admin areas
| Admin area | Purpose |
|---|---|
| Speech Insights -> Expression Search | A reference page for the expression syntax. |
| Speech Insights -> Search: Index Name | Search completed calls for a Speech Insights index. |
| Speech Insights -> Ingestion Sources | Create and manage recording sources. |
| Speech Insights -> Transcript Tags | Maintain tag definitions. |
| Speech Insights -> Transcript Tag Groups | Organize tags into groups. |
| Speech Insights -> Indicators | Maintain indicator definitions. |
| Speech Insights -> Indicator Groups | Organize indicators into groups. |
| Speech Insights -> Retro Processing | Apply tags and indicators to calls that were already processed. |
| Speech Insights -> Rule Evaluations | Follow tag and indicator test runs. |
| Speech Insights -> Redaction Rules | Control how sensitive content is masked. |
| Speech Insights -> Saved Lookups | Save reusable searches. |
| Speech Insights -> Processing Queue | Monitor processing and retry abandoned calls. |
| Settings -> Speech Insights | Export retention, search result fields, and silence highlighting. |
Redaction rules
Go to Speech Insights -> Redaction Rules to add instructions to the redaction prompt.
| Option | What it does |
|---|---|
| Name | The rule name. Required. |
| Description | Explains the rule. |
| Rule Instructions | Free-form text about what to redact or what not to redact. It is added to the base redaction prompt. Required. |
| Clients, Ingestion sources | Only applies the rule to calls from these clients or sources. Leave empty to apply it to all calls. |
| Active? | When on, the rule is used during redaction of matching calls. |
Permissions
| Permission | Purpose |
|---|---|
| Manage Ingestion Sources | Create, edit, test, clone, run, and delete ingestion sources. |
| Manage Speech Transcript Tags | Maintain transcript tags and tag groups. |
| Manage Speech Indicators | Maintain indicators and indicator groups. |
| View Speech Rule Evaluations | Open Rule Evaluations and see your own evaluation runs. To start an evaluation, a user needs Manage Speech Transcript Tags or Manage Speech Indicators. |
| View Speech Rule Evaluations Created By Others | See evaluation runs created by other users. |
| Manage Speech Transcript Retro Processing | Create and control retro-processing runs. |
| Delete Speech Transcript Retro Processing Runs | Delete finished retro-processing run records. |
| Manage Speech Lookups | Maintain saved lookups. |
| Manage Sensitive Data Redaction Rules | Maintain redaction rules. |
| Manage speech transcript processing queue | Monitor the processing queue and retry abandoned work. |
| Manage Speech Insights Settings | Edit Settings -> Speech Insights and open the Expression Search reference page. Added by the Speech Insights Export feature. |
| Export speech transcript | Generate and download exports. Added by the Speech Insights Export feature. |
| Access indexed data using API | Read indexed call data through the OData endpoint. |
| Download redacted speech audio file | Play the redacted recording on the call page. |
Tune sensitive-data redaction
If the AI deployment is slow, throttled, or unreachable, redaction requests time out, and the call is retried and can be abandoned in the Processing Queue with an error such as (sensitive-data-redaction) The AI endpoint did not respond within the configured timeout ....
A system administrator can tune redaction in the CloudSolutions_SpeechInsights_Redaction configuration section (for example in appsettings.json or the tenant configuration):
{
"CloudSolutions_SpeechInsights_Redaction": {
"RequestTimeout": "00:01:00",
"ProcessingTimeBudget": "00:02:00",
"MaxPhrasesPerRequest": 50,
"RedactionToken": "*****",
"WordSensitivityThreshold": 0.85,
"ReasoningEffort": "Low"
}
}
| Setting | Default | Purpose |
|---|---|---|
RequestTimeout | 00:01:00 (60s) | Maximum time for one redaction or embedding request. Increase it if a healthy deployment needs longer to respond. |
ProcessingTimeBudget | 00:02:00 (120s) | Total time for one group of phrases, including retries. After that, no more retries are scheduled for that group. |
MaxPhrasesPerRequest | 50 | Number of phrases sent per AI request. Smaller groups give shorter requests that are less likely to time out. |
RedactionToken | ***** | The text that replaces redacted content. |
WordSensitivityThreshold | 0.85 | Threshold (0 to 1) used when mapping redactions back to individual words. |
ReasoningEffort | Low | Reasoning effort sent with the redaction request. See Reduce latency on reasoning-capable models. |
If redaction keeps timing out while other AI features work, first check the health and quota of the tenant's default utility or chat deployment. Continued timeouts usually mean an overloaded, throttled, or misconfigured deployment, not a data problem.
Tune real-time tagging and rule evaluations
The CloudSolutions_SpeechInsights_Evaluation section controls tags and indicators on new calls and the Evaluate against existing transcripts runs:
{
"CloudSolutions_SpeechInsights_Evaluation": {
"ReasoningEffort": "Low",
"IndicatorConcurrency": 4,
"MaxSemanticTagsPerCall": 20,
"MaxSemanticTagCharsPerCall": 8000,
"SemanticTagConcurrency": 4,
"KeywordTagHardLimit": 5000,
"SemanticTagHardLimit": 250,
"IndicatorHardLimit": 5000
}
}
| Setting | Default | Purpose |
|---|---|---|
ReasoningEffort | Low | Reasoning effort for semantic tag and indicator requests. |
IndicatorConcurrency | 4 | Number of indicator requests sent at the same time for one call. |
MaxSemanticTagsPerCall | 20 | Maximum number of Semantics tags in one AI request. 0 removes the limit. |
MaxSemanticTagCharsPerCall | 8000 | Maximum combined length of tag names and descriptions in one AI request. 0 removes the limit. |
SemanticTagConcurrency | 4 | Number of semantic tag requests sent at the same time for one call. |
KeywordTagHardLimit | 5000 | Maximum sample size when evaluating a keyword tag. |
SemanticTagHardLimit | 250 | Maximum sample size when evaluating a semantics or expression tag. |
IndicatorHardLimit | 5000 | Maximum sample size when evaluating an indicator. |
Retro processing has its own settings. See Speech Insights Retro Processing.
Reduce latency on reasoning-capable models
Redaction, semantic tags, and indicators all send the call to the default utility or chat deployment. When that deployment is a reasoning model (for example the GPT-5 family, including GPT-5 nano), most of the response time goes into reasoning before the answer. These tasks are classification and extraction, not deep reasoning, so a lower reasoning effort usually makes them much faster with little or no loss in quality.
Each processing path has its own ReasoningEffort setting:
| Section | Applies to |
|---|---|
CloudSolutions_SpeechInsights_Redaction | Sensitive-data redaction requests. |
CloudSolutions_SpeechInsights_Evaluation | Semantic tags and indicators on new calls, and rule evaluations. |
CloudSolutions_SpeechInsights_RetroProcessing | Every retro-processing request. |
Valid values are None, Low, Medium, High, and ExtraHigh. Each section defaults to Low. Leave the value empty ("") when the deployment uses a non-reasoning model (for example GPT-4o), because those models reject requests that include a reasoning effort.
Semantics tags are split into batches, so a call with many tags is classified in several smaller requests instead of one oversized request. A new batch starts when MaxSemanticTagsPerCall or MaxSemanticTagCharsPerCall is reached. Each request still carries the whole conversation, because semantic tagging needs the full context. Indicators always use one request per indicator. If a model returns indicator values in an unexpected shape (for example objects instead of plain text), WinnerWare converts them to key: value text instead of dropping the result.
Usage
Example rollout
- Enable Speech Insights, Speech Service (Azure), the source features you need, and Speech Insights Export if you want exports and the settings page.
- Set the default utility or chat deployment and the embedding deployment in Settings -> Artificial Intelligence.
- Create a Speech Insights index in Search -> Indexes.
- Review the Default PII Redaction rule and add redaction rules if you need them.
- Create an ingestion source, fill in the Metadata File Map (at least
AudioFileNameandCallDate), and use Test connection and Get statistics to check it. - Turn on Active? for the source, then click Ingest recordings or wait for the hourly run.
- Watch the Processing Queue until calls are Completed, then open the search page to review them.
- Add tags and indicators, test them with Evaluate against existing transcripts, and activate them.
- Use Retro Processing to apply the new rules to older calls.
- Add saved lookups and exports when the review pattern is stable.
What to do when something fails
| Symptom | Where you see it | What to do |
|---|---|---|
| A source shows Last error, or ingests nothing | Ingestion Sources | Use Test connection to check the host, credentials, and path. Fix the source, then click Ingest recordings. |
| A source shows failed recordings but no error | Hover over Processed | Those recordings stay at the source and are retried on the next run. Check the logs if the number does not go down. Unreadable recordings are in the failed folder at the source. |
| Get statistics shows files Waiting to settle | Source row | The files changed in the last 2 minutes. They are ingested on a later run. |
| Recordings stay at the source | Get statistics | Check that a metadata file lists them in the AudioFileName column and that the source is active. |
| Calls are Abandoned | Processing Queue | Open Error details. Timeouts and throttling usually mean the AI deployment is overloaded or has no quota. Fix the cause, then use Retry all abandoned. |
| A tag or indicator is missing on older calls | Call page | Rules apply only to new calls. Use Retro Processing. |
| An indicator is empty on one call | Call page | The AI provider may have refused the call's text because of its content policy. The logs show a warning for that indicator. |
| No audio player on the call page | Call page | The user needs the Download redacted speech audio file permission. |
Operational Notes
- Ingestion runs every hour. Retro-processing recovery runs every 5 minutes. Export clean-up runs weekly.
- A recording with a CallId that was already ingested for the same source is removed from the source and not ingested again.
- Tags, tag groups, indicators, indicator groups, and redaction rules can be moved between environments with deployment plans (Speech Transcript Tags, Speech Transcript Tag Groups, Speech Indicators, Speech Indicator Groups, Sensitive Data Redaction Rules) and matching recipe steps. On import, items are matched by ID and then by name, so re-importing updates the existing item.
- The OData endpoint is available under
speech/insights/odata/{indexName}. - After a recording is accepted, it is removed from the source. After processing, the shared copy used for Azure transcription is deleted. Reviewers hear only the redacted audio.