Skip to main content
Calibration Sessions help QA teams identify and resolve scoring differences by having multiple QA Analysts independently evaluate the same set of customer conversations. QA Analysts may score the same customer conversation differently because they interpret evaluation criteria differently. Teams lack visibility into scoring inconsistencies, analyst alignment, and adherence to team standards without a structured way to compare evaluations. The Calibration Manager creates the session, selects the conversations and evaluation form, and assigns analysts. Analysts then score the conversations independently in Calibration Mode, where peer scores remain hidden. After all analysts submit their evaluations, the manager uses Debrief View to compare scores, resolve differences, and set Reference Scores for contested metrics. The manager then closes the session, generating a shareable Calibration Session Report with agreement results, analyst accuracy, and resolution notes. The workflow gives teams a consistent way to measure analyst alignment, establish reference outcomes, and improve how they apply evaluation criteria over time.

Key Capabilities

Calibration Sessions provide the following capabilities.

Prerequisites

Before creating a Calibration Session, confirm that you have the following:

Access Calibration Sessions

Navigate to Quality AI > AutoQA > Calibration Sessions. Calibration Session Landing Page
Users with both the Audit Allocations permission and Cross Queue Data Access can view calibration sessions across the entire application, not only the sessions they created.

Calibration Sessions Landing Page

The Calibration Sessions landing page provides a centralized view of calibration sessions. Calibration Managers can create sessions, monitor active sessions, review closed sessions, and continue working with draft sessions. The page includes three tabs: Session Status Use the Search field to find a session by name.

Session List

The session list displays the following information: Agreement uses traffic-light coloring:
  • Green: 90%-100%
  • Amber: 70–89%
  • Red: Below 70%

Session Status

Create a Calibration Session

Select + Create New Session to start a calibration exercise. Create Calibration Session The wizard has three steps: General, Select Conversations, and Assign Auditors. A Calibration Manager creates a session by configuring the session, selecting conversations, and assigning QA Analysts. After the manager activates the session, assigned analysts receive the calibration work in their audit queue.

Step 1: General

Define the session configuration.
  1. Enter a Name for the calibration exercise.
  2. Enter a Description with instructions or context for participating analysts.
  3. Turn on Analyst Justification > Mandatory to require analysts to add a comment before submitting each metric score.
  4. Select an AutoQA Visibility option.
  5. Select the Evaluation Form analysts will use to score the conversations.
  6. Select Next, or Save as Draft to continue later.

Configure AutoQA Visibility

The Calibration Manager can control when analysts see AutoQA recommendations during calibration. The available options are: Create New Session

Step 2: Select Conversations

Build the set of conversations analysts must evaluate
  1. Select a Conversation Type to define which conversations are available to add.
    • Audited: Displays audited conversations.
    • Unaudited: Displays conversations that aren’t audited.
  2. Select a tab to choose the conversation source:
    • All Conversations: Conversations that match the selected filters.
    • Marked for Calibration: Conversations that analysts or supervisors previously flagged using Mark for Calibration on the audit screen. See Mark a Conversation for Calibration.
  3. Use the Search field to find specific conversations.
  4. Select Filter to narrow the list of available conversations.
  5. Configure the required filters and select Apply.
  6. Select Select all or individual conversations, then select Add Interactions then select Add Interactions to add it to the Selected Conversations panel.
  7. Review the Selected Conversations panel. Select View Conversation to preview a conversation, or remove a conversation before continuing.
  8. Select Next.
Managers can preview or remove conversations before continuing. They must select at least one conversation to proceed. Select Conversations

Conversation Filters

The filter panel helps narrow conversations before adding them to the session. Select Conversations

Mark a Conversation for Calibration

Supervisors and auditors can flag a conversation for calibration directly from the standard audit screen, without opening the Calibration workflow first.
  1. Open a conversation on the audit screen.
  2. Select the more options menu () in the header.
  3. Select Mark for Calibration.
Flagged conversations appear under the Marked for Calibration tab in Step 2: Select Conversations of session creation, so a Calibration Manager can add them to a session later.

Step 3: Assign Auditors (QA Analysts)

Assign the QA Analysts who can evaluate every selected conversation independently.
  1. Use Search Auditors to Add to find and select one or more analysts.
  2. Review the selected analysts listed in the Auditors section. Remove an analyst using the delete icon if needed.
  3. Each assigned analyst receives the same set of conversations and evaluates them independently.
  4. Select Save as Draft to save the session and continue configuring it later.
  5. Select Create and Notify Participants to create and activate the session, or Save as Draft to continue later.
  6. Select Save to save the current configuration.
A Calibration Session requires a minimum of two analysts. Individual evaluations remain hidden from other analysts until the session reaches the comparison stage.
Selecting Create and Notify Participants sets the session status to Active and sends an email notification to each assigned analyst. Select Conversations

My Calibration Sessions

Access Assigned Calibration Sessions

QA Analysts access assigned Calibration Sessions from the Calibration tab in Audit Allocations. The system also sends an email notification when it assigns a session to an analyst. The Calibration tab appears when the analyst has at least one assigned session. Each session displays the session name, assigned-by name, scoring progress (for example, 2 of 4 conversations scored), and session status (Not Started, In Progress, or Submitted). Each session displays the Session name, Assigned-by name, Scoring progress (for example 2 of 4 conversations scored), and Session status: (Not Started, In Progress, or Submitted). From an assigned session, an analyst can:
  • Review the conversations included in the session, with a per-conversation status of Score Now, Scored, or Locked.
  • Monitor evaluation progress.
  • Open pending conversations and complete calibration evaluations.
  • Review the calibration report after the manager closes the session.

Monitor an Active Session

Opening an active session from the landing page displays the manager’s monitoring view. The session header displays the Created Date, Auditors avatar stack, Conducted by, and Status.
  • Submission Progress: Shows the overall submission count (for example, 8/8 Submitted) and a card for each assigned analyst. Each card displays the analyst’s name, number of Conversations Scored, and a progress bar. The card turns green when the analyst reaches 100%.
  • Conversation-Level Summary: Lists every conversation in the session with the following details:
:The Flows/Queues column displays the Experience Flow for conversations handled by an AI Agent, and the queue for conversations handled by a Human Agent.

Calibration Evaluation

Calibration Mode is the scoring experience for QA Analysts participating in a Calibration Session. It uses the selected evaluation form and keeps peer scores hidden so analysts can independently evaluate each conversation. A Calibration Mode indicator identifies that the analyst is completing a calibration evaluation rather than a production audit. During calibration:
  • Peer scores remain hidden.
  • AutoQA visibility follows the session’s configured AutoQA Visibility setting.
  • Analysts evaluate each metric independently.
  • Add comments when Analyst Justification has the Mandatory option enabled.
  • Audit Progress displays evaluation completion status.
  • Analysts can’t modify an evaluation after submission.
  • Calibration Session conversations aren’t routed to the agent for review, acknowledgement, or dispute.

Complete a Calibration Evaluation

  1. Open an assigned Calibration Session.
  2. Select a pending conversation.
  3. Review the conversation transcript.
  4. Evaluate each metric using the selected evaluation form.
  5. Add notes to provide additional scoring context.
  6. Review the completed evaluation.
  7. Confirm that the evaluation is final.
  8. Select Submit Score.
After submission, the evaluation becomes final. The analyst can’t modify it unless a separate correction workflow is available. Submitted calibration scores don’t affect agent dashboards, leaderboard rankings, or scorecard scores.

Debrief View

After all assigned analysts submit their evaluations, the Calibration Manager can open Debrief View to compare analyst scores side by side for each evaluation metric before setting a reference score. Debrief View opens from a conversation in the session and displays a metric-by-metric comparison grid. The closed Calibration Session Report uses the same grid when you select the view icon for a conversation in the Conversation-Level Summary, as described under Calibration Session Report.

Debrief View Components

A Participant Summary at the bottom of Debrief View shows, for each analyst: Metrics Scored, Aligned with Consensus (%), and Accuracy.

Review a Disagreement

  1. Open Debrief View after all required analyst evaluations are complete.
  2. Review the analyst scores for each metric.
  3. Identify metrics where the scores differ.
  4. Expand the metric to review analyst notes and AutoQA AI Justification.
  5. Determine the appropriate reference score.
  6. Add resolution notes explaining the decision.
  7. Continue until you resolve all required disagreements.

Reference Score

The Reference Score represents the Calibration Manager’s final determination for a metric after reviewing the analyst scores and supporting context. The manager uses the Reference Score to establish the expected scoring outcome for the calibration exercise. The reference score should reflect the team’s agreed interpretation of the evaluation criterion.

Close a Calibration Session

The Calibration Manager closes the session after reviewing the submitted evaluations and completing the debrief. Before closing the session, the manager should:
  1. Review all analyst scores.
  2. Resolve the required scoring differences.
  3. Set reference scores.
  4. Add resolution notes.
  5. Review the calibration results.
  6. Close the session.
After the manager closes the session, the final calibration results become available to participants through the session report. Select Conversations

Calibration Session Report

Closing a session generates a read-only Calibration Session Report, which is available to the Calibration Manager and all participating QA Analysts. The report summarizes calibration outcomes, analyst alignment, scoring consistency, and areas that require additional calibration. The report header displays the session name, date closed, conducted by, participants, and conversation count.

Report Sections

Conversation-Level Comparison

Select the view icon on a Conversation-Level Summary row to expand the row inline and reveal the metric-by-analyst comparison grid directly beneath it, without leaving the report. This expanded grid uses the same structure as Debrief View.

Accuracy

In the reviewed report data, Accuracy measures how well an analyst’s scores match the final Reference Scores established by the Calibration Manager during Debrief. For example, if an analyst’s scores match the Reference Scores for 8 out of 10 metrics, their accuracy is 80%. If the Calibration Manager hasn’t established Reference Scores, the report can’t calculate Accuracy.

Export the Report

Select Export CSV to download the calibration results as a CSV file.

Exported CSV Fields

The exported CSV contains calibration results from the report.
The exact CSV column names and export structure may vary based on the final report implementation.
Select Conversations QA Analysts can use the report to identify differences between their evaluations and the final Reference Scores and review any Resolution Notes from the calibration process. The report also compares AutoQA recommendations with the final consensus to help QA Analysts assess AutoQA’s alignment with human evaluators.

Calibration Results and Formulas

After the manager completes the debrief, the system calculates calibration results. The results provide an overall view of scoring consistency and identify differences between individual analyst evaluations and the established Reference Score. Every calculation treats N/A as a third verdict alongside Adhered and Not Adhered. The system includes every analyst response in each metric’s denominator. The system calculates agreement and alignment metrics using the following calculations.

Variables

Metric Agreement Rate

A metric reaches consensus when one verdict receives a majority of analyst responses: count>N/2\text{count} > N / 2 Otherwise, the system labels the metric No Consensus and excludes it from further calculations. Metric Agreement Rate=max(A,B,C)N\text{Metric Agreement Rate} = \frac{\max(A, B, C)}{N} The system computes the Metric Agreement Rate only when the metric reaches consensus.

Session Agreement Rate

The Session Agreement Rate averages the agreement rates for all metrics that reach consensus. The system excludes metrics labeled No Consensus. sum of all metric agreement rates / number of consensus metrics Session Agreement Rate=sum of all metric agreement ratesnumber of consensus metrics\text{Session Agreement Rate} = \frac{\text{sum of all metric agreement rates}}{\text{number of consensus metrics}}

Agreement Distribution

The system groups metrics into the following agreement categories:

Session N/A Rate

The Session N/A Rate measures the proportion of metric responses that analysts marked N/A. Session N/A Rate=total N/A verdicts across all metricsC×M×N\text{Session N/A Rate} = \frac{\text{total N/A verdicts across all metrics}}{C \times M \times N}

Auditor Agreement Rate (Aligned with Consensus)

For each analyst, the proportion of their verdicts that matched the consensus verdict, excluding No Consensus metrics from both the numerator and denominator. Auditor Agreement Rate=number of matching verdictsnumber of consensus metrics\text{Auditor Agreement Rate} = \frac{\text{number of matching verdicts}}{\text{number of consensus metrics}}

Auditor N/A Rate

The Auditor N/A Rate measures the proportion of metrics for which an analyst selected N/A. Auditor N/A Rate=metrics where the analyst responded N/AC×M\text{Auditor N/A Rate} = \frac{\text{metrics where the analyst responded N/A}}{C \times M} The Participant Summary flags an analyst when their N/A rate exceeds the session average by more than 15 percentage points.

Auditor Inter-Auditor Agreement (Cohen’s Kappa)

Measures whether an analyst’s alignment with the group consensus reflects genuine agreement rather than shared scoring bias, using the three verdict categories (Adhered, Not Adhered, and N/A). Where: P_o = the analyst’s Auditor Agreement Rate. P_chance = the probability the analyst and the consensus would match by chance, using each side’s verdict distribution. Kappa=P_oP_chance1P_chance\text{Kappa} = \frac{\text{P\_o} - \text{P\_chance}}{1 - \text{P\_chance}} The report displays Kappa only when the analyst scored 10 or more consensus metrics, and shows a dash otherwise.

Kappa Interpretation

Per-Conversation Agreement Rate

The Conversation-Level Summary calculates the agreement rate for each conversation using the same method as the Session Agreement Rate, scoped to that conversation.

Per-Analyst Accuracy

A recommended calculation for individual analyst accuracy is: Analyst Accuracy (%)=Number of metrics matching the Reference ScoreTotal metrics evaluated by the analyst×100\text{Analyst Accuracy (\%)} = \frac{\text{Number of metrics matching the Reference Score}}{\text{Total metrics evaluated by the analyst}} \times 100 For example, if an analyst’s scores match the Reference Score for 42 of 50 metrics: Analyst Accuracy=4250×100=84%\text{Analyst Accuracy} = \frac{42}{50} \times 100 = 84\% This result helps the Calibration Manager identify analysts who consistently score differently from the established team standard.

Most Contested Metric

The Most Contested Metric is the metric with the lowest agreement rate, shown per conversation and at the session level. When two metrics share the lowest rate, the report displays the one that appears first in the evaluation form.

Metric Calibration Gaps — Thresholds

The Metric Calibration Gaps panel sorts metrics ascending by agreement rate, lowest first, and applies these thresholds across the Calibration module.

AutoQA in Calibration Calculations

AutoQA appears in the Consensus Alignment Summary as a reference row. It isn’t included in consensus or agreement calculations based on human analyst verdicts.
  • Consensus calculation: Human analyst verdicts determine consensus. The system excludes the AutoQA verdict from N and the A, B, and C counts.
  • AutoQA Consensus Alignment: Uses the Auditor Agreement Rate formula and treats AutoQA as an additional participant. The system compares its verdict with the human-only consensus.
  • AutoQA Cohen’s Kappa: Computed using the Auditor Kappa formula, treating AutoQA as an additional analyst, with the consensus distribution drawn from human analysts only.
The report displays the AutoQA row below the human analyst rows with the following note: AutoQA evaluated against human consensus; not included in consensus calculation.

Calibration Session Lifecycle

A Calibration Session follows a structured workflow from setup to reporting. Each stage helps the Calibration Manager prepare the session, collect independent analyst evaluations, compare scoring differences, establish reference scores, and review the final calibration results.

Roles and Permissions

The Calibration tab in Audit Allocations is visible to any QA Analyst with at least one assigned Calibration Session, regardless of whether they have Calibration Manager permissions. The Calibration Session Report is available to all session participants after the session closes, including analysts who don’t have the Audit Allocations permission.