> ## Documentation Index
> Fetch the complete documentation index at: https://koreai-content-gov.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Calibration Sessions

Calibration Sessions help QA teams identify and resolve scoring differences by having multiple QA Analysts independently evaluate the same set of customer conversations.

QA Analysts may score the same customer conversation differently because they interpret evaluation criteria differently. Teams lack visibility into scoring inconsistencies, analyst alignment, and adherence to team standards without a structured way to compare evaluations.

The **Calibration Manager** creates the session, selects the conversations and evaluation form, and assigns analysts. Analysts then score the conversations independently in **Calibration Mode**, where peer scores remain hidden.

After all analysts submit their evaluations, the manager uses **Debrief View** to compare scores, resolve differences, and set **Reference Scores** for contested metrics. The manager then closes the session, generating a shareable **Calibration Session Report** with agreement results, analyst accuracy, and resolution notes.

The workflow gives teams a consistent way to measure analyst alignment, establish reference outcomes, and improve how they apply evaluation criteria over time.

## Key Capabilities

Calibration Sessions provide the following capabilities.

| Capability                         | Description                                                                                                                                                               |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Create calibration sessions**    | Create a session with selected conversations, an evaluation form or scorecard, and assigned QA Analysts.                                                                  |
| **Assign QA Analysts**             | Assign multiple analysts to the same conversations so they can evaluate them independently.                                                                               |
| **Control AutoQA visibility**      | Choose whether analysts can view AutoQA scores while completing their calibration evaluations.                                                                            |
| **Calibration Mode**               | Provide a dedicated calibration experience where peer scores remain hidden during evaluation scoring.                                                                     |
| **Submission confirmation**        | Allow analysts to confirm their evaluation before submitting it as final.                                                                                                 |
| **Monitor evaluation progress**    | Track analyst submissions and conversation completion throughout the session.                                                                                             |
| Debrief and compare results        | Review analyst scores side by side for each evaluation metric after all assigned evaluations are complete.                                                                |
| **Analyze scoring gaps**           | Review metric-level disagreements, including analyst notes and AutoQA AI Justification, to understand differences in scoring.                                             |
| **Set reference score**            | Allow the manager to establish a reference score after reviewing analyst evaluations and discussion context.                                                              |
| **Measure agreement and accuracy** | Measure agreement between analysts and the reference score, and identify analysts whose scores differ from the established reference.                                     |
| **Calibration report**             | Generate a shareable or exportable report containing calibration outcomes, agreement results, analyst accuracy, conversation-level results, and manager resolution notes. |
| **Session history**                | Review closed calibration sessions and their final results.                                                                                                               |
| **Support analyst coaching**       | Use calibration results to identify scoring inconsistencies and coaching opportunities.                                                                                   |

## Prerequisites

Before creating a Calibration Session, confirm that you have the following:

| Requirement                | Description                                                                                                                                                  |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Quality AI permissions** | You must have the **Audit Allocations** permission to create and manage Calibration Sessions, access Debrief View, set Reference Scores, and close sessions. |
| **Evaluation form**        | A published **Human Agent** or **AI Agent** evaluation form must be available and suitable for the conversations you want to calibrate.                      |
| **Conversations**          | Conversations must be available for evaluation and match the selected evaluation form and configured filters.                                                |
| **QA Analysts**            | At least two QA Analysts must be available for assignment so you can compare scoring consistency.                                                            |
| **Calibration scenarios**  | Identify conversations that represent the scoring scenarios your team needs to calibrate.                                                                    |
| **AutoQA evaluation**      | If you plan to use AutoQA recommendations during the session, the selected conversations must have AutoQA evaluation results available.                      |

## Access Calibration Sessions

Navigate to **Quality AI** > **AutoQA** > **Calibration Sessions**.

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/calibration-landingpage.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=3cf7251484f7191fce318b1dd437fa73" alt="Calibration Session Landing Page" width="1903" height="852" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/calibration-landingpage.png" />

<Note> Users with both the **Audit Allocations** permission and **Cross Queue Data Access** can view calibration sessions across the entire application, not only the sessions they created.</Note>

## Calibration Sessions Landing Page

The **Calibration Sessions** landing page provides a centralized view of calibration sessions. Calibration Managers can create sessions, monitor active sessions, review closed sessions, and continue working with draft sessions.

The page includes three tabs:

| Tab        | Description                                                                    |
| ---------- | ------------------------------------------------------------------------------ |
| **All**    | Displays all calibration sessions available to the user, regardless of status. |
| **Active** | Displays calibration sessions that are still in progress.                      |
| **Closed** | Displays completed calibration sessions.                                       |

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/calibration-active.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=d6fa74701995e02a8f4332537e550605" alt="Session Status" width="1671" height="609" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/calibration-active.png" />

Use the **Search** field to find a session by name.

## Session List

The session list displays the following information:

| Column            | Description                                                                                                                                                                                                                 |
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Name**          | Displays the calibration session name.                                                                                                                                                                                      |
| **Conversations** | Displays the number of conversations included in the session.                                                                                                                                                               |
| **Auditors**      | Displays the analysts assigned to the session.                                                                                                                                                                              |
| **Created**       | Displays the session creation date.                                                                                                                                                                                         |
| **Created By**    | Identifies the user who created the session, when available.                                                                                                                                                                |
| **Status**        | Displays the current session status.                                                                                                                                                                                        |
| **Agreement**     | Displays the overall agreement percentage after the session closes.                                                                                                                                                         |
| **Actions**       | Opens the session, or provides the actions available for it based on status and your permissions. The view icon opens Active and Closed sessions in read-only or monitoring view, and the pencil icon edits Draft sessions. |

Agreement uses traffic-light coloring:

* **Green:** 90%-100%
* **Amber:** 70–89%
* **Red:** Below 70%

## Session Status

| Status     | Description                                                              |
| ---------- | ------------------------------------------------------------------------ |
| **Draft**  | The manager has started configuring the session but hasn't activated it. |
| **Active** | The session is in progress and assigned evaluations remain incomplete.   |
| **Closed** | The manager has completed the debrief and closed the session.            |

## Create a Calibration Session

Select **+ Create New Session** to start a calibration exercise.

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/create-new-calibration-session.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=a67ffbde200cc92a0d19af650d6139a4" alt="Create Calibration Session" width="1677" height="375" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/create-new-calibration-session.png" />

The wizard has three steps: **General**, **Select Conversations**, and **Assign Auditors**.

A Calibration Manager creates a session by configuring the session, selecting conversations, and assigning QA Analysts. After the manager activates the session, assigned analysts receive the calibration work in their audit queue.

### Step 1: General

Define the session configuration.

1. Enter a **Name** for the calibration exercise.
2. Enter a **Description** with instructions or context for participating analysts.
3. Turn on **Analyst Justification** > **Mandatory** to require analysts to add a comment before submitting each metric score.
4. Select an **AutoQA Visibility** option.
5. Select the **Evaluation Form** analysts will use to score the conversations.
6. Select **Next**, or **Save as Draft** to continue later.

### Configure AutoQA Visibility

The Calibration Manager can control when analysts see AutoQA recommendations during calibration.

The available options are:

| Option                                 | Description                                                                                    |
| -------------------------------------- | ---------------------------------------------------------------------------------------------- |
| **Hidden until all analysts submit**   | AutoQA recommendations remain hidden until every assigned analyst submits the session.         |
| **Visible throughout**                 | AutoQA recommendations remain visible during the complete evaluation.                          |
| **Visible after each analyst submits** | AutoQA recommendations become visible only after an individual analyst submits the evaluation. |

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/create-new-session.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=9a9d37c00dc8632ea59d4b2fcfbe38cd" alt="Create New Session" width="1908" height="937" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/create-new-session.png" />

### Step 2: Select Conversations

**Build the set of conversations analysts must evaluate**

1. Select a **Conversation Type** to define which conversations are available to add.
   * **Audited**: Displays audited conversations.
   * **Unaudited**: Displays conversations that aren't audited.
2. Select a tab to choose the conversation source:
   * **All Conversations:** Conversations that match the selected filters.
   * **Marked for Calibration:** Conversations that analysts or supervisors previously flagged using **Mark for Calibration** on the audit screen. See [Mark a Conversation for Calibration](#mark-a-conversation-for-calibration).
3. Use the **Search** field to find specific conversations.
4. Select **Filter** to narrow the list of available conversations.
5. Configure the required filters and select **Apply**.
6. Select **Select all** or individual conversations, then select **Add Interactions** then select Add Interactions to add it to the **Selected Conversations** panel.
7. Review the **Selected Conversations** panel. Select **View Conversation** to preview a conversation, or remove a conversation before continuing.
8. Select **Next**.

Managers can preview or remove conversations before continuing. They must select at least one conversation to proceed.

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/view-conversation.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=673328aab2a935f21b7e3bdedcd5b94e" alt="Select Conversations" width="1911" height="891" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/view-conversation.png" />

#### Conversation Filters

The filter panel helps narrow conversations before adding them to the session.

| Filter              | Description                                               |
| ------------------- | --------------------------------------------------------- |
| **Date Range**      | Filters conversations using evaluation date.              |
| **Evaluation Form** | Displays conversations evaluated with the selected form.  |
| **Queues**          | Filters conversations by queue or experience flow.        |
| **Agents**          | Filters conversations by evaluated agents.                |
| **Duration**        | Filters conversations using minimum and maximum duration. |
| **AutoQA Score**    | Filters conversations using an AutoQA score range.        |
| **Sentiment Score** | Filters conversations using a sentiment score range.      |

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/new-session-conversation-selection-filter.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=27f574557caeec8de97a871e4dc48fad" alt="Select Conversations" width="1908" height="934" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/new-session-conversation-selection-filter.png" />

#### Mark a Conversation for Calibration

Supervisors and auditors can flag a conversation for calibration directly from the standard audit screen, without opening the Calibration workflow first.

1. Open a conversation on the audit screen.
2. Select the more options menu (**⋮**) in the header.
3. Select **Mark for Calibration**.

Flagged conversations appear under the **Marked for Calibration** tab in [Step 2: Select Conversations](#step-2-select-conversations) of session creation, so a Calibration Manager can add them to a session later.

### Step 3: Assign Auditors (QA Analysts)

Assign the QA Analysts who can evaluate every selected conversation independently.

1. Use Search Auditors to Add to find and select one or more analysts.
2. Review the selected analysts listed in the Auditors section. Remove an analyst using the delete icon if needed.
3. Each assigned analyst receives the same set of conversations and evaluates them independently.
4. Select **Save as Draft** to save the session and continue configuring it later.
5. Select **Create and Notify Participants** to create and activate the session, or **Save as Draft** to continue later.
6. Select **Save** to save the current configuration.

<Note> A Calibration Session requires a minimum of two analysts. Individual evaluations remain hidden from other analysts until the session reaches the comparison stage.</Note>

Selecting **Create and Notify Participants** sets the session status to **Active** and sends an email notification to each assigned analyst.

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/assign-auditors.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=3f1246f2cc744b62aca766ae391bfaac" alt="Select Conversations" width="1905" height="895" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/assign-auditors.png" />

## My Calibration Sessions

### Access Assigned Calibration Sessions

QA Analysts access assigned Calibration Sessions from the Calibration tab in Audit Allocations. The system also sends an email notification when it assigns a session to an analyst.

The **Calibration** tab appears when the analyst has at least one assigned session. Each session displays the session name, assigned-by name, scoring progress (for example, 2 of 4 conversations scored), and session status (Not Started, In Progress, or Submitted).

Each session displays the Session name, Assigned-by name, Scoring progress (for example **2 of 4 conversations scored**), and Session status: (**Not Started**, **In Progress**, or **Submitted**).

From an assigned session, an analyst can:

* Review the conversations included in the session, with a per-conversation status of **Score Now**, **Scored**, or **Locked**.
* Monitor evaluation progress.
* Open pending conversations and complete calibration evaluations.
* Review the calibration report after the manager closes the session.

## Monitor an Active Session

Opening an active session from the landing page displays the manager's monitoring view.

The session header displays the **Created Date**, **Auditors** avatar stack, **Conducted by**, and **Status**.

* **Submission Progress**: Shows the overall submission count (for example, 8/8 Submitted) and a card for each assigned analyst. Each card displays the analyst's name, number of Conversations Scored, and a progress bar. The card turns green when the analyst reaches 100%.

* **Conversation-Level Summary**: Lists every conversation in the session with the following details:

| Column              | Description                                                                                                                     |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| **Handled By**      | Identifies the agent who handled the conversation.                                                                              |
| **Flows/Queues**    | Displays the Experience Flow for conversations handled by an AI Agent and the queue for conversations handled by a Human Agent. |
| **Conversation ID** | Identifies the conversation.                                                                                                    |
| **Date**            | Displays the conversation date.                                                                                                 |
| **Duration**        | Displays the conversation duration.                                                                                             |
| **AutoQA Score**    | Displays the AutoQA score when available.                                                                                       |

<Note>:The **Flows/Queues** column displays the Experience Flow for conversations handled by an AI Agent, and the queue for conversations handled by a Human Agent.</Note>

## Calibration Evaluation

**Calibration Mode** is the scoring experience for QA Analysts participating in a Calibration Session. It uses the selected evaluation form and keeps peer scores hidden so analysts can independently evaluate each conversation.

A **Calibration Mode** indicator identifies that the analyst is completing a calibration evaluation rather than a production audit.

During calibration:

* Peer scores remain hidden.
* AutoQA visibility follows the session's configured **AutoQA Visibility** setting.
* Analysts evaluate each metric independently.
* Add comments when **Analyst Justification** has the **Mandatory** option enabled.
* **Audit Progress** displays evaluation completion status.
* Analysts can't modify an evaluation after submission.
* Calibration Session conversations aren't routed to the agent for review, acknowledgement, or dispute.

### Complete a Calibration Evaluation

1. Open an assigned Calibration Session.
2. Select a pending conversation.
3. Review the conversation transcript.
4. Evaluate each metric using the selected evaluation form.
5. Add notes to provide additional scoring context.
6. Review the completed evaluation.
7. Confirm that the evaluation is final.
8. Select **Submit Score**.

After submission, the evaluation becomes final. The analyst can't modify it unless a separate correction workflow is available. Submitted calibration scores don't affect agent dashboards, leaderboard rankings, or scorecard scores.

## Debrief View

After all assigned analysts submit their evaluations, the Calibration Manager can open **Debrief View** to compare analyst scores side by side for each evaluation metric before setting a reference score.

Debrief View opens from a conversation in the session and displays a metric-by-metric comparison grid. The closed Calibration Session Report uses the same grid when you select the view icon for a conversation in the Conversation-Level Summary, as described under [Calibration Session Report](#calibration-session-report).

### Debrief View Components

| Component                   | Description                                                                                                                                                         |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Conversation**            | Identifies the conversation under review. Conversations display as tabs, each summarized by Handled By, Flows/Queues, Agreement Rate, and Most Disputed Metric.     |
| Metric                      | Displays the evaluation criterion that analysts compare.                                                                                                            |
| **Analyst scores**          | Displays each analyst's score response for the metric, shown side by side.                                                                                          |
| **Agreement indicator**     | Identifies whether analysts reached the same result or have differing scores.                                                                                       |
| **AutoQA Verdict**          | Displays the AutoQA result for the metric, when available. Displays as **Adhered**, **Not Adhered**, **N/A**, or **Not Found** when AutoQA didn't produce a result. |
| **Expand**                  | Expands a metric to provide additional scoring context.                                                                                                             |
| **Accuracy**                | Displays the metric-level accuracy result, when available.                                                                                                          |
| **AutoQA AI Justification** | Displays the AI-generated justification associated with the AutoQA result when available.                                                                           |
| **Reference score**         | Records the manager's established score for the metric.                                                                                                             |
| **Resolution notes**        | Records the manager's explanation for the final resolution.                                                                                                         |

A **Participant Summary** at the bottom of Debrief View shows, for each analyst: Metrics Scored, Aligned with Consensus (%), and Accuracy.

### Review a Disagreement

1. Open **Debrief View** after all required analyst evaluations are complete.
2. Review the analyst scores for each metric.
3. Identify metrics where the scores differ.
4. Expand the metric to review analyst notes and AutoQA AI Justification.
5. Determine the appropriate reference score.
6. Add resolution notes explaining the decision.
7. Continue until you resolve all required disagreements.

### Reference Score

The **Reference Score** represents the Calibration Manager's final determination for a metric after reviewing the analyst scores and supporting context.

The manager uses the Reference Score to establish the expected scoring outcome for the calibration exercise. The reference score should reflect the team's agreed interpretation of the evaluation criterion.

## Close a Calibration Session

The Calibration Manager closes the session after reviewing the submitted evaluations and completing the debrief.

Before closing the session, the manager should:

1. Review all analyst scores.
2. Resolve the required scoring differences.
3. Set reference scores.
4. Add resolution notes.
5. Review the calibration results.
6. Close the session.

After the manager closes the session, the final calibration results become available to participants through the session report.

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/calibration-close-status.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=52a2bbf56d7252d7fa20edba9639b4b6" alt="Select Conversations" width="1681" height="205" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/calibration-close-status.png" />

## Calibration Session Report

Closing a session generates a read-only **Calibration Session Report**, which is available to the Calibration Manager and all participating QA Analysts. The report summarizes calibration outcomes, analyst alignment, scoring consistency, and areas that require additional calibration.

The report header displays the **session name**, **date closed**, **conducted by**, **participants**, and **conversation count**.

### Report Sections

| Section                         | Description                                                                                                                                                                                                                                                                     |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Agreement Breakdown**         | Displays the distribution of **Full Agreement**, **Partial Agreement**, and **No Agreement** results, along with the **Overall Agreement Rate** for the session.                                                                                                                |
| **Metric Calibration Gaps**     | Lists evaluation metrics ranked by agreement rate, from lowest to highest, and identifies metrics that may require additional calibration.                                                                                                                                      |
| **Conversation-Level Summary**  | Displays each conversation included in the session, along with its agreement results and the most disputed metric. Select the **View** icon to expand a conversation and review detailed metric-level comparisons.                                                              |
| **Consensus Alignment Summary** | Displays per-analyst results for Parameters Evaluated, Consensus Alignment, Inter-Auditor Agreement, and Accuracy. An AutoQA benchmark row appears below the analyst rows and compares AutoQA results with human consensus without including them in the consensus calculation. |

### Conversation-Level Comparison

Select the view icon on a **Conversation-Level Summary** row to expand the row inline and reveal the metric-by-analyst comparison grid directly beneath it, without leaving the report.

This expanded grid uses the same structure as [Debrief View](#debrief-view).

#### Accuracy

In the reviewed report data, **Accuracy** measures how well an analyst's scores match the final **Reference Scores** established by the Calibration Manager during Debrief.

For example, if an analyst's scores match the Reference Scores for **8 out of 10 metrics**, their accuracy is **80%**.

If the Calibration Manager hasn't established Reference Scores, the report can't calculate Accuracy.

### Export the Report

Select [Export CSV](https://raw.githubusercontent.com/Koredotcom/docs-v2/refs/heads/main/ai-for-service/assets/agentai-interaction-calibration.csv) to download the calibration results as a CSV file.

### Exported CSV Fields

The exported CSV contains calibration results from the report.

| **Field**                | **Description**                                                         |
| ------------------------ | ----------------------------------------------------------------------- |
| **Session Details**      | Calibration session name and relevant session information.              |
| **Conversation**         | Conversations included in the calibration session.                      |
| **Analyst Scores**       | Scores submitted by participating QA Analysts.                          |
| **Reference Scores**     | Final scores established by the Calibration Manager during Debrief.     |
| **Agreement Rate**       | Agreement result for the calibration session.                           |
| **Per-Analyst Accuracy** | Accuracy of each analyst compared with the established reference score. |
| **Resolution Notes**     | Manager's explanation for resolved scoring differences.                 |
| **Calibration Outcome**  | Final results of the calibration exercise.                              |

<Note> The exact CSV column names and export structure may vary based on the final report implementation.</Note>

<img src="https://mintcdn.com/koreai-content-gov/tcqMVqczuj0NFd3C/ai-for-service/quality-ai/analyze/calibration-sessions/images/closed-session-view.png?fit=max&auto=format&n=tcqMVqczuj0NFd3C&q=85&s=aa93ec1bf2c39395f077e36be8cea32e" alt="Select Conversations" width="1896" height="946" data-path="ai-for-service/quality-ai/analyze/calibration-sessions/images/closed-session-view.png" />

QA Analysts can use the report to identify differences between their evaluations and the final Reference Scores and review any **Resolution Notes** from the calibration process.

The report also compares **AutoQA recommendations** with the final consensus to help QA Analysts assess AutoQA's alignment with human evaluators.

## Calibration Results and Formulas

After the manager completes the debrief, the system calculates calibration results.

The results provide an overall view of scoring consistency and identify differences between individual analyst evaluations and the established Reference Score.

Every calculation treats N/A as a third verdict alongside Adhered and Not Adhered. The system includes every analyst response in each metric's denominator.

The system calculates agreement and alignment metrics using the following calculations.

### Variables

| Variable    | Definition                                                                                                     |
| ----------- | -------------------------------------------------------------------------------------------------------------- |
| **N**       | Total number of analysts in the session.                                                                       |
| **A, B, C** | Number of analysts who responded **Adhered**, **Not Adhered**, and **N/A** on a given metric. `A + B + C = N`. |
| **M**       | Number of metrics per conversation.                                                                            |
| **C**       | Number of conversations in the session.                                                                        |

### Metric Agreement Rate

A metric reaches consensus when one verdict receives a majority of analyst responses:

$$
\text{count} > N / 2
$$

Otherwise, the system labels the metric **No Consensus** and excludes it from further calculations.

$$
\text{Metric Agreement Rate} = \frac{\max(A, B, C)}{N}
$$

The system computes the Metric Agreement Rate only when the metric reaches consensus.

### Session Agreement Rate

The Session Agreement Rate averages the agreement rates for all metrics that reach consensus. The system excludes metrics labeled No Consensus.

sum of all metric agreement rates / number of consensus metrics

$$
\text{Session Agreement Rate} = \frac{\text{sum of all metric agreement rates}}{\text{number of consensus metrics}}
$$

### Agreement Distribution

The system groups metrics into the following agreement categories:

| Category               | Condition                                             |
| ---------------------- | ----------------------------------------------------- |
| **Full Agreement**     | Metric Agreement Rate = **100%**                      |
| **Majority Agreement** | Metric Agreement Rate is between **50%** and **100%** |
| **No Consensus**       | No verdict receives more than **N / 2** responses.    |

### Session N/A Rate

The Session N/A Rate measures the proportion of metric responses that analysts marked N/A.

$$
\text{Session N/A Rate} = \frac{\text{total N/A verdicts across all metrics}}{C \times M \times N}
$$

### Auditor Agreement Rate (Aligned with Consensus)

For each analyst, the proportion of their verdicts that matched the consensus verdict, excluding **No Consensus** metrics from both the numerator and denominator.

$$
\text{Auditor Agreement Rate} = \frac{\text{number of matching verdicts}}{\text{number of consensus metrics}}
$$

### Auditor N/A Rate

The **Auditor N/A Rate** measures the proportion of metrics for which an analyst selected **N/A**.

$$
\text{Auditor N/A Rate} = \frac{\text{metrics where the analyst responded N/A}}{C \times M}
$$

The Participant Summary flags an analyst when their N/A rate exceeds the session average by more than 15 percentage points.

### Auditor Inter-Auditor Agreement (Cohen's Kappa)

Measures whether an analyst's alignment with the group consensus reflects genuine agreement rather than shared scoring bias, using the three verdict categories (`Adhered`, `Not Adhered`, and `N/A`).

Where:

`P_o` = the analyst's Auditor Agreement Rate.

`P_chance` = the probability the analyst and the consensus would match by chance, using each side's verdict distribution.

$$
\text{Kappa} = \frac{\text{P\_o} - \text{P\_chance}}{1 - \text{P\_chance}}
$$

The report displays Kappa only when the analyst scored 10 or more consensus metrics, and shows a dash otherwise.

### Kappa Interpretation

| Kappa Range   | Interpretation |
| ------------- | -------------- |
| **Below 0.4** | Poor           |
| **0.4–0.6**   | Moderate       |
| **0.6–0.8**   | Substantial    |
| **Above 0.8** | Near-perfect   |

### Per-Conversation Agreement Rate

The **Conversation-Level Summary** calculates the agreement rate for each conversation using the same method as the **Session Agreement Rate**, scoped to that conversation.

### Per-Analyst Accuracy

A recommended calculation for individual analyst accuracy is:

$$
\text{Analyst Accuracy (\%)} = \frac{\text{Number of metrics matching the Reference Score}}{\text{Total metrics evaluated by the analyst}} \times 100
$$

For example, if an analyst's scores match the Reference Score for 42 of 50 metrics:

$$
\text{Analyst Accuracy} = \frac{42}{50} \times 100 = 84\%
$$

This result helps the Calibration Manager identify analysts who consistently score differently from the established team standard.

### Most Contested Metric

The Most Contested Metric is the metric with the lowest agreement rate, shown per conversation and at the session level.

When two metrics share the lowest rate, the report displays the one that appears first in the evaluation form.

### Metric Calibration Gaps — Thresholds

The **Metric Calibration Gaps** panel sorts metrics ascending by agreement rate, lowest first, and applies these thresholds across the Calibration module.

| Label                    | Agreement Rate | Color |
| ------------------------ | -------------- | ----- |
| **Calibrated**           | 80% or above   | Green |
| **Needs Attention**      | 65–79%         | Amber |
| **Calibration Required** | Below 65%      | Red   |

### AutoQA in Calibration Calculations

AutoQA appears in the **Consensus Alignment Summary** as a reference row. It isn't included in consensus or agreement calculations based on human analyst verdicts.

* **Consensus calculation**: Human analyst verdicts determine consensus. The system excludes the AutoQA verdict from **N** and the **A**, **B**, and **C** counts.
* **AutoQA Consensus Alignment**: Uses the Auditor Agreement Rate formula and treats AutoQA as an additional participant. The system compares its verdict with the human-only consensus.
* **AutoQA Cohen's Kappa:** Computed using the Auditor Kappa formula, treating AutoQA as an additional analyst, with the consensus distribution drawn from human analysts only.

The report displays the AutoQA row below the human analyst rows with the following note:

`AutoQA evaluated against human consensus; not included in consensus calculation`.

## Calibration Session Lifecycle

A Calibration Session follows a structured workflow from setup to reporting.

| Stage       | Calibration Manager                                   | QA Analyst                                         |
| ----------- | ----------------------------------------------------- | -------------------------------------------------- |
| **Create**  | Creates the session and selects the evaluation scope. | —                                                  |
| **Assign**  | Assigns conversations to analysts.                    | Receives the assignment and email notification.    |
| **Score**   | Monitors evaluation progress.                         | Scores assigned conversations in Calibration Mode. |
| **Submit**  | Waits for all required submissions.                   | Completes and submits the final evaluation.        |
| **Debrief** | Compares scores and reviews disagreements.            | —                                                  |
| **Resolve** | Sets Reference Scores and Resolution Notes.           | —                                                  |
| **Results** | Reviews agreement and analyst accuracy.               | —                                                  |
| **Close**   | Closes the session and generates the report.          | Reviews the closed-session report.                 |

Each stage helps the Calibration Manager prepare the session, collect independent analyst evaluations, compare scoring differences, establish reference scores, and review the final calibration results.

## Roles and Permissions

| Permission                                      | Grants Access To                                                                                                                                        |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Audit Allocations**                           | The **Calibration Sessions** navigation item, creating and managing Calibration Sessions, Debrief View, setting Reference Scores, and closing sessions. |
| **Audit Allocations + Cross Queue Data Access** | Calibration Sessions created across the application, rather than only sessions created by the user.                                                     |

The **Calibration** tab in **Audit Allocations** is visible to any QA Analyst with at least one assigned Calibration Session, regardless of whether they have Calibration Manager permissions.

The **Calibration Session Report** is available to all session participants after the session closes, including analysts who don't have the **Audit Allocations** permission.
