> ## Documentation Index
> Fetch the complete documentation index at: https://help.ciarem.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# The Evaluation center: reviewing real replies

> Evaluate the responses your agent actually sent, approve or correct them by category, and turn your feedback into lasting behavior.

The playground tells you how your agent *would* answer. The **Evaluation center** (under **AI configuration** in the sidebar) covers what it *actually* said. It collects the replies your agents sent to real customers, has the AI flag the ones worth a second look, and walks you through reviewing them. A short review session becomes lasting corrections.

## Run an evaluation

Click **New evaluation**. Ciarem scans every reply your agents have sent **since your last evaluation**, across all channels (WhatsApp, Instagram, Messenger, and website chat), pairing each reply with the customer message it answered. It then keeps only the responses that deserve review and groups them into named **categories**, such as "Missing factual support" or "Tone too cold", written in your workspace's primary language.

Two outcomes need no work from you. If your agents have not sent anything since the last run, you see "No new responses". If everything it scanned looks fine, the evaluation completes with nothing flagged.

<img src="https://mintcdn.com/ciaremaai/cNtTcjF_a4GWWvTG/images/ciarem-help-evaluation.png?fit=max&auto=format&n=cNtTcjF_a4GWWvTG&q=85&s=c42128c8d67c6c4a73813ca4c90a370b" alt="The Evaluation center: evaluation history with responses evaluated, adequate versus corrected, and categories per run" width="1680" height="1114" data-path="images/ciarem-help-evaluation.png" />

## Review category by category

A prepared evaluation walks you through its categories one step at a time ("1 of 3"). Each step shows the real exchanges, with the contact's message next to the response your agent generated, and asks for one verdict on the category:

* **Adequate responses**: these replies are right. Move to the next category.
* **Needs correction**: something should change. You then say what, in one of two ways:
  * **Improve response**: describe in plain language how the agent should have answered ("don't confirm surgery coverage, only the Premium plan includes minor surgery"). Ciarem folds that instruction into the agent's learned behavior. If your description contains a business fact (a price, a coverage area, a policy), Ciarem says so and points you to the [knowledge base](/ai-agent/knowledge-bases) instead. Facts are answered from there, not learned as behavior.
  * **Should transfer**: this moment belonged with a human. Ciarem creates a real [transfer rule](/ai-agent/human-handoff-and-escalation) from your description, so next time the conversation reaches a person.

Every correction also leaves an automatic test behind, so Ciarem [keeps checking](/ai-agent/agent-health) that the fix holds as your configuration evolves.

<img src="https://mintcdn.com/ciaremaai/cNtTcjF_a4GWWvTG/images/ciarem-help-evaluation-review.png?fit=max&auto=format&n=cNtTcjF_a4GWWvTG&q=85&s=bcf4ddabad8742c143ca3d2a94fcd887" alt="Reviewing one category: the contact's message beside the generated response, with the Adequate responses and Needs correction verdicts" width="1680" height="1114" data-path="images/ciarem-help-evaluation-review.png" />

## The evaluation history

The Evaluation center's home is a history of every run: when it ran, how many responses were evaluated, the split between **Adequate** and **Corrected**, how many categories you reviewed, and its status. Open any run with **View results** to revisit what was flagged and what you decided. Over time this table is your quality trend. A rising Adequate percentage means your agent is getting better.

## Corrections from the playground too

You do not have to wait for a batch review. In the playground, every analysed reply carries **Good response** / **Needs correction** actions, and a correction opens the same two-way dialog: improve the reply or turn the moment into a transfer rule. See [The playground](/ai-agent/playground).

## When to run it

Run an evaluation after your first days live, then whenever something meaningful changes: new prices, a new agent on the team, a campaign that brings a new kind of customer. Reviewing a few categories each week is faster than reading every conversation.
