Skip to main content
The playground tells you how your agent would answer. The Evaluation center (under AI configuration in the sidebar) covers what it actually said. It collects the replies your agents sent to real customers, has the AI flag the ones worth a second look, and walks you through reviewing them. A short review session becomes lasting corrections.

Run an evaluation

Click New evaluation. Ciarem scans every reply your agents have sent since your last evaluation, across all channels (WhatsApp, Instagram, Messenger, and website chat), pairing each reply with the customer message it answered. It then keeps only the responses that deserve review and groups them into named categories, such as “Missing factual support” or “Tone too cold”, written in your workspace’s primary language. Two outcomes need no work from you. If your agents have not sent anything since the last run, you see “No new responses”. If everything it scanned looks fine, the evaluation completes with nothing flagged. The Evaluation center: evaluation history with responses evaluated, adequate versus corrected, and categories per run

Review category by category

A prepared evaluation walks you through its categories one step at a time (“1 of 3”). Each step shows the real exchanges, with the contact’s message next to the response your agent generated, and asks for one verdict on the category:
  • Adequate responses: these replies are right. Move to the next category.
  • Needs correction: something should change. You then say what, in one of two ways:
    • Improve response: describe in plain language how the agent should have answered (“don’t confirm surgery coverage, only the Premium plan includes minor surgery”). Ciarem folds that instruction into the agent’s learned behavior. If your description contains a business fact (a price, a coverage area, a policy), Ciarem says so and points you to the knowledge base instead. Facts are answered from there, not learned as behavior.
    • Should transfer: this moment belonged with a human. Ciarem creates a real transfer rule from your description, so next time the conversation reaches a person.
Every correction also leaves an automatic test behind, so Ciarem keeps checking that the fix holds as your configuration evolves. Reviewing one category: the contact's message beside the generated response, with the Adequate responses and Needs correction verdicts

The evaluation history

The Evaluation center’s home is a history of every run: when it ran, how many responses were evaluated, the split between Adequate and Corrected, how many categories you reviewed, and its status. Open any run with View results to revisit what was flagged and what you decided. Over time this table is your quality trend. A rising Adequate percentage means your agent is getting better.

Corrections from the playground too

You do not have to wait for a batch review. In the playground, every analysed reply carries Good response / Needs correction actions, and a correction opens the same two-way dialog: improve the reply or turn the moment into a transfer rule. See The playground.

When to run it

Run an evaluation after your first days live, then whenever something meaningful changes: new prices, a new agent on the team, a campaign that brings a new kind of customer. Reviewing a few categories each week is faster than reading every conversation.