Run an evaluation
Click New evaluation. Ciarem scans every reply your agents have sent since your last evaluation, across all channels (WhatsApp, Instagram, Messenger, and website chat), pairing each reply with the customer message it answered. It then keeps only the responses that deserve review and groups them into named categories, such as “Missing factual support” or “Tone too cold”, written in your workspace’s primary language. Two outcomes need no work from you. If your agents have not sent anything since the last run, you see “No new responses”. If everything it scanned looks fine, the evaluation completes with nothing flagged.
Review category by category
A prepared evaluation walks you through its categories one step at a time (“1 of 3”). Each step shows the real exchanges, with the contact’s message next to the response your agent generated, and asks for one verdict on the category:- Adequate responses: these replies are right. Move to the next category.
- Needs correction: something should change. You then say what, in one of two ways:
- Improve response: describe in plain language how the agent should have answered (“don’t confirm surgery coverage, only the Premium plan includes minor surgery”). Ciarem folds that instruction into the agent’s learned behavior. If your description contains a business fact (a price, a coverage area, a policy), Ciarem says so and points you to the knowledge base instead. Facts are answered from there, not learned as behavior.
- Should transfer: this moment belonged with a human. Ciarem creates a real transfer rule from your description, so next time the conversation reaches a person.
