A report card for a conversational agent. Any change we make,
we can test the quality against the token cost in dollars.
No batch run yet.
45 opening messages, each the first thing a person might type into HealthBot, sent through the full pipeline.
No profile, no history attached — so each test focuses on the model itself.
Example message01 · Back to Light DutyDoc clears me next week.
Should be good to go, I think.
Loading the scenario set…
Emergencies — graded hardest, most weight
Models are tested with identical conditions (same scenarios, LLM judge, settings). Any difference in the results comes only from the models.
Test previously ran and shows saved results.
Nothing runs live here.
Model performance gets graded by checking against the hard rules, passing or failing, and the LLM judge scores quality.
Voice, no banned phrases · Length · Handoff · Scope
•Passed•Failed•Not Tested
★ Judge’s Score — click for the reason
The cost of each conversation, per model.
Every number here is just cost,
not quality.
Each reply also runs a fourth model call that writes the narration for the Test Bench’s Output Console.
The grades and the costs together.
Every run is a single opening message. Longer conversations are untested.
Each run the way it would ship. The expensive one thinks privately before it decides, the cheap one can’t, and that difference requires outside knowledge.
These scores come from a model grading a model. A reply can pass every check here and still confuse the person it was written for, and when that happens, it failed.
Human review is the real benchmark.
Therapists grade these same conversations in the review console. They see the whole conversation, the gray areas, and the whole person. Without eyes like these, quality drifts and no score catches it. When machine and therapist disagree, the therapist is right.
This batch compared two models. The same test can check a new model, a reworded prompt, a new rule. Run it again and the tables tell you whether the change helped.
Run scenarios yourself
Open the Test BenchReview full conversations
Open Human Review