Evaluation map

Separate the signal that should move from the evidence that should stay put.

Dynamic evals are eight recurring Gemini-backed goal conversations. Point-in-time evals are closed incident records: useful to reproduce a named failure, deliberately excluded from the recurring score.

DynamicContinuous goal conversations
Point-in-timeFixed historical cases
Mechanical
Recurring signal

Execution mechanics

Goal adherence, first audible response, dropped turns, and turn-guard counts across completed runs.

8Gemini-backed evals — count is pinned
Case evidence

Mechanical regressions

Named failures, their original conditions, and the latest rerun evidence. Browse broadly; rerun explicitly.

223186 closed · 36 open · 1 unverified
Prompt quality
Recurring signal

Generated prompt quality

Gricean, PALBench, EQBench, faithfulness, prompt craft, and capability axes shown by run over time.

0–20Judge-axis scale, with N/A excluded
Case evidence

Prompt regression library

Historical prompt and configuration failures with their requested axes and latest judgement evidence.

223cases kept out of the recurring trend

Model bake-off remains a family, not a ninth dynamic eval. Its model arms reuse docent-full to answer which builder LLM performs best.

Open model comparisons →
Columns describe eval inputs. Rows describe result measurements.