Persona Builder · Adaptive voice bake-off

Which model should drive Charlie?

Compare the same ten reasoning-paired models plus the historical GLM 4.7 baseline, the standalone Kimi Fireworks provider arm, and two standalone non-reasoning Crusoe Gemma 4 arms (its dedicated gemma-4-31b-response deployment and the serverless google/gemma-4-31b-it catalog model), the standalone non-reasoning Wafer Gemma 4 dedicated-endpoint arms (two runs), and the same-stack Crusoe managed and Cerebras direct prod-rerun anchors, and the standalone non-reasoning Wafer GLM 5.2 arms (two runs) across two Builder architectures. Every era has 32 registered arms through the same adaptive docent-full goal on the real voice CVI path. Provider/configuration failures and unrun arms are shown explicitly and never converted into model scores.

The standings now rank product outcomes, not a count of binary gates. Autonomous handoffs and guard avoidance carry the most weight; final build completeness, discovery, reliability, and continuous first-audio latency make up the rest. The ten deterministic checks remain visible as diagnostic evidence. Click any completed card for its exact shared timeline and transcript.
Standings click a completed model to inspect its timeline and transcript
Head-to-head pick any two arms · the page link updates so the comparison is shareable
Scenario matrix

The original at-a-glance diagnostic chart, showing each reasoning arm independently. Its passed/total count is not the standings score: hard thresholds and overlapping checks would otherwise distort the ordering. A timeout or rejected provider credential is harness/configuration evidence; it is not presented as a model-quality failure.

all checks clean partial pass never called tools dropped a turn ×provider/runtime error, not scored ·not run yet
Tool-call matrix

Every tool observed in either the model’s original turn or the turn-guard recovery. Counts are actual calls, not unique field updates.

tool called never called ×provider/runtime error, not scored
Run behavior one adaptive conversation per reasoning arm
Reasoning is a per-run setting, not a model capability score. Each pair records the effective setting and its separately verified trace evidence. “Lowest” is used where the provider does not support fully disabling reasoning. Blank Charlie turns are counted from the transcript and should be interpreted alongside the config.
Scenario: docent-full. The adaptive creator pursues the same museum-docent requirements across every arm while responding to what each Charlie actually said. Every run preflights the same live readable/presentable PDF and disables the generic per-turn acknowledgement; explicit long-running draft narration remains. Standings score = 45% autonomy + 35% final build completeness + 10% discovery before drafting + 5% reliability + 5% continuous first-audio latency. Autonomy itself is 70% self-handoff coverage and 30% guard-free turns. Binary checks are diagnostic gates, not additive weights. Provider/runtime failures are excluded instead of receiving a zero. Timeline and transcript inspectors are rendered by the same core/timing_viz.py component used by the Timeline family.