Persona Builder · Adaptive voice bake-off
Compare the same ten reasoning-paired models plus the historical GLM 4.7 baseline, the standalone Kimi Fireworks provider arm, and two standalone non-reasoning Crusoe Gemma 4 arms (its dedicated gemma-4-31b-response deployment and the serverless google/gemma-4-31b-it catalog model), the standalone non-reasoning Wafer Gemma 4 dedicated-endpoint arms (two runs), and the same-stack Crusoe managed and Cerebras direct prod-rerun anchors, and the standalone non-reasoning Wafer GLM 5.2 arms (two runs) across two Builder architectures. Every era has 32 registered arms through the same adaptive docent-full goal on the real voice CVI path. Provider/configuration failures and unrun arms are shown explicitly and never converted into model scores.
The original at-a-glance diagnostic chart, showing each reasoning arm independently. Its passed/total count is not the standings score: hard thresholds and overlapping checks would otherwise distort the ordering. A timeout or rejected provider credential is harness/configuration evidence; it is not presented as a model-quality failure.
Every tool observed in either the model’s original turn or the turn-guard recovery. Counts are actual calls, not unique field updates.
tavus-glm-4.7 managed baseline alias
{
"model": "tavus-glm-4.7",
"speculative_inference": true
}
Evidence: the managed baseline alias is explicit; no reasoning control or trace is exposed
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
tavus-gemma-4-thinking managed alias (Cerebras)
{
"model": "tavus-gemma-4-thinking",
"speculative_inference": true
}
Evidence: 685 reasoning tokens in bake-off conversation c62f39421e7d8477
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
tavus-gemma-4 managed alias (Cerebras)
{
"model": "tavus-gemma-4",
"speculative_inference": true
}
Evidence: non-thinking managed alias selected explicitly
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-2.5-flash via Google AI Studio
{
"model": "gemini-2.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/"
}
Evidence: 285 hidden thought tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-2.5-flash via Google AI Studio
{
"model": "gemini-2.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: 0 hidden thought tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-3.5-flash via Google AI Studio
{
"model": "gemini-3.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/"
}
Evidence: 261 hidden thought tokens and a thought signature in the matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-3.5-flash via Google AI Studio
{
"model": "gemini-3.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
"extra_body": {
"reasoning_effort": "minimal"
}
}
Evidence: 0 billed hidden tokens on a simple probe, but a thought signature remained
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-luna via OpenAI, reasoning_effort=medium
{
"model": "gpt-5.6-luna",
"speculative_inference": true,
"base_url": "https://forecast-accessible-relationships-ebook.trycloudflare.com/v1",
"extra_body": {
"reasoning_effort": "medium"
}
}
Evidence: 26 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-luna via OpenAI, reasoning_effort=none
{
"model": "gpt-5.6-luna",
"speculative_inference": true,
"base_url": "https://api.openai.com/v1",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: 0 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-sol via OpenAI, reasoning_effort=high
{
"model": "gpt-5.6-sol",
"speculative_inference": true,
"base_url": "https://forecast-accessible-relationships-ebook.trycloudflare.com/v1",
"extra_body": {
"reasoning_effort": "high"
}
}
Evidence: 3400 reasoning tokens on a complex matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-sol via OpenAI, reasoning_effort=none
{
"model": "gpt-5.6-sol",
"speculative_inference": true,
"base_url": "https://api.openai.com/v1",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: 0 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-opus-4-6 via Anthropic, extended thinking
{
"model": "claude-opus-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/",
"extra_body": {
"thinking": {
"type": "enabled",
"budget_tokens": 1024
}
}
}
Evidence: signed thinking block and 41 thinking tokens in the native trace
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-opus-4-6 via Anthropic
{
"model": "claude-opus-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/"
}
Evidence: no thinking parameter was sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-sonnet-4-6 via Anthropic, extended thinking
{
"model": "claude-sonnet-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/",
"extra_body": {
"thinking": {
"type": "enabled",
"budget_tokens": 1024
}
}
}
Evidence: signed thinking block and 8 thinking tokens in the native trace
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-sonnet-4-6 via Anthropic
{
"model": "claude-sonnet-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/"
}
Evidence: no thinking parameter was sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
moonshotai/kimi-k3 via Cloudflare
{
"model": "moonshotai/kimi-k3",
"speculative_inference": true,
"base_url": "https://api.cloudflare.com/client/v4/accounts/19cbe0d2e934c6dc43052886b7cbd8f1/ai/v1"
}
Evidence: reasoning_content and non-zero reasoning tokens in matching probes
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
moonshotai/kimi-k3 via Cloudflare, reasoning_effort=none
{
"model": "moonshotai/kimi-k3",
"speculative_inference": true,
"base_url": "https://api.cloudflare.com/client/v4/accounts/19cbe0d2e934c6dc43052886b7cbd8f1/ai/v1",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: no reasoning_content or reasoning tokens in the matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
accounts/fireworks/routers/kimi-k3-fast via Fireworks
{
"model": "accounts/fireworks/routers/kimi-k3-fast",
"speculative_inference": true,
"base_url": "https://api.fireworks.ai/inference/v1"
}
Evidence: reasoning_content and non-zero reasoning tokens in the matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
mercury-2 medium reasoning via Inception
{
"model": "mercury-2 medium reasoning via Inception",
"speculative_inference": true
}
Evidence: 44 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
mercury-2 instant reasoning via Inception
{
"model": "mercury-2 instant reasoning via Inception",
"speculative_inference": true
}
Evidence: 5 reasoning tokens remained in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
grok-4.5 via xAI
{
"model": "grok-4.5",
"speculative_inference": true,
"base_url": "https://api.x.ai/v1"
}
Evidence: 192 reasoning tokens and reasoning_content in the matching high probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
grok-4.5 via xAI, reasoning_effort=low
{
"model": "grok-4.5",
"speculative_inference": true,
"base_url": "https://api.x.ai/v1",
"extra_body": {
"reasoning_effort": "low"
}
}
Evidence: 92 reasoning tokens and reasoning_content remained in the low probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
tavus-glm-4.7 managed baseline alias
{
"model": "tavus-glm-4.7",
"speculative_inference": true
}
Evidence: the managed baseline alias is explicit; no reasoning control or trace is exposed
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
tavus-gemma-4-thinking managed alias (Cerebras)
{
"model": "tavus-gemma-4-thinking",
"speculative_inference": true
}
Evidence: 685 reasoning tokens in bake-off conversation c62f39421e7d8477
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
tavus-gemma-4 managed alias (Cerebras)
{
"model": "tavus-gemma-4",
"speculative_inference": true
}
Evidence: non-thinking managed alias selected explicitly
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-2.5-flash via Google AI Studio
{
"model": "gemini-2.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/"
}
Evidence: 285 hidden thought tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-2.5-flash via Google AI Studio
{
"model": "gemini-2.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: 0 hidden thought tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-3.5-flash via Google AI Studio
{
"model": "gemini-3.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/"
}
Evidence: 261 hidden thought tokens and a thought signature in the matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemini-3.5-flash via Google AI Studio
{
"model": "gemini-3.5-flash",
"speculative_inference": true,
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai/",
"extra_body": {
"reasoning_effort": "minimal"
}
}
Evidence: 0 billed hidden tokens on a simple probe, but a thought signature remained
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-luna via OpenAI, reasoning_effort=medium
{
"model": "gpt-5.6-luna",
"speculative_inference": true,
"base_url": "https://coaches-bones-railway-historical.trycloudflare.com/v1",
"extra_body": {
"reasoning_effort": "medium"
}
}
Evidence: 26 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-luna via OpenAI, reasoning_effort=none
{
"model": "gpt-5.6-luna",
"speculative_inference": true,
"base_url": "https://api.openai.com/v1",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: 0 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-sol via OpenAI, reasoning_effort=high
{
"model": "gpt-5.6-sol",
"speculative_inference": true,
"base_url": "https://coaches-bones-railway-historical.trycloudflare.com/v1",
"extra_body": {
"reasoning_effort": "high"
}
}
Evidence: 3400 reasoning tokens on a complex matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gpt-5.6-sol via OpenAI, reasoning_effort=none
{
"model": "gpt-5.6-sol",
"speculative_inference": true,
"base_url": "https://api.openai.com/v1",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: 0 reasoning tokens in the matching direct probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-opus-4-6 via Anthropic, extended thinking
{
"model": "claude-opus-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/",
"extra_body": {
"thinking": {
"type": "enabled",
"budget_tokens": 1024
}
}
}
Evidence: signed thinking block and 41 thinking tokens in the native trace
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-opus-4-6 via Anthropic
{
"model": "claude-opus-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/"
}
Evidence: no thinking parameter was sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-sonnet-4-6 via Anthropic, extended thinking
{
"model": "claude-sonnet-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/",
"extra_body": {
"thinking": {
"type": "enabled",
"budget_tokens": 1024
}
}
}
Evidence: signed thinking block and 8 thinking tokens in the native trace
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
claude-sonnet-4-6 via Anthropic
{
"model": "claude-sonnet-4-6",
"speculative_inference": true,
"base_url": "https://api.anthropic.com/v1/"
}
Evidence: no thinking parameter was sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
moonshotai/kimi-k3 via Cloudflare
{
"model": "moonshotai/kimi-k3",
"speculative_inference": true,
"base_url": "https://api.cloudflare.com/client/v4/accounts/19cbe0d2e934c6dc43052886b7cbd8f1/ai/v1"
}
Evidence: reasoning_content and non-zero reasoning tokens in matching probes
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
moonshotai/kimi-k3 via Cloudflare, reasoning_effort=none
{
"model": "moonshotai/kimi-k3",
"speculative_inference": true,
"base_url": "https://api.cloudflare.com/client/v4/accounts/19cbe0d2e934c6dc43052886b7cbd8f1/ai/v1",
"extra_body": {
"reasoning_effort": "none"
}
}
Evidence: no reasoning_content or reasoning tokens in the matching probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
grok-4.5 via xAI
{
"model": "grok-4.5",
"speculative_inference": true,
"base_url": "https://api.x.ai/v1"
}
Evidence: 192 reasoning tokens and reasoning_content in the matching high probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
grok-4.5 via xAI, reasoning_effort=low
{
"model": "grok-4.5",
"speculative_inference": true,
"base_url": "https://api.x.ai/v1",
"extra_body": {
"reasoning_effort": "low"
}
}
Evidence: 92 reasoning tokens and reasoning_content remained in the low probe
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemma-4-31b-response via Crusoe
{
"model": "gemma-4-31b-response",
"speculative_inference": true,
"base_url": "https://api.inference.crusoecloud.com/v1"
}
Evidence: reasoning was null in the matching direct response and reasoning_effort was not sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
google/gemma-4-31b-it via Crusoe serverless
{
"model": "google/gemma-4-31b-it",
"speculative_inference": true,
"base_url": "https://api.inference.crusoecloud.com/v1/"
}
Evidence: tool-calling probe returned finish_reason=tool_calls with no reasoning_content and reasoning_effort was not sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
google/gemma-4-31b-it via Crusoe serverless
{
"model": "google/gemma-4-31b-it",
"speculative_inference": true,
"base_url": "https://api.inference.crusoecloud.com/v1/"
}
Evidence: tool-calling probe returned finish_reason=tool_calls with no reasoning_content and reasoning_effort was not sent
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
google/gemma-4-31b-it via Crusoe serverless, throttled by the provider
{
"model": "google/gemma-4-31b-it",
"speculative_inference": true,
"base_url": "https://api.inference.crusoecloud.com/v1/"
}
Evidence: tool-calling probe returned finish_reason=tool_calls with no reasoning_content and reasoning_effort was not sent; the provider returned HTTP 429 max_requests_limit during this run; 12 turns produced no output
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
Gemma-4-31B via Wafer dedicated B200 endpoint, prod RQH
{
"model": "Gemma-4-31B",
"speculative_inference": true,
"base_url": "https://tavus.wafer.ai/v1",
"extra_body": {
"thinking": {
"type": "disabled"
}
}
}
Evidence: default and thinking-disabled probes returned no reasoning_content with completion_tokens_details.reasoning_tokens=0; a thinking-disabled tool-calling probe returned finish_reason=tool_calls with valid arguments
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
Gemma-4-31B via Wafer dedicated B200 endpoint, prod RQH
{
"model": "Gemma-4-31B",
"speculative_inference": true,
"base_url": "https://tavus.wafer.ai/v1",
"extra_body": {
"thinking": {
"type": "disabled"
}
}
}
Evidence: same-day probe returned reasoning_tokens=0 by default and with thinking disabled, and a thinking-disabled tool probe returned finish_reason=tool_calls with valid arguments
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
tavus-gemma-4 managed alias, served by Crusoe gemma-4-responsive, prod RQH
{
"model": "tavus-gemma-4",
"speculative_inference": true
}
Evidence: every logged request for conversation cb0ccde83c6a84fc was served by [email protected] on attempt 1 with no fallback and no reasoning parameter
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
gemma-4-31b via the Cerebras API directly, prod RQH
{
"model": "gemma-4-31b",
"speculative_inference": true,
"base_url": "https://api.cerebras.ai/v1"
}
Evidence: probe returned completion_tokens_details.reasoning_tokens=0 with no reasoning parameter sent; the run sends none
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
GLM-5.2 via Wafer dedicated B200 endpoint, prod RQH
{
"model": "GLM-5.2",
"speculative_inference": true,
"base_url": "https://tavus.wafer.ai/v1",
"extra_body": {
"thinking": {
"type": "disabled"
}
}
}
Evidence: default, thinking-disabled, and reasoning_effort=none probes all returned reasoning_tokens=0 with clean content; a thinking-enabled probe cleanly emitted reasoning_content (169 reasoning tokens), so off-by-default is configuration, not a capability gap; a thinking-disabled streamed two-tool probe returned valid per-call JSON
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
GLM-5.2 via Wafer dedicated B200 endpoint, prod RQH
{
"model": "GLM-5.2",
"speculative_inference": true,
"base_url": "https://tavus.wafer.ai/v1",
"extra_body": {
"thinking": {
"type": "disabled"
}
}
}
Evidence: same-day probe returned reasoning_tokens=0 by default and with thinking disabled, and a thinking-disabled tool probe returned finish_reason=tool_calls with valid arguments
A blank turn with a small max_tokens or reasoning budget is configuration evidence, not automatically a model-quality failure.
docent-full. The adaptive creator pursues the same museum-docent requirements across every arm while responding to what each Charlie actually said. Every run preflights the same live readable/presentable PDF and disables the generic per-turn acknowledgement; explicit long-running draft narration remains. Standings score = 45% autonomy + 35% final build completeness + 10% discovery before drafting + 5% reliability + 5% continuous first-audio latency. Autonomy itself is 70% self-handoff coverage and 30% guard-free turns. Binary checks are diagnostic gates, not additive weights. Provider/runtime failures are excluded instead of receiving a zero. Timeline and transcript inspectors are rendered by the same core/timing_viz.py component used by the Timeline family.