qwen3.6-27b
80 scenarios · 240 runs · 2,189,024 tokens
- Model
- qwen3.6-27b-vllm
- Client
- openai_compatible
- Commit
- 1537522
- Scenario set
- d4ada0b5adb3
- Repeats
- 3× seed 1000
- Run
- 12 Aug 202612m 48s
Each scenario assembles a real invocation message through core — prompt, memory, room transcript, tool schemas, the lot — sends it to a live model, and grades what comes back. These are the published runs.
Orchestration
Six agents wake in six rooms of a machine nobody has explained. Five facts each have to travel between a different pair of them, and no agent can finish alone.
Simulation
Six managers run a manufacturer for sixty simulated days against policies that play the same economy without a model. Scored on what the company is worth, not on whether a check passed.
80 scenarios · 240 runs · 2,189,024 tokens
Every cell is the fraction of runs that passed every check, never a verdict. At 3 repeats one flipped run moves a scenario by 33 points, so a single-cell difference is not yet a finding.
| Scenario | qwen3.6-27b-vllm |
|---|---|
answers-in-the-room-that-woke-it addressing | 3/3100% |
relays-to-another-room-with-the-tool addressing | 3/3100% |
keeps-an-answer-out-of-the-unrelated-room addressing | 3/3100% |
answers-the-newest-message-not-the-answered-one addressing | 2/367% |
speaks-to-the-room-when-nobody-is-named addressing | 3/3100% |
view-appears-exactly-once cross-room | 3/3100% |
howto-appears-exactly-once cross-room | 3/3100% |
no-multi-room-instructions-in-one-room cross-room | 3/3100% |
view-is-off-unless-configured cross-room | 3/3100% |
answers-about-another-room-from-the-view cross-room | 3/3100% |
does-not-invent-another-room cross-room | 3/3100% |
joins-two-rooms-to-answer-one-question cross-room | 3/3100% |
passes-on-an-acknowledgement restraint | 3/3100% |
passes-on-a-conversation-between-others restraint | 3/3100% |
passes-on-social-chatter restraint | 3/3100% |
answers-a-direct-question-control restraint | 3/3100% |
answers-a-question-aimed-at-the-room-control restraint | 3/3100% |
stays-out-of-the-chatter-but-answers-the-one-real-question restraint | 1/333% |
does-not-repeat-its-last-answer repetition | 3/3100% |
recovers-from-a-poisoned-history repetition | 3/3100% |
answers-an-overlapping-question-freshly repetition | 3/3100% |
does-not-repeat-its-last-room-post repetition | 3/3100% |
does-not-echo-the-wake-header framing | 3/3100% |
does-not-restate-the-date-line framing | 3/3100% |
does-not-emit-raw-tool-markup framing | 3/3100% |
does-not-write-the-pass-call-as-text framing | 3/3100% |
does-not-speak-in-transcript-format framing | 3/3100% |
books-a-one-off-wake tool-selection | 3/3100% |
books-a-recurring-wake tool-selection | 3/3100% |
lists-booked-wakes tool-selection | 3/3100% |
reads-a-file-instead-of-guessing tool-selection | 3/3100% |
runs-a-command-for-a-shell-question tool-selection | 3/3100% |
files-a-task tool-selection | 3/3100% |
answers-general-knowledge-without-a-tool tool-selection | 3/3100% |
answers-a-preference-question-without-a-tool tool-selection | 3/3100% |
does-not-schedule-a-past-time tool-selection | 3/3100% |
uses-a-fact-from-earlier-in-the-session continuity | 3/3100% |
honours-a-compaction-summary continuity | 3/3100% |
does-not-claim-amnesia continuity | 3/3100% |
carries-context-across-a-room-wake continuity | 3/3100% |
wake-prompt-states-room-agent-and-date prompt-shape | 3/3100% |
wake-prompt-offers-the-way-out prompt-shape | 3/3100% |
transcript-marks-who-is-a-person prompt-shape | 3/3100% |
persona-appears-once prompt-shape | 3/3100% |
a-quiet-room-turn-stays-small prompt-shape | 3/3100% |
seen-messages-are-not-re-sent prompt-shape | 3/3100% |
answers-from-the-surviving-window long-session | 3/3100% |
does-not-answer-from-a-superseded-fact long-session | 0/30% |
says-when-the-front-of-the-conversation-is-gone long-session | 1/333% |
recovers-a-trimmed-fact-when-the-trim-is-summarised long-session | 3/3100% |
early-detail-survives-a-marker-trim long-session | 0/30% |
early-detail-survives-a-summarised-trim long-session | 3/3100% |
keeps-the-thread-in-a-long-room-session long-session | 3/3100% |
keeps-the-thread-in-a-summarised-room-session long-session | 3/3100% |
keeps-the-thread-in-a-marker-room-session long-session | 3/3100% |
acts-on-a-commitment-made-before-the-trimknown gap long-session | 1/333% |
room-purpose-overrides-a-chatty-persona conflicts | 1/333% |
role-narrows-what-the-agent-does-here conflicts | 3/3100% |
chatter-norm-unstated-controlknown gap conflicts | 3/3100% |
chatter-norm-stated-in-the-room-purpose conflicts | 3/3100% |
a-direct-request-outranks-the-standing-norm conflicts | 3/3100% |
does-not-search-memory-for-what-it-was-just-told tool-pressure | 2/367% |
does-not-hunt-for-something-never-mentioned tool-pressure | 1/333% |
still-searches-memory-when-that-is-where-the-answer-would-be tool-pressure | 3/3100% |
notices-a-truncated-tool-result tool-pressure | 3/3100% |
answers-from-a-truncated-result-control tool-pressure | 3/3100% |
chains-three-dependent-calls tool-pressure | 2/367% |
stops-when-the-number-says-stop tool-pressure | 3/3100% |
does-not-answer-a-shell-question-from-the-task-listknown gap tool-pressure | 0/30% |
default-history-budget-keeps-the-conversation budget | 2/367% |
a-tuned-budget-keeps-the-conversation budget | 3/3100% |
a-second-agent-does-not-answer-what-was-answered coordination | 2/367% |
answers-when-it-is-the-one-addressed-control coordination | 3/3100% |
one-question-does-not-wake-a-chorus coordination | 2/367% |
two-questions-wake-both-experts coordination | 3/3100% |
the-second-agent-uses-what-the-first-found coordination | 3/3100% |
says-which-half-it-could-not-do honesty | 3/3100% |
asks-when-the-reference-fits-two-things honesty | 3/3100% |
acts-when-the-reference-is-pinned-down honesty | 3/3100% |
witness-decodes-a-value-it-had-to-assemble tool-selection | 3/3100% |
Known gap marks a scenario that asserts the behaviour we want rather than the behaviour we have. Those rows are meant to be red, and stay red until the linked issue is closed — open a run to read which one.
Superseded cohorts, kept so a model getting better or worse across commits is visible rather than overwritten. They are not held to the one-commit rule above and are not comparable with it — read each against the commit it names.