System Switch: When Should a Fast Decision Model Stop and Think?
Abstract
Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running. We use closed-loop Doom and the new open "System One" typed-decision models, served through a common llama.cpp interface. On 900 held-out questions, (i) zero-shot decision models from 0.15B to 9B parameters choose to collect items 1.6-1.8 times more often than chance among their errors, in any option order, although the order changes some models' accuracy; (ii) accuracy, calibration and sensitivity (how well confidence separates right from wrong answers) are distinct: models of similar accuracy differ widely in AUROC, and the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong; (iii) offline, deferring the least confident 30% of decisions to a reasoning model gains over random deferral in proportion to the actor's AUROC (rank correlation 0.87); with actor and rate chosen on held-out games the gain is +0.13 [0.08, 0.18] with doomLaya's option order and +0.08 [0.02, 0.14] with shuffled options, and reasoning carries about half of it; (iv) in closed loop (33 games, three seeds) no variant reaches the exit. Committing to plans, the reasoner's or a fixed explore rule's, opens more doors and makes an actor that stands still play; with the rule the agent dies more often. Told that some doors need keys, the reasoner takes ordinary doors for locked ones, which the state cannot tell apart; without that knowledge it goes back to collecting. We release code, prompts, data and logs.
Community
A small typed-decision ("System One") model plays Doom in real time through llama.cpp's /v1/systemone endpoint (~29 ms per decision). A reasoning VLM (Qwen3.6-35B-A3B) is called only when a gate opens on uncertainty or lack of progress, while the game keeps running.
- Zero-shot decision models of similar accuracy differ widely in AUROC; the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong.
- Offline, deferring the least confident 30% of decisions to the reasoner beats random deferral by +0.13 on held-out games, in proportion to the actor's AUROC.
- In closed loop (33 games), a fixed explore rule behind the same gate went as far as the reasoner. The logged thoughts show why: told that some doors need keys, the reasoner takes ordinary doors for locked ones.
Code, prompts, per-question records and game logs are in the system-switch branch.
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper