Papers
arxiv:2610.09683

System Switch: When Should a Fast Decision Model Stop and Think?

Published on Oct 7
· Submitted by
Gian Luca Bailo
on Oct 8
Authors:

Abstract

Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running. We use closed-loop Doom and the new open "System One" typed-decision models, served through a common llama.cpp interface. On 900 held-out questions, (i) zero-shot decision models from 0.15B to 9B parameters choose to collect items 1.6-1.8 times more often than chance among their errors, in any option order, although the order changes some models' accuracy; (ii) accuracy, calibration and sensitivity (how well confidence separates right from wrong answers) are distinct: models of similar accuracy differ widely in AUROC, and the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong; (iii) offline, deferring the least confident 30% of decisions to a reasoning model gains over random deferral in proportion to the actor's AUROC (rank correlation 0.87); with actor and rate chosen on held-out games the gain is +0.13 [0.08, 0.18] with doomLaya's option order and +0.08 [0.02, 0.14] with shuffled options, and reasoning carries about half of it; (iv) in closed loop (33 games, three seeds) no variant reaches the exit. Committing to plans, the reasoner's or a fixed explore rule's, opens more doors and makes an actor that stands still play; with the rule the agent dies more often. Told that some doors need keys, the reasoner takes ordinary doors for locked ones, which the state cannot tell apart; without that knowledge it goes back to collecting. We release code, prompts, data and logs.

Community

Paper author Paper submitter

A small typed-decision ("System One") model plays Doom in real time through llama.cpp's /v1/systemone endpoint (~29 ms per decision). A reasoning VLM (Qwen3.6-35B-A3B) is called only when a gate opens on uncertainty or lack of progress, while the game keeps running.

  • Zero-shot decision models of similar accuracy differ widely in AUROC; the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong.
  • Offline, deferring the least confident 30% of decisions to the reasoner beats random deferral by +0.13 on held-out games, in proportion to the actor's AUROC.
  • In closed loop (33 games), a fixed explore rule behind the same gate went as far as the reasoner. The logged thoughts show why: told that some doors need keys, the reasoner takes ordinary doors for locked ones.

Code, prompts, per-question records and game logs are in the system-switch branch.

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.09683 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.09683 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.09683 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.