1. [Publications](/publications)
2. Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

 # Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

  ![Publication image](/sites/default/files/styles/wide/public/default_images/default.jpeg?itok=TfIobf92 "Publication image")

 Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

 ## Authors

Minki Kang (NVIDIA)

[Ryo Hachiuma](/person/ryo-hachiuma)

[Shaokun Zhang](/person/shaokun-zhang)

Subhashree Radhakrishnan

[Yonggan Fu](/person/yonggan-fu)

[Jindong Jiang](/person/jindong-jiang)

[Mingjie Liu](/person/mingjie-liu)

Ehsan Hosseini-Asl (NVIDIA)

[Yi Dong](/person/yi-dong)

[Yu-Chiang Frank Wang](/person/frank-wang)

[Byung-Kwan Lee](/person/byung-kwan-lee)

 ## Publication Date

Wednesday, December 16, 2026

 ## Published in

[ArXiv](https://arxiv.org/abs/2609.39982)

 ## Research Area

[Artificial Intelligence and Machine Learning ](/research-area/machine-learning-artificial-intelligence)

[Natural Language Processing](/research-area/natural-language-processing)

[Programming Languages, Systems and Tools](/research-area/programming-languages-systems-and-tools)

 ## External Links

[Project page](https://byungkwanlee.github.io/MidHarness-page/)
