SciGroveBeta
Machine Learning

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang

Featured August 22, 2026

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Instead of making AI models smarter, this paper shows that improving the control system around them, like giving them a clear checklist and rules for each step, makes them much more reliable and cheaper to run for complex tasks.

In depth
The paper introduces StateM, an agent-native runtime that significantly boosts the reliability of long-horizon agents without altering their core model weights. It achieves this by externalizing procedural knowledge into a shared, machine-operable YAML runbook, which defines durable states, explicit transition contracts, and recoverable execution paths. This approach, termed harness scaling, allows agents and users to jointly inspect, audit, and revise the control layer, effectively closing critical execution gaps.

Key Takeaways

  • 1
    Harness scaling is presented as a crucial capability axis, complementary to model scaling, that dramatically improves agent reliability and performance by enhancing the execution system around a fixed model.
  • 2
    StateM provides an agent-native control layer through a YAML runbook, offering durable states, explicit transition contracts, and a shared, auditable interface for both agents and users.
  • 3
    The methodology enables failure-driven harness optimization, allowing procedural knowledge to accumulate in the external control layer, leading to substantial accuracy gains and significant cost reductions for agent deployment.

Conceptual Flow

HIGH LEVEL
1
Methodology (The 'Logic')

The system gives an AI agent a clear, step-by-step plan with rules to follow, making sure it finishes complex jobs correctly.

AI Agent
Complex Task
Follow Rules
Finished Job
2
Results (The 'Impact')

This new way helps AI agents achieve much higher accuracy on tough tasks and costs significantly less money to operate.

Old Way
New Way
Compare Performance
Higher Accuracy
Lower Cost

This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.