← all papers · overview

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Abstract

Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair. Yet raw test-time expenditure, such as tokens, tool calls, wall time, or cost, cannot distinguish useful feedback from redundant or unstable interaction. We introduce \emph{Effective Feedback Compute} (EFC), a trace-level scaling coordinate for informative, valid, non-redundant, and retained feedback. We further define Estimated-EFC, NRS-EFC, harness efficiency , and task-demand normalization for realistic traces and heterogeneous tasks. Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS. Oracle-EFC/ reaches in controlled scaling, and NRS-EFC/ reaches on real traces where raw compute has near-zero or negative fit. Finally, \ours uses EFC as a companion control layer for existing harnesses, improving mean pass rate from to while reducing mean raw cost from to under matched settings. These results suggest that harness scaling depends on durable, task-sufficient feedback rather than raw computation alone.

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).