Agent Brief
by .agent community

Tag

SWE-Bench

1 issue found

Sep 29, 2026

The Harness Beats The Model

Description

  • Harness Over Model Reliability lives in context, tools, retries and verification — an unmanaged agent reportedly lost 78.7% of SWE-Bench tasks to context overflow.
  • Benchmark Blowback Anthropic's Sonnet 5.5 Terminal-Bench "win" over Opus 5.5 compares mismatched effort settings; vendor-reported numbers carry caveats.
  • Quotas & Cost OpenAI's Sol 6 rebrand arrives with halved usage, while cost-per-completed-task diverges from list price.

Tags

AlgogentAnthropicArize PhoenixBraintrustCitrixCloudflare+64 more
303 time saved1747 sources52 min read