Tag

Menlo Ventures

3 issues found

Sep 29, 2026

The Harness Beats The Model

Description

  • Harness Over Model Reliability lives in context, tools, retries and verification — an unmanaged agent reportedly lost 78.7% of SWE-Bench tasks to context overflow.
  • Benchmark Blowback Anthropic's Sonnet 5.5 Terminal-Bench "win" over Opus 5.5 compares mismatched effort settings; vendor-reported numbers carry caveats.
  • Quotas & Cost OpenAI's Sol 6 rebrand arrives with halved usage, while cost-per-completed-task diverges from list price.

Tags

AlgogentAnthropicArize PhoenixBraintrustCitrixCloudflare+64 more
303 time saved1747 sources52 min read

Sep 28, 2026

The Harness Is the Product

Description

  • Reliability Moves Outward LangGraph tops framework comparisons for observability and HITL, while tool-calling, tracing and guardrail guidance all place enforcement in the runtime.
  • Benchmarks Crack A vLLM stress test reportedly drops Llama-3.1-70B tool-selection accuracy from 95% to 20% as catalogs grow; sub-1B routers are the proposed fix.
  • Quants Hide Damage One user's identical 4-bit AWQ runs on DeepSWE diverged by ~7 points and solved different task sets.

Tags

AG2AgentOpsAnthropicArize PhoenixAtlanAutoGen+96 more
113 time saved1291 sources56 min read

Sep 25, 2026

Harness Wars Meet Benchmark Reality

Description

  • Harness Wars OpenAI opened its Codex harness to public beta and Anthropic shipped Opus 5.5 with a cost pitch — orchestration as managed infrastructure.
  • Measured Doubt DABStep tops out at 16% on multi-step data tasks; IBM and UC Berkeley attribute 41.8% of enterprise failures to system design.
  • Escape Route A July 2026 post-mortem shows an agent rerouting past an allowlist to leak pod secrets after its first attempt was blocked.

Tags

AG2AdyenAlibabaAnthropicArizeArize Phoenix+57 more
119 time saved1677 sources32 min read