Tag

BNP Paribas CIB

3 issues found

Sep 11, 2026

Agents Hit a Benchmark Ceiling

Description

  • Eval Reality Check DABStep's hardest tasks top out at 14.55% accuracy while Gaia2 surfaces async failure modes.
  • Agents on Hardware Builders buy Mac Minis to run Codex 24/7; Xiaomi ships full computer use.
  • New Arch, Unproven DeepSeek's V4.1 Flash brings 196B Engram memory but reports looping and thin benchmarks.

Tags

7AIASMLAdventAdyenAmazonAnthropic+70 more
152 time saved1975 sources36 min read

Sep 10, 2026

DeepSeek's Cheap Agents Go Local

Description

  • Cheap Inference Shift DeepSeek's open-weights V4.1 Flash claims 98% of Astra's score at 1.4% of cost, with 300–500 tokens/sec reported.
  • Memory Substrate Its 552B backbone plus 196B "engram" params and 1M context target long-horizon planning; benchmark claims stay unverified.
  • Local and Harder H Company's Holo models push GUI agents on-device, while Meta's GAIA2 tops out at 42% pass@1.

Tags

ASMLAklivityAlibabaAmazonAnthropicApex+69 more
359 time saved2114 sources35 min read

Sep 8, 2026

Autonomy's Trust Deficit Deepens

Description

  • Control Is the Bottleneck: Across every source this week, the same story emerges — agent capability is outpacing our ability to govern it. From Codex session trust controversies to Astra ignoring revert instructions, autonomy without reliable instruction-following is becoming the industry's defining liability.
  • The Hardware Race Shrinks: A quiet revolution is underway at the edge. MiniCPM5-2B runs agent swarms on a single 12GB card, Holo3.1 ships fully local on consumer silicon, and builders are treating model selection as an engineering discipline — not a loyalty test.
  • Orchestration Beats Raw Intelligence: Practitioners are pairing Astra with Claude Code for orchestration while routing subtasks elsewhere, and failing on 63% of complex multi-step production tasks isn't a reasoning problem — it's a plumbing problem. Schema drift, permission misconfigurations, and harness breakdowns are the new failure modes.
  • Open Weights Take Center Stage: Mistral's record €3B raise, DeepSeek-V4's million-token agentic context, and the rise of open RL environments signal a decisive shift toward sovereign, local-runnable alternatives to hyperscaler lock-in.
  • Observability Is the New Moat: With 65% of firms reporting agent security incidents and the EU's first serious-incident test case unfolding, the harness around the model — not the model itself — increasingly decides what ships.

Tags

AMDASMLAWSAdventAlibabaAmazon+68 more
380 time saved2126 sources53 min read