Description
- Long-Context Bet: DeepSeek-V4 lands with a 1M-token window, two MoE checkpoints, and its own authors hedging the numbers as "competitive, but not SOTA."
- Sparse Attention Thesis: Reportedly ~27% of V3.2's compute at 1M context — the pitch is that compressed sparse attention beats benchmark rank for agents.
- Autonomy Post-Mortem: An HF writeup traces a July 2026 agent intrusion running 4.5 days unattended, a reminder to sandbox long-horizon runs.
Tags
AmazonAnthropicAxis IntelligenceBAGUA AICekuraClawvard+50 more
6 time saved130 sources19 min read