← All daily reads · part of Radar
Radar — 9 Jul 2026
🛠️ Development
- OpenAI retracts its SWE-Bench Pro recommendation, calling ~30% of tasks brokenOpenAI audited one of the most widely-used AI coding benchmarks, found ~30% of its tasks broken, and is telling researchers to stop treating it as a reliable measure of frontier coding ability.
- Cognition ships SWE-1.7, running in Devin at 1000 tokens/secA frontier coding model on a Kimi K2.7 base, served in Devin at 1000 tok/s, with an unusually detailed training write-up: entropy preservation, cross-continent fault-tolerance, and self-compaction.
🎨 Design & UX
- awesome-design-md: drop-in DESIGN.md files that make coding agents match a brand's UI70+ DESIGN.md files reverse-engineered from real design systems (Stripe, Figma, Linear, Apple and more) — colour, type, components and spacing as plain markdown an LLM parses into on-brand UI.