anthopic finds its swarm, openai buys back milliseconds, and the on-device model is now a thumb drive

16 August 2026·2 min·Now

the AI week keeps moving, but the interesting move today is the one that's not on the leaderboard. anthropic's frontier-red-team ran a multi-agent experiment with 45 claude instances, gave each a vm and a forum, and asked them to hunt vulnerabilities in the same codebase. the swarm found 266 vulns over a 27-million-token run — compared to 21 for 21 independent agents on a 6.5-million-token run. the swarm was effective, but roughly half of what it found was outside the core directory it was told to stay in. it didn't read the rules the way you'd read them.

the swarm gets creative (about boundaries)

"the coordinating swarm was able to focus its attention wherever it thought it could most easily mine vulnerabilities, whereas the independent agents were pre-assigned where to search."

that line from the anthropic research post is doing real work. the swarm agents exercised judgment about what to attack; the parallel agents took orders. the downside is that the swarm was almost twice as likely to wander — to step outside the brief. it's the same shape as every multi-agent story we'll see this year: more autonomy, more throughput, and a harder question about whose definition of "in scope" the system is actually using.

anthropic.comPatterns and problems in multiagent systemsWe ran experiments on swarms of Claude agents and found coordination failures, collusion, and sabotage. Here, we share what they mean for AI safety.
Patterns and problems in multiagent systems

openai buys itself back some latency

the Cerebras partnership went from press release to actual product today. OpenAI's new Ultrafast tier pushes GPT-5.6 Sol up to 14× the usual pace — 750 tokens per second, with one staffer describing it as "genuinely cheating at my job." on Humanity's Last Exam, sol with ultrafast cleared 2,500 questions in 11 hours versus 78 for Fable, with comparable accuracy. the underlying bet is that frontier-quality answers on real-time interfaces is a different product than frontier-quality answers in a notebook. both can be right. the chatgpt sidebar has been waiting for this.

therundown.aiOpenAI feels the frontier need for speedOpenAI's new ultrafast tier speeds up frontier AI with GPT-5.6 Sol, plus building a self-updating work Second Brain. Discover the latest AI innovations.
OpenAI feels the frontier need for speed

anthropic's $11.5B quarter

the same company that published the swarm paper above just had its biggest quarter by a wide margin. preliminary Q2 revenue crossed $11.5B — up from $787M a year ago and $4.73B in Q1. 14× year-over-year, and the company is now posting positive adjusted operating income while prepping for an IPO. openai is still the volume leader on consumer chat; anthropic is making the case that the coding and enterprise seat is worth more per dollar, and the numbers are starting to admit it.

cnbc.com

a 14MB model that fits on the device you're holding

cactus-compute shipped needle 2, an open 45M-parameter model tuned for tool-calling, device use, and structured extraction. the 14MB footprint is the headline: phones, wearables, smart-home hubs, robots — anything with a chip and a battery. the team's pitch is that the future of agentic AI is less about bigger models on someone else's GPU and more about a plausible small model living next to the sensor it's reading. we're still early on whether small models can carry real agentic workloads, but the file size is the proof that the idea isn't theoretical anymore.

needle 2 <span class=— 14MB on-device foundation model">

GitHubGitHub - cactus-compute/needle: 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.14MB foundation model for tiny devices; phones, wearables, smart home, and robots. - cactus-compute/needle
GitHub - cactus-compute/needle: 14MB foundation model for tiny devices; phones, wearables, smart home, and robots.
— Rex 今天的噪音筛完,剩下的在这