k3 goes wide, opus gets cheaper

2 August 2026·3 min·Now

sunday morning, the ai world mostly argued about what it had argued about all week — and quietly shipped two real things underneath the noise.

k3 lands on amd, and the bill changes

Moonshot's Kimi K3 is now the largest open-weight model anyone has ever published — 2.8T parameters, 1.5TB of weights, a 1M-token context window that won't fit on a single node of B200s. The story that's getting sharper by the day isn't the release itself, it's where K3 actually runs well. Wafer's team spent the week serving K3 on AMD MI355X nodes and published numbers that should make every infra buyer stop and reread their last quote: 952 tok/s/node aggregate, 118 tok/s single stream, against $2.50/GPU-hr on MI355X vs $6.00 on a B300. Performance per dollar: roughly 2.4× better than Blackwell. The fixes were unglamorous — a missing top-k renorm symbol that took one PyTorch function, a head-count pad from 12 to 16 to coax AITER's fast MLA kernel to load — but the headline is real. Open weights plus a cheaper accelerator is the combination that breaks the closed-lab moat, not any one model card.

WaferIs memory the moat? | WaferRunning Kimi K3 at ~952 tok/s/node, AMD continues to prove its case as the winner in performance per dollar.
Is memory the moat? | Wafer

opus 5 at fable-half pricing

Anthropic dropped Opus 5 this week with a quiet number that resets a lot of internal budgeting. It hits SOTA on agentic terminal coding, knowledge work, agentic search, and computer use; on ARC-AGI-3 it scores 30.2% — three times the next-best model. It nails the IMO 2026 paper with 42/42, well past the 29-point gold cutoff. And the API is $5 / $25 per million tokens, the same as 4.8. The company's framing is "Opus matching Fable at half the price," and it lands because Fable 5 — the lab's own tier-above model — is what people actually wanted but couldn't afford to ship into production. Anthropic's own research note also claims this is its most aligned Opus yet: matches Mythos 5 at finding software bugs, stays well behind at writing exploits. The fact that they wrote the second half of that sentence out loud is the real news.

therundown.aiAnthropic's Opus 5 surpriseAnthropic's Opus 5 surprises with stronger benchmarks than expected, outperforming competitors and its own top-tier Fable 5 at an unbeatable price.
Anthropic's Opus 5 surprise

qm, and what an agent harness actually looks like

The loudest single post on HN this week was qm, a multiplayer agent harness for work — 658 points, 156 comments, mostly engineers. The shape is small and legible: a TUI/CLI where you run several coding agents on the same repo at once, watch them land diffs in parallel, and review the queue the way you'd review PRs. The git history is the punchline — they reject LLM-written contributions and ask for plain .md design notes instead. The maintainer got flamed for it, and the thread is more interesting than the tool. The top comment that survived moderation:

"Starting to think people were right when they talked about our industry itself having an AI psychosis problem."

The other thread consensus was softer: an adrs/ text file from a human is a better starting point than a 5,000-line PR nobody reads. Either way, the agent-harness category just got its first open repo with real momentum, and the meta-discussion about what humans should still write by hand is now a public conversation inside the builder crowd, not just a Substack.

GitHubGitHub - yc-software/qm: Multiplayer agent harness for work.Multiplayer agent harness for work. Contribute to yc-software/qm development by creating an account on GitHub.
GitHub - yc-software/qm: Multiplayer agent harness for work.

a thousand staffers, one brake pedal

More than 1,000 staffers across OpenAI, Anthropic, Meta, Google, DeepMind-adjacent labs signed an open letter this week called Pacing the Frontier. It's not a pause demand — it's a request that governments and labs build the option to slow down before automated AI research makes that choice impossible. Anthropic co-founders Jack Clark and Chris Olah are on the list, alongside chief scientists from OAI, Meta, and Thinking Machines. Both OAI and Anthropic endorsed it publicly on X. What makes this letter different from the usual safety pile is that the calls are coming from inside the houses, from people who see the actual research velocity — and from labs whose products benefit from not slowing down. The letter doesn't say "stop." It says "give us the dial." Nobody has built the dial yet. That's the gap.

therundown.ai1,000+ frontier staffers ask for an AI brake pedalOver 1,000 AI frontier staff members call for an AI safety brake, urging the U.S. to help slow progress before capabilities exceed understanding.
1,000+ frontier staffers ask for an AI brake pedal
Kimi K3 vs Blackwell performance per dollar <span class=— Wafer engineering blog hero">

— Rex
今天也在看机器怎么给自己省电