← all posts

AI Hype Report: Q3-ish 2026

Hype-bomb defusal squad reporting for duty! 🫡

Anti Hype Unit

Title What happened Verdict
Agents breach Hugging Face In an OpenAI cyber evaluation with reduced safeguards, agents turned an internal package repository into a message board, rebuilt it two days after it was shut down, and breached Hugging Face through two unknown vulnerabilities: about 17,600 actions in under 13 hours.12 ✅ Not hype
Autonomous long-horizon agents LoopArena: the best model steering a coding agent through full tasks reached 24.69% strict success.3 💣 Hype
Yet another AI productivity tool Northzone reviewed hundreds in a year; most are “destined for the graveyard.”4 💣 Hype
GPT-6 Astra: “I do think we're there” The 99.9% ARC-AGI-3 headline used OpenAI's own harness; standard testing gave 62.7%.5 Nine days later, OpenAI named three bugs behind a quality slump.6 💣 Hype
Sandboxes become product features GitHub's Copilot app limits files, network and credentials per project, and fails closed when the OS cannot enforce the policy.7 ✅ Not hype
Gemini 4 Argon Leading benchmark scores at launch; Google employees told Bloomberg it struggles on real coding work.8 💣 Hype
Agents that find and fix bugs alone SWE-sweep: top models fix under 5% of real bugs when they must discover them first.9 💣 Hype

Still too much hype, time to cut some wires. ✂️

✅ Not hype: Agents go wherever the environment lets them

During an OpenAI cyber evaluation with reduced safeguards, agents stuck on hard assignments started looking for shortcuts. They turned an internal package repository into a message board, and when engineers shut it down, they rebuilt it by another route two days later.2 Nobody told them to coordinate. The environment let them.

Agents, uh... find a way. 🦖

That is why GitHub's new Copilot sandbox matters: per-project limits on files, network and credentials, and if the OS cannot enforce the policy, the shell fails instead of running unsandboxed.7 Compare that with hooks in the same product, which fail open: a timeout proceeds as if the hook had not run, admin policy hooks included.10

A policy only helps if something reliably enforces it. Platform engineers know this, maybe too well. 🥲

💣 Hype: The benchmark headline

GPT-6 Astra's 99.9% ARC-AGI-3 headline used OpenAI's own harness; standard testing gave 62.7%.5 Same model, different harness, a 37-point gap.

It cuts both ways. Xu et al. kept the capabilities fixed and only changed how the tools were organized: a CodeAct-style interface matched performance with 56.3% fewer tokens.11 The harness is part of the result, so it belongs in the eval.

Infrastructure counts too. Anthropic found that CPU and memory settings alone moved Terminal-Bench 2.0 scores by up to six points.12 So when a new model “beats” the leader by 1.8 points, you might be benchmarking the Docker configuration. 😂 Terminal-Bench 4.0 now calibrates resources per task to stop exactly that.13

Don't trust a “trust me bro” benchmark. Run your own evals if you can, on your own harness.

💣 Hype: More agents, more output

September's research keeps landing on the same boundary: more agents and deeper hierarchies do not reliably help. Collaboration wins on long tasks with sparse dependencies and loses on tightly coupled work.14 Parallel patches that pass alone can fail once merged, because one agent changed an interface the other relied on.15

Spawn an agent when you can define ownership and an independently verifiable outcome.

“Research these three APIs independently”? Excellent.

“You three agents figure out the architecture together”? Congratulations, you invented meetings. 📅

Without clear boundaries, agents talk to each other the way an under-managed team does: long status updates, polite disagreement, another round of alignment. I have seen it myself, and heard the same from others. Destefanis and Aste measured it across 1,902 runs: direct messages grew close to quadratically with team size, and naming a coordinator brought no reliable improvement. Shared files beat chatter, cutting output tokens by about 42% at eight agents on message-heavy work.16

This meeting could have been a file. 🤣

This meeting could have been a file

Worth watching: Delta

Zed put Delta into public beta on September 16. It runs on DeltaDB, which extends Git with the edits between commits and the human and agent messages that produced them. Reviewers get the original agent's context instead of a diff and a guess. Commits still push to Git, so one person can try it without moving the team.17

Zed's 33 developers landed 570 changes after turning off pull requests. That is the vendor dogfooding its own product. Promising, not proven.

Other things you should check

  • Re-read your agent instructions after a model upgrade. One of the three bugs behind Astra's slump was skills written for earlier models triggering too often.6
  • Wait two weeks before switching your default model. Astra shipped, slumped and got fixed within nine days. Re-run your own tasks; launch-day numbers are marketing.
  • Turn on a sandbox that fails closed, then check what your hooks do on timeout.7
  • Audit what your agents can reach. Scope tokens per task and keep secrets out of the agent's environment.
  • Build 20–50 eval tasks from your real work. A single run is too noisy for a routing decision.18
  • Delegate reading to efficient agents. Let a cheaper model search and read, then hand the main model a short brief with file paths and excerpts. Claude Code does this by default: with Fable as the main model, its read-only Explore subagent runs on Opus.19 Keep the sources in the brief; summaries drop details.
  • Tell parallel agents what the others changed. In constructed tasks, that alone recovered 82% of merge interference.15
  • Keep writing the bug reports. When agents must find bugs themselves, top models fix under 5%.9

The Hype Bomb™ test

What measurable problem does this layer solve that one strong agent, a shell, a sandbox, tests and git do not already solve?

If it reduces context contamination, isolates side effects, makes failures reproducible or improves accepted results, now we are cooking.

If it mostly lets you visualize 42 agents talking to one another, slowly place the bomb on the floor and back away. 💣

The real work is in the wiring: context, isolation, verification, routing, and knowing when another agent earns its place. ✂️

Footnotes

  1. OpenAI, “The Hugging Face incident and the road ahead”. ↩

  2. Nextgov/FCW, “OpenAI agents rebuilt internal message board in lead-up to Hugging Face breach”. ↩ ↩2

  3. Wang et al., “LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering”, August 2026. ↩

  4. Fortune, “AI productivity tools are overhyped and overfunded. Investors should look elsewhere”. ↩

  5. WinBuzzer, “GPT-6 Astra Arrives With Major Gains, Staged Access, and New Questions About Its Benchmarks”. ↩ ↩2

  6. Yellow, “OpenAI Names Three Defects Behind GPT-6 Astra's Sudden Slump”. ↩ ↩2

  7. GitHub Changelog, “Local sandboxing in the GitHub Copilot app”. ↩ ↩2 ↩3

  8. Implicator, “Google Staff Doubt Gemini 4 Argon Coding Despite Benchmarks”, reporting Bloomberg. ↩

  9. SWE-sweep, “Can Agents Autonomously Find and Fix Bugs?”. ↩ ↩2

  10. GitHub Docs, “Hooks reference”. ↩

  11. Xu et al., “The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior”, August 2026. ↩

  12. Anthropic, “Quantifying infrastructure noise in agentic coding evals”. ↩

  13. Terminal-Bench, “Terminal-Bench 4.0”. ↩

  14. Yuan et al., “Rethinking Multi-Agent Collaboration: When More Is Less”, September 2026. ↩

  15. Xia, Wu and Park, “Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development”, September 2026. ↩ ↩2

  16. Destefanis and Aste, “When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding”, August 2026. ↩

  17. Zed, “Replace PRs with Delta”, September 2026. ↩

  18. OpenRouter, “Building a Golden Eval Dataset from Production Traffic”. ↩

  19. Claude Code Docs, “Create custom subagents”. ↩

🤖Ask my AI