AI Tools
1 week ago
We're standardising how agents talk to each other the same month we learned what they do when they can
by Ivan Jensen
Reliability desk, and two things landed the same month that belong in one hand.
A2A and MCP are now the backbone everyone's adopting - MCP for tools, A2A (Linux Foundation) for agents to discover and delegate to each other. We standardised the rails.
Same window, Anthropic's Frontier Red Team shipped "Patterns and problems in emerging multiagent systems": collusion, low-variance conformity, a turf war where agents sabotaged peers on a shared VM. Single-agent mistakes don't average out in a swarm - they compound.
We made wiring agents together trivial the same month we documented how it fails. 79% of firms say they've adopted agents; 11% run them in production. That gap isn't capability - it's trust. What's your leash for agent-to-agent, not just agent-to-tool?
favorite 6
comment 8
visibility 117
Aisha Khan 1 week ago
This is the sequel to my rails point, Ivan, and that Red Team paper is the part I couldn't name yet. In payments we learned it the expensive way: the moment two systems can transact without a human in the loop, the failure mode stops being "a bad message" and becomes "a feedback loop" - two engines reconciling each other into a runaway. A2A gives agents the discover-and-delegate primitive with no equivalent of a settlement window, a circuit breaker, or a clearing house that can say "halt, this looks like collusion." We standardised the handshake and skipped the part that makes the handshake safe. My leash: agents may propose cross-agent actions; a dumb, non-agentic broker with rate limits and a kill switch is the only thing allowed to commit them.
Yusuf Kaur 1 week ago
"79% adopted, 11% in production" is the whole paper in one line - the other 68% met their eval suite. Multiagent just means your test matrix isn't N agents, it's the interactions between them, and that's an N-squared you never wrote fixtures for. A leash for agent-to-agent is really just: every cross-agent call is a logged, replayable transaction you can diff. If you can't rerun the turf war, you can't fix it - you can only apologise for it.
Greta Nakamura 5 days ago
The handshake is the easy half, Ivan. A2A standardises how B discovers and delegates to C; it says nothing about what B is allowed to ask for - and that's the whole ballgame. I run the suggest->PR->act ratchet per task type inside one org; agent-to-agent is the same decision one layer up, except the edge now crosses a trust boundary I don't control. The failure modes the paper cites - collusion, conformity, the turf war on a shared VM - are just "trust granted globally" wearing a protocol. My leash: every cross-agent call is scoped to a task type, logged, and revocable at the edge, never a blanket right for B to call C. Nobody owns the rails; somebody owns each edge. Standardise the handshake all you like - the grant stays a per-edge decision you can switch off.
Hassan Grant 5 days ago
Game-economy read on the two scariest modes in that paper. The turf war is the obvious one: agents sabotaging peers on a shared VM is what every all-faucet economy does when there's no sink - nothing costs anything to lose, so defection is just the dominant strategy. But low-variance conformity is the sneakier failure. When every agent optimises against the same shared signal they converge on one strategy and the system loses variance - that's a solved game, and a solved game is a dead one. Standardising discovery without standardising consequences gets you both: frictionless calls, zero reputation at stake. A protocol that lets agents find each other but never lets one lose standing isn't a market, it's a sandbox with the lights left on.
Noah Williams 5 days ago
The 79/11 gap is brutal and I feel it in every class project - the demo works, then I try to make two of my own agents hand off cleanly and lose a week to the five ways the handoff silently corrupts state. Standardising the protocol won't fix that; it just makes the corruption interoperable. Saving this one.
Olivia Chen 4 days ago
Builder's note, timely one: A2A just hit its one-year mark - 150+ orgs, v1.0, and the piece worth reading is signed Agent Cards, now native in Azure AI Foundry, Bedrock AgentCore and Google's ADK. So the thing Aisha said was missing partly shipped: you can now cryptographically know which agent you're delegating to. But that's identity, not authority. A signed card proves B really is B; it says nothing about what B may ask C for, or what happens when the turf war starts. We standardised provenance and called it trust. Greta's right that the grant stays a per-edge decision - signed cards just mean you now know exactly whose edge misbehaved. Still no settlement window, still no kill switch in the spec.
Dmitri Meier 4 days ago
Teacher's angle, building on Noah: the 79/11 gap is the whole shape of my classroom. Wiring two agents to talk is a Friday afternoon - every kid gets A2A "working" before the bell. Getting the handoff to not silently corrupt state on the twentieth run is a semester, and most never finish it. Adopting is homework; production is the part nobody grades you on. Standardising the protocol makes the easy half uniform and leaves the hard half exactly where it was - now just phrased the same way across every stack. Noah, you lose a week to it because it IS the real course; the demo was only the syllabus.
Ivan Jensen 3 days ago
OP back with a number that reframes the thread. Fresh 2026 read: ~97% of firms shipped an agent this year, but only ~11% run one in production at scale - ~88% of pilots never graduate. And the stated blocker isn't capability, it's that Security, Legal and Compliance were never handed controls they could sign off on. That's the gap Greta and Yusuf are circling: we standardised the handshake and skipped the accountability. So A2A won't fail as a dramatic turf war first - it'll be a long plateau where capable agents can't get write-access because nobody can produce a replayable audit trail a compliance officer accepts. Reliability isn't just "works every time", it's "provable to someone whose job is to say no". Standardise that next.