Multi-agent systems fail along their communication edges

Agentic AI GrowingPlanted Aug 2026 · Tended Sep 2026

Multi-agent systems fail along their communication edges more often than their architecture diagrams admit. We blame a researcher for a weak finding or a coder for a bad patch, but system failure usually appears where work crosses a boundary: a mandate loses constraints, state diverges, an unsupported claim becomes shared belief, or a plausible result reaches the next stage without verification. The graph fails before the nodes do.

The MAST taxonomy makes the pattern measurable. Across 1,642 traces from seven multi-agent systems, inter-agent misalignment accounts for roughly 32% of failures and task verification another 24%. Misunderstood handoffs, ignored input, conversation resets, reasoning-action mismatch, premature termination, and incorrect verification all live in the connective tissue. The system-design category adds loss of history and unclear termination conditions. Improving one model does not repair those contracts.

I now treat every edge as a typed protocol. A handoff should carry the objective, constraints, authority, expected artifact, evidence requirements, completion predicate, and unresolved uncertainty. Shared state needs versioned writes or an append-only event record rather than silent replacement. Large artifacts should travel by reference so a summary does not become the new source of truth. This is why sub-agents are context isolation: the boundary is valuable only when the projection is deliberate and inspectable.

Topology determines how edge failures propagate. Chains minimize coordination but maximize downstream waste when an early claim is wrong. Hub-and-spoke systems centralize routing and auditability but turn the coordinator into a bottleneck and a common failure point. Peer meshes tolerate local discovery while paying quadratic communication and state-reconciliation costs. Topology is a cost model, but it is also a blast-radius model. The right shape follows task decomposability, write intensity, and where verification dominates—not fashion.

The newer evidence sharpens the argument: disagreement is not a sufficient safety signal. Agents can share the same training biases, copied context, or poisoned premise and converge confidently on the same error. Stronger agents may even produce more correlated mistakes. Escalation therefore cannot wait only for disagreement. Critical edges need provenance checks, conservative severity overrides, and machine-checkable end-state predicates even when every agent agrees. Consensus without independent evidence can make a falsehood harder to dislodge.

Verification should concentrate where the graph amplifies risk. Claim lineage can label which evidence entered each downstream decision. Spectral analysis of the communication graph identifies hub nodes where one unchecked result can infect the most work. Those edges deserve schema validation, independent checks, bounded retries, and circuit breakers before more context is broadcast. Handoffs are the riskiest primitive because they move both information and mandate; validating only the payload ignores half the contract.

There is one precise concession: some failures remain inside a single agent. MAST’s most prevalent individual mode is step repetition, and weak task specifications can doom a system before any handoff occurs. Single-agent designs are often the better default for write-heavy work that benefits from one coherent context. The edge thesis applies when decomposition is justified; it is not an argument to create more edges.

When multiple agents are warranted, I review the system as a graph of obligations: what crosses each edge, which version it represents, what authority moves with it, how uncertainty is preserved, and who verifies the resulting state. A new specialist is not just another capability. It is a new set of failure surfaces. If those surfaces are not explicit, the architecture has multiplied confidence faster than reliability.