10 min read

Closing the Push Gap: A2A Completion Notifications Across LangGraph Agent Servers

Closing the Push Gap: A2A Completion Notifications Across LangGraph Agent Servers

This article follows Langstrata: A blueprint for deploying LangGraph with dual-agent servers, published 9 September 2026.

Readers who prefer to follow along in code will find hands-on tutorials in the Langstrata README. Four progressive labs run from a single development server to a self-hosted Docker Compose stack. Lab 3 demonstrates the completion notifier described here, using just a2a-demo.

The earlier article proposed a split deployment, in which a supervisor and its workers run on separate LangGraph Agent Servers, the server processes that host LangGraph graphs. Here they are called the supervisor's Agent Server and the worker's Agent Server. The earlier article also committed to building a completion notifier, so that a worker could tell its supervisor when its work was done, and proposed packaging that notifier as an MCP server.

The notifier now exists as the a2a_completion_notifier bite in langshark-bites, and Langstrata integrates it. It has two halves. The emitter is middleware on each worker that posts a notice when the worker finishes. The A2A receiver is a small FastAPI application, hosted on the supervisor's Agent Server, that accepts the notice and hands it to the supervisor.

This article covers four topics: how the A2A receiver coexists with the supervisor's built-in A2A server, the A2A implementation that every LangGraph Agent Server includes, what the emitter does, how a worker is configured to push, and what the supervisor's model sees and decides when a notice arrives. The README tutorials demonstrate the machinery itself.

1. Co-location with the built-in A2A server

Every LangGraph Agent Server includes a built-in A2A server, which serves the Agent Server's assistants over A2A. It handles the inbound half of the protocol: sending messages, streaming responses, and querying or cancelling tasks. In a split deployment, the supervisor acts as the A2A client, and the worker's Agent Server executes the task as a run. In LangGraph, a run is one execution of an assistant's graph on a thread, the thread holding the conversation's state. It proceeds from its input until it reaches a final status such as success or error.

A2A's answer to long-running tasks is push, and in a split deployment it would work in two steps:

  1. The supervisor hosts a webhook and registers its URL with the worker's Agent Server.
  2. When the task finishes, the worker's Agent Server posts a notice to that webhook.

Neither step is possible with today's Langgraph server:

  • Its built-in A2A server does not implement A2A's push methods, so the supervisor's attempt to register its webhook returns "method not found". And, nothing is sent when a worker's run ends.
  • Neither the supervisor Agent Server nor its built-in A2A server provides an endpoint where an incoming notice could land.

The supervisor is therefore left to poll the worker's Agent Server for status. The notifier supplies both missing halves: the emitter on the worker's side and the A2A receiver on the supervisor's.

The original plan filled that gap with a standalone MCP server. A simpler arrangement proved possible. A LangGraph Agent Server can host a custom web application in its own process and on its own port, so the A2A receiver is mounted directly on the supervisor's Agent Server. Its webhook routes sit beside those of the built-in A2A server without colliding: the built-in A2A server handles inbound A2A requests, and the A2A receiver handles completion notices.

The supervisor dispatches work with a push configuration attached. When the worker's run ends, its emitter posts the completion notice to the A2A receiver, which shares the supervisor's Agent Server with the built-in A2A server.

Co-location removes a separate process to deploy and secure. It also gives the A2A receiver direct access to the supervisor's Store, where notices wait until the supervisor reads them.

2. The emitter

The emitter is middleware on the worker graph. Its one job is to announce that the worker's run has ended, whatever the outcome.

Making it middleware rather than a tool is the essential choice. A tool runs only if the model decides to call it. Middleware runs unconditionally, once when the agent finishes and also when a model call fails, so a failed run is reported as reliably as a successful one. A worker cannot forget to report.

The notice itself is a standard A2A task object, sparsely populated: the task ID, its terminal status, a short summary of the result, and the routing token the worker was given. The notice can be signed so the A2A receiver can verify the sender. Delivery is retried, and the A2A receiver discards duplicates. A delivery failure is logged rather than raised, so a broken webhook never takes down the worker's own run.

3. Configuring a worker to push

A worker can push only if it knows where to send the notice and which conversation it belongs to. Because the worker's built-in A2A server offers no way to register a webhook, that information travels with the dispatch. When the supervisor starts a run on the worker's Agent Server over Agent Protocol, it attaches a push configuration: the A2A receiver's URL and an opaque token that only the supervisor can read.

On the worker's Agent Server, a graph factory builds the worker afresh for each run from that configuration. If a push configuration is present, the factory attaches the emitter; if not, the worker runs as an ordinary agent. A per-run factory is necessary because middleware cannot read a run's configuration on its own, so the decision has to be made when the graph is built.

Langstrata's supervisor is a Deep Agent, and Deep Agents launches workers as async subagents through its AsyncSubAgentMiddleware. That middleware gives the model five tools: start_async_task to launch a worker, and four more to check, update, cancel and list tasks. When start_async_task calls runs.create on the worker's Agent Server, it sends only a thread, an assistant and the task description. The supervisor run's configuration is not forwarded, so the push configuration never reaches the worker.

Langstrata closes the gap with ConfigForwardingAsyncSubAgentMiddleware, defined in its a2a_dispatch_forwarder.py module. The class is a drop-in replacement for Deep Agents' own AsyncSubAgentMiddleware. It is a subclass that behaves identically in every respect but one: its launch tool forwards the push configuration to the worker. The supervisor passes it to create_deep_agent alongside the same subagent definitions, and Deep Agents then uses it in place of the built-in middleware:

from langstrata.a2a.a2a_dispatch_forwarder import ConfigForwardingAsyncSubAgentMiddleware

supervisor = create_deep_agent(
    model=model,
    subagents=async_subagents,
    middleware=[ConfigForwardingAsyncSubAgentMiddleware(async_subagents=async_subagents)],
)

4. What the supervisor sees

The A2A receiver stores each notice against the supervisor's conversation. If the supervisor is idle, the A2A receiver also starts a new run so the notice is not left waiting. Middleware on the supervisor then injects the pending notices into the prompt history immediately before its next model call. Two kinds of message appear, both as ordinary user turns:

  • A wake message, beginning [completion notifier] wake:. It carries no content; its only purpose is to give an idle supervisor a reason to run.
  • A completion notice, headed [SUBAGENT COMPLETION NOTICE]. It has three fields: the task_id of the delegated task, its status (such as completed or failed), and a short summary of the result. Notices that arrive together are combined into one message.
[SUBAGENT COMPLETION NOTICE]
task_id: dispatch_1_abc123
status: completed
summary: Researched LangGraph store namespaces; key finding in §04.

Nothing visually marks these messages as injected. The model recognizes them by their content, which is why the supervisor's system prompt should describe them.

5. Decisions the supervisor can make

The notice arrives as context, not as the result of a tool the model chose to call, so the supervisor decides what it means. The useful choices are few:

  • Proceed on the summary. If the summary answers the sub-task, the supervisor continues without fetching anything further.
  • Fetch the full result once. If the summary is insufficient, a Deep Agents supervisor calls its own check_async_task, which reads the result directly from the worker's Agent Server. A supervisor built without Deep Agents has no such tool, so the notifier provides get_async_result instead. It fetches the result through the A2A receiver, which already knows which worker sent the notice and holds the credentials to reach it. Either way it is a single fetch, not a poll, because the notice has already established that the task is finished.
  • Respond to failure. A failed status tells the supervisor not to wait for a result that will never arrive. It can re-dispatch the task, send it to a different worker, or report the failure.
  • Wait for the rest, or move on. With several workers outstanding, the supervisor can hold its synthesis until every notice has arrived, or begin the next step with partial results.

A short instruction in the supervisor's system prompt is enough to establish this behaviour:

When you see a [SUBAGENT COMPLETION NOTICE]:
1. If the summary fully answers the sub-task, continue without fetching.
2. Otherwise call check_async_task(task_id) once -- never poll, never re-fetch.
3. Use the returned result and continue.

The result is a supervisor that reacts to its workers rather than asking after them. Each turn begins with whatever has finished, and the model spends its reasoning on what to do next.

6. Next steps: a federated fleet experiment

The work so far connects a LangGraph supervisor to LangGraph workers. The next step is an experiment with a federated fleet: agents built by different teams, with different harnesses, delegating work to one another inside a single organization.

The premise is that organizations will not standardize on one agent framework. Teams will choose Deep Agents, PydanticAI, Agno or whatever suits the job, and agents will appear organically. What those agents need in common is not a shared runtime, but a trusted way to hand off work and report back. The experiment's promise is therefore bring your own harness: each team keeps the framework it chose, and the federation supplies only the hand-off and the trust.

The experiment has three parts:

  1. A Deep Agents A2A client. Today, a Deep Agents supervisor can delegate asynchronously only to Agent Protocol servers, and an open Deep Agents issue asks for A2A support. The proposed client keeps the five async-subagent tools the model already knows, but lets each subagent use either Agent Protocol or A2A. One supervisor can then delegate to the whole fleet with a single vocabulary.
  2. Push configuration in native A2A. The supervisor's webhook and routing token would be registered in the A2A request itself, as the specification defines, and the A2A receiver would accept standard pushes from any compliant server. Its result fetch, get_async_result, would likewise learn to use A2A's own tasks/get, making it the framework-neutral way to collect a finished worker's output. The notifier's extras become an optional A2A extension, advertised on the agent card: a signed sender, a short summary, and delivery that is safe to retry.
  3. A push-capable server where a framework lacks one. PydanticAI's A2A server does not yet send push notifications, so the experiment includes a small kit that serves PydanticAI agents through the official A2A SDK, which does. Workers without push can still fall back to streaming or polling.

Security runs through all three parts. Every completion notice carries a routing token that only the supervisor can read and, with the extension, a signature identifying the agent that sent it. An agent joins the fleet by publishing an agent card and being registered as a trusted issuer, not by adopting a shared framework.

Sharing tools: an MCP gateway or per-agent OAuth

A federated fleet must also share tools. Agents from different teams need the same MCP servers, and how they reach those servers decides whether a bring-your-own-harness deployment scales.

The direct approach gives each agent its own MCP connections, secured with OAuth or API keys. It works for one agent but multiplies poorly. Every harness has its own MCP client and its own auth support, so credentials, allowlists and audit end up scattered across teams. Unattended agents cannot complete an interactive OAuth consent flow, and stdio servers or servers on a private network offer no endpoint a remote agent can reach.

The alternative is an MCP gateway acting as a reverse proxy: one front door through which every agent reaches every tool. Gateways such as IBM ContextForge, Docker MCP Gateway and Microsoft mcp-gateway already take this shape. Each agent authenticates to the gateway with its own short-lived, scoped credential. The gateway holds the upstream credentials, and the MCP authorization specification forbids passing the agent's token through to the server. A human completes OAuth consent once, out of band; the gateway stores the token, and the agent only ever holds a handle.

Per-agent MCP with OAuth

MCP gateway as reverse proxy

Credentials

Held by each agent

Held by the gateway; each agent holds a scoped handle

OAuth consent

Interactive, per agent

Once, out of band, stored at the gateway

stdio and private servers

Unreachable from a remote agent

Hosted or fronted by the gateway

Policy and audit

Per harness, per team

Organization-wide at the gateway

Adding a new harness

Rebuild auth and policy for it

Point its MCP client at the gateway

The gateway does not remove the need for per-agent rules. Each agent still needs its own tool allowlist and context budget, supplied by a thin client-side middleware in front of the gateway. That pairing, a gateway for the organization and a small client for each agent, is the tool-side counterpart of the A2A receiver. In both cases the federation supplies the trust, and each team keeps its harness.

The experiment succeeds if a supervisor can delegate across harnesses with one set of tools, and learn of every completion by push with no polling turns. No completion may be lost or duplicated, a stock A2A server without the extension must still interoperate, and every agent must reach its tools through the gateway under its own credential. Whichever harness produced a result, the supervisor sees the same completion notice in its prompt history and makes the same decisions described above.

Summary

The notifier changes very little about the agents themselves. What it supplies is the trusted hand-off between them: which task finished, with what status, for which conversation, and from which sender. That hand-off is exactly what a federated fleet needs. Whether a worker is built with Deep Agents, PydanticAI or Agno, the supervisor needs only a notice it can trust and a result it can fetch. The tools those agents share call for the same arrangement: one trusted front door, rather than credentials spread across every agent.

The next steps are an experiment, and their outcome is still open. To try the current design, the Langstrata README's labs build it up one step at a time, and Lab 3 shows a completion notice arriving. Feedback, and results from other harnesses, are welcome.