4 min read

Langstrata - A blueprint for deploying Langgraph with dual-agent servers

Langstrata - A blueprint for deploying Langgraph with dual-agent servers
Langstrata - A curated collection of Langchain/Langgraph features for building highly scalable multi-agent solutions.

Establishing a blueprint for how Langchain/Langgraph multi-agents are implemented, deployed and scaled is possible. The approach requires a bit more upfront infrastructure, but removes the constraints that make multi-agent systems difficult to scale and observe. Bespoke Langchain multi-agent development becomes a design choice rather than a requirement.

I propose a novel dual-agent server pattern consisting of a supervisor server and one or more subagent servers. The supervisor server hosts a dedicated deep agent. A second and additional agent servers host the subagents.

Support for structured responses and asynchronous agents are essential to the design. The former improves supervisor orchestration of subagents and the latter simplifies scalability.

Deploying a Langchain dual-agent server requires implementing a secure and standardized mechanism that subagent's integrate as middleware. In Langchain parlance, the supervisor and subagent infrastructure that a subagent employs to update the supervisor is called a completion notifier. The A2A notifications specification fills this gap. A2A was developed to facilitate interoperability and communication between independently designed agents. This proposal flips the use case to facilitate communication between agents designed for internal use. In lieu of building bespoke tools to fill gaps in existing frameworks, adopt A2 whenever possible. In this case, its choosing A2A notifications to build the completion notifier.

💡
Please refer to the code comments in the Langchain Deep agents example completion_notifier.py file to learn more about completion notifiers.

The proposed A2A completion notifier will be packaged as a MCP server. Ensuring it integrates using the same configuration pattern as any other MCP tool and is not tied to any particular agent implementation. The implementation will automatically mitigate duplicate and bursts of subagent notifications. I plan to post the implementation to my stokomax/langshark-bites repo when it is available. You simply incorporate the completion notifier in your agent the same way as existing langshark-bites.

The completion notifier hardens the supervisor and reduces it token consumption by making the supervisor event-driven.

Langshark-bites is a collection of bite-size add-ons and wrappers for building durable and scalable LangChain multi-agent solutions. I created the collection to solve common, everyday, infrastructure issues when building multi-agent solutions.

Learn more

The dual-agent server architecture isolates the supervisor server runtime from the subagent server runtime. A supervisor can freely launch one or more ReAct agent threads as needed without those threads affecting the supervisor's own scheduling. This guards the supervisor agent performance from runaway subagent activity.

The main classes of agent design patterns within Langchain are: workflows, which have multiple variants, ReAct agents, and Deep Agents. You mix-and-match these as you choose. Workflows are efficient but rigid. That rigidity means error handling has to be designed in from the start, which is workable when the failure modes are known. Given LLM outputs are non-deterministic, hard coding error handling is a best effort at best.

A Deep Agent, designed well, provides a built-in a multi-agent management interface through its Todo list feature. If a multi-agent execution happens to abort midway through several agent handoffs, a Deep Agent can recover with human-in-the-loop assistance. Well defined error patterns are incrementally added to the agent as needed – reducing the need for human intervention in the process. This is why I am all-in on building deployments with Deep Agents.

If you have interest, as I do, in experimenting with prompt versioning, Reflexion-style agent and evaluation frameworks, the deployment challenge is managing and observing that activity. I avoided experimenting with these features until I landed on a standardized deployment model. The dual-agent server solution ensures a robust runtime no matter the subagent activity and therefore is capable of hosting more complicated agent features.

Subagent servers should be deployed using a workflow (i.e. StateGraph) combined with map-reduce and the Send API, or as one or more ReAct agents. The workflow approach is the better choice for keeping resource requirements low while scaling horizontally — each worker is a lightweight graph node rather than a full agent process. The ReAct approach is simpler to build initially, though heavier in terms of thread and checkpoint resource usage.

The asynchronous agent requirement is a design mandate rather than a technical requirement. The goal is agents that are independent and cooperative from the start, not as an afterthought. Implementing your agents asynchronously is the best path toward that and provides natural scalability to your solution. I learned from my IoT and embedded background that when dealing with non-deterministic endpoints, an asynchronous listener implementation wins over synchronous approaches every time.

For self-hosted deployments, the Redis server and PostgreSQL server can be shared between agent servers. Adding a new agent server is a matter of adding a service to your compose file and assigning it a non-overlapping port. Multiple langgraph dev servers can be launched for prototyping without a full Postgres and Redis stack.

Self-hosted Langgraph Servers

By adopting the Langstrata blueprint you can scale agent experiments without wrestling with Langchain infrastructure. The architecture accommodates more agents, more complex subagent designs, and heavier workloads without requiring structural changes to the deployment.

CTA Image

Clone the langstrata repo and run the examples.

Learn more

Next steps:

  • Make the langstrata demo agents your own. Replace the researcher, coder and analyst workers with graphs and tools that fit your domain,
    that regenerate the config files and Dockerfiles as described in the langstrata REAME.md file.
  • Harden your workers. As your workers fan out and call more
    external APIs, the langshark-bites collection addresses exactly those problems — rate limiting, visible backoff, provider failover, tolerant JSON output parsing, and reducers that merge parallel Send results without duplicates.
  • Make the supervisor event-driven. Today the supervisor polls
    for worker completion — the deep-agent flow launches a subagent with
    start_async_task and then repeatedly checks check_async_task until
    it finishes, which burns supervisor turns and adds latency under load.
    The hardening step is to be notified when a worker completes instead
    of asking. That event-driven completion is the motivation behind the
    planned completion_notify bite; it isn't available yet, so it remains
    a known gap in this template.