Back

Multi-Agent Conversational AI

Supervisor-pattern agent orchestration

Ask a building a question in plain English — five specialist agents over a 4 TB operations database, instead of five admin screens.

Stack

PythonLangChainLangGraphGPT-4o / GPT-4o-mini / GPT-5Azure SQLCosmos DBAzure SignalRReact + Redux

Built at Apra Labs

What made it hard

A 4 TB operations database behind a product serving 300+ facilities, with answers split across that database, a documents archive, and 100+ report types. Non-technical operators could not query any of it without SQL or five separate admin screens, so experienced engineers were spending the day as a human router between systems. Generating SQL was never the hard part. Doing it without letting a generated query read across a tenant boundary was.

Problem

  • 4 TB database with no usable natural-language interface
  • Users can't query without SQL expertise or five admin screens
  • Answers split across the ops database, a docs archive, and 100+ report types
  • New operators took weeks to learn where anything lived
  • Experienced engineers spent the day as a human router between systems
  • Manual report generation taking days

Solution

  • Supervisor agent routes to 5 specialized agents — never answers itself
  • Registry-driven tool system: ~30 tools, decorator-registered, prompt generated at startup
  • SignalR streaming of routing, plan, tool calls, observations, and citations
  • Seven-collection graph and vector retrieval searched concurrently
  • Auto-generated SQL with read-only validation and post-generation permission rewrite
  • 100+ report types via twenty dedicated tools plus a generic fallback

Outcome

  • Natural language queries fully operational across 300+ facilities
  • Non-technical staff querying data independently
  • Real-time streaming of the answer, citations, and reasoning trace
  • Five specialized agents running in production
  • Report generation reduced from days to seconds
  • Live reasoning trace made misroutes obvious before the final answer landed

5

specialized agents

4 TB

database queried

~30

registry-driven tools

300+

facilities served

Architecture Review

System design · Pipeline · Decisions

System Architecture

How It Works

  1. 1

    Dispatch

    The web tier authenticates the caller, resolves which sites and tenants they can see, writes a cancellation row, then posts the question and that permission set into the agent tier and drops the connection. Dispatch and response are separate transactions — a deep question chaining four or five tool calls runs well past a normal HTTP timeout, and a page refresh should not kill work already in flight. Results stream back over SignalR on two channels: answer text and citations.

  2. 2

    Context & routing

    Prior turns and a rolling conversation summary are fetched in parallel and packed newest-first into a token budget, so a long conversation degrades by dropping the oldest turns rather than by failing. Greetings short-circuit on a keyword check. Everything else goes through three parallel classifiers — depth, cost tier, and owning agent — which resolve to a concrete model (a cheap small model for a device lookup, a frontier model for multi-step troubleshooting) and one of five specialists: Reporting, Admin, System Setup, Knowledge, or Support Tickets. The supervisor never answers anything itself.

  3. 3

    Agent loop

    The chosen agent pulls its tool registry and runs a reason-act loop: call the model, decide whether to continue, run tools, repeat. Query tools generate SQL from a schema description — a validator rejects anything that isn't a read, then a permission service rewrites the query with the WHERE clauses that user's scope allows. Retrieval tools hit all seven collections concurrently (chunks, entities, relationships, topics, community summaries, historical Q&A, document metadata), merge and rank, then synthesise with a source manifest. Report tools parse dates and filters out of the question, call the reporting service, and summarise the returned file.

  4. 4

    Finalise & stream

    History is saved, the conversation summary is regenerated incrementally, and the client commits the message once both streams close. Multi-turn state is maintained with context windowing. Fallback chains activate when the primary agent can't resolve. The streamed reasoning trace — routing choice, plan, tool calls, observations and citations — is visible while the answer is still forming.

Constraints

  • Generated SQL runs against live production data, so read-only validation and post-generation permission rewrite are non-negotiable
  • Per-tenant access rules already existed and could not be bypassed or re-implemented in the agent layer
  • 4 TB across the operations database means query shape decides whether an answer takes seconds or never returns
  • Answers must carry citations — an unsourced answer about a building is worse than no answer
  • Users need to see progress before the final answer lands, or a multi-agent route feels broken
  • Model cost across ~30 tools required routing cheap models to most calls and reserving the expensive one for reasoning

Key Decisions

Supervisor pattern, not a single monolithic agent

A single agent with every tool attached picked the wrong tool constantly — thirty tool descriptions in one prompt give the model almost no signal to separate them. Reporting, admin operations, system setup, knowledge retrieval, and ticket management all have different tool sets and context needs. Narrower agents with about five tools each fixed it. The router classifies depth, cost tier, and owner, then hands off. The cost is one extra model round trip before any real work starts.

Dispatch and response are separate transactions

The web tier validates a request, hands it to the agent tier, and returns immediately. Everything after that streams back over a socket channel. A page refresh does not kill work already in flight. The cost is two failure surfaces instead of one, plus a client that has to commit a message it never got a direct HTTP response for.

Authorisation is applied after generation, never requested in the prompt

The model drafts SQL from a schema description. Before execution, a validator rejects anything that isn't a read, and a permission service rewrites the query with the WHERE clauses that user's scope allows. Asking the prompt to respect a tenant boundary would work most of the time — and most of the time is how you leak one site's access records into another's answer.

Registries instead of a hand-maintained routing prompt

Each agent declares its capabilities in a registry rather than hardcoding tool access. Agents and tools register themselves with a decorator, and the supervisor's routing prompt is generated from that registry at startup. The hand-edited list drifted from reality within two weeks of the first new tool shipping, and the failure was silent: routing just quietly got worse.

Answer text and citations stream on separate SignalR channels

LLM responses over a 4 TB database take time. SignalR streaming shows users partial results as they are generated. Sources resolve on a different clock from the prose, and interleaving them into one stream meant a citation could land against a sentence that had already scrolled past. The message commits only when both finish.

Multi-Agent Conversational AI — Architecture

Multi-Agent Conversational AI architecture

Evaluation

This shipped without a formal evaluation harness, and that's the biggest gap in the project. What existed instead: a fixed set of known-answer questions per agent, re-run by hand after any prompt or schema change, and manual review of generated SQL against the tables I expected it to touch. The streamed reasoning trace helped more than I expected — because routing, plan and tool calls are all visible, a misroute is obvious in a way it never is when you only see the final answer. If I rebuilt this I'd measure routing accuracy against a labelled question set first, then query correctness, retrieval relevance, and cost per resolved question broken down by tier.

What I'd Change

  • Build the evaluation harness before the second agent
  • Generate SQL for the tail, not the head
  • Better agent handoff protocol
  • Cancellation deserved a real primitive
  • Two chat clients was one too many
  • Adaptive context windows — and versioned rolling summaries

Get in Touch

Let's build
something great.