# Multi-Agent Conversational AI

> Supervisor-pattern agent orchestration

Ask a building a question in plain English — five specialist agents over a 4 TB operations database, instead of five admin screens.

- **Page:** https://siddharthdeshpande.com/projects/multi-agent
- **Built at:** Apra Labs
- **Tech stack:** Python, LangChain, LangGraph, GPT-4o / GPT-4o-mini / GPT-5, Azure SQL, Cosmos DB, Azure SignalR, React + Redux

## Impact

- **5** — specialized agents
- **4 TB** — database queried
- **~30** — registry-driven tools
- **300+** — facilities served

## Problem

- 4 TB database with no usable natural-language interface
- Users can't query without SQL expertise or five admin screens
- Answers split across the ops database, a docs archive, and 100+ report types
- New operators took weeks to learn where anything lived
- Experienced engineers spent the day as a human router between systems
- Manual report generation taking days

## Solution

- Supervisor agent routes to 5 specialized agents — never answers itself
- Registry-driven tool system: ~30 tools, decorator-registered, prompt generated at startup
- SignalR streaming of routing, plan, tool calls, observations, and citations
- Seven-collection graph and vector retrieval searched concurrently
- Auto-generated SQL with read-only validation and post-generation permission rewrite
- 100+ report types via twenty dedicated tools plus a generic fallback

## Outcome

- Natural language queries fully operational across 300+ facilities
- Non-technical staff querying data independently
- Real-time streaming of the answer, citations, and reasoning trace
- Five specialized agents running in production
- Report generation reduced from days to seconds
- Live reasoning trace made misroutes obvious before the final answer landed

## How it works

1. **Dispatch** — The web tier authenticates the caller, resolves which sites and tenants they can see, writes a cancellation row, then posts the question and that permission set into the agent tier and drops the connection. Dispatch and response are separate transactions — a deep question chaining four or five tool calls runs well past a normal HTTP timeout, and a page refresh should not kill work already in flight. Results stream back over SignalR on two channels: answer text and citations.
2. **Context & routing** — Prior turns and a rolling conversation summary are fetched in parallel and packed newest-first into a token budget, so a long conversation degrades by dropping the oldest turns rather than by failing. Greetings short-circuit on a keyword check. Everything else goes through three parallel classifiers — depth, cost tier, and owning agent — which resolve to a concrete model (a cheap small model for a device lookup, a frontier model for multi-step troubleshooting) and one of five specialists: Reporting, Admin, System Setup, Knowledge, or Support Tickets. The supervisor never answers anything itself.
3. **Agent loop** — The chosen agent pulls its tool registry and runs a reason-act loop: call the model, decide whether to continue, run tools, repeat. Query tools generate SQL from a schema description — a validator rejects anything that isn't a read, then a permission service rewrites the query with the WHERE clauses that user's scope allows. Retrieval tools hit all seven collections concurrently (chunks, entities, relationships, topics, community summaries, historical Q&A, document metadata), merge and rank, then synthesise with a source manifest. Report tools parse dates and filters out of the question, call the reporting service, and summarise the returned file.
4. **Finalise & stream** — History is saved, the conversation summary is regenerated incrementally, and the client commits the message once both streams close. Multi-turn state is maintained with context windowing. Fallback chains activate when the primary agent can't resolve. The streamed reasoning trace — routing choice, plan, tool calls, observations and citations — is visible while the answer is still forming.

## Key decisions

- **Supervisor pattern, not a single monolithic agent** — A single agent with every tool attached picked the wrong tool constantly — thirty tool descriptions in one prompt give the model almost no signal to separate them. Reporting, admin operations, system setup, knowledge retrieval, and ticket management all have different tool sets and context needs. Narrower agents with about five tools each fixed it. The router classifies depth, cost tier, and owner, then hands off. The cost is one extra model round trip before any real work starts.
- **Dispatch and response are separate transactions** — The web tier validates a request, hands it to the agent tier, and returns immediately. Everything after that streams back over a socket channel. A page refresh does not kill work already in flight. The cost is two failure surfaces instead of one, plus a client that has to commit a message it never got a direct HTTP response for.
- **Authorisation is applied after generation, never requested in the prompt** — The model drafts SQL from a schema description. Before execution, a validator rejects anything that isn't a read, and a permission service rewrites the query with the WHERE clauses that user's scope allows. Asking the prompt to respect a tenant boundary would work most of the time — and most of the time is how you leak one site's access records into another's answer.
- **Registries instead of a hand-maintained routing prompt** — Each agent declares its capabilities in a registry rather than hardcoding tool access. Agents and tools register themselves with a decorator, and the supervisor's routing prompt is generated from that registry at startup. The hand-edited list drifted from reality within two weeks of the first new tool shipping, and the failure was silent: routing just quietly got worse.
- **Answer text and citations stream on separate SignalR channels** — LLM responses over a 4 TB database take time. SignalR streaming shows users partial results as they are generated. Sources resolve on a different clock from the prose, and interleaving them into one stream meant a citation could land against a sentence that had already scrolled past. The message commits only when both finish.

## Evaluation

This shipped without a formal evaluation harness, and that's the biggest gap in the project. What existed instead: a fixed set of known-answer questions per agent, re-run by hand after any prompt or schema change, and manual review of generated SQL against the tables I expected it to touch. The streamed reasoning trace helped more than I expected — because routing, plan and tool calls are all visible, a misroute is obvious in a way it never is when you only see the final answer. If I rebuilt this I'd measure routing accuracy against a labelled question set first, then query correctness, retrieval relevance, and cost per resolved question broken down by tier.

## What I'd change

- **Build the evaluation harness before the second agent** — Every other item below is downstream of not measuring. A labelled set of a few hundred real questions with expected agent, expected tables and expected answer would have taken a week and paid for itself immediately.
- **Generate SQL for the tail, not the head** — The same twenty questions account for most traffic, and free-form generation over a wide schema is a fragile way to answer them. Parameterised query templates the model selects between would be faster, cheaper and verifiable. Keep generation for the genuinely open-ended remainder.
- **Better agent handoff protocol** — When a query spans two agents, handoff is clunky. A cleaner multi-agent coordination protocol would handle compound requests more naturally.
- **Cancellation deserved a real primitive** — A flag row polled at checkpoints works, but a request can only die where I remembered to check, and I did not always remember. Long tool calls could run several seconds past a user hitting stop.
- **Two chat clients was one too many** — An embedded widget and a full-page experience shipped separately and duplicated the streaming logic. Every protocol change then needed doing twice, and the second copy was usually the one with the bug.
- **Adaptive context windows — and versioned rolling summaries** — Fixed window sizes per agent type. Adaptive windowing based on query complexity would improve accuracy on longer sessions. Rolling summaries were one file per conversation, rewritten every turn, with last write winning — cheap until two tabs were open on the same conversation.
