# Designing a multi-agent pipeline

> The agent is rarely the hard part of an agent system. The queue, the budget, the durable record and the independent check are — here is what I keep across them.

In an agent system the agent is rarely the hard part. The queue in front of it, the budget around it, the durable record underneath it and the independent check after it are. I've built this a few times — the runtime behind an AI-visibility platform, a research cascade for startup validation, a content conveyor that waits for human approval before it ships — and the scaffolding is what carries over between them.

## Why a pipeline, not a prompt

A prompt returns an answer. To return the same answer reliably, at a cost you can predict, you need the stages around it: a plan you can inspect, state you can replay, and a step that can reject the result. That is the pipeline, and it is most of the engineering.

### The shape of a run

Every run moves through the same stages and writes to the same record. If the only way to find out what happened is to run it again, the record isn't doing its job.

## Put slow work behind a queue

Agent calls are slow and sometimes expensive, so the API doesn't make them inline. It enqueues the work and background workers run it. The API stays responsive, a failed step retries without holding a request open, and a duplicate request collapses onto the run already in flight instead of paying twice.

Give each run a budget and stop when it's spent:

```ts
interface RunBudget {
  maxSteps: number;
  maxTokens: number;
  deadline: number; // epoch ms
}

function withinBudget(spent: RunBudget, limit: RunBudget): boolean {
  return (
    spent.maxSteps <= limit.maxSteps &&
    spent.maxTokens <= limit.maxTokens &&
    Date.now() <= limit.deadline
  );
}
```

### Retries without duplication

Retries are fine; repeated side effects are not. Key each step so a retry is idempotent, and a second attempt confirms the first instead of redoing it.

## Make state explicit

Each run is a durable record, and each step writes its inputs and outputs. The trace then is the source of truth, not a reconstruction after the fact.

| Concern      | Where it lives      | Why                                  |
| ------------ | ------------------- | ------------------------------------ |
| Run metadata | `runs` table        | One row per run, status and budget   |
| Step I/O     | `steps` table       | Inputs, outputs, timing per step     |
| Artifacts    | object storage      | Large payloads kept out of the DB    |

## The producer can't grade itself

An agent reviewing its own output tells you nothing. The pipelines I trust run an independent pass: a verifier that didn't produce the thing it checks, a review the author can't approve. Bound it with a budget so it can't loop, and when something is still open after the budget runs out, send it to a human rather than mark it done.

## Observability

If you can't see why an agent did what it did, you can't run it in production. A span per step turns a run into a trace you can read top to bottom, so "the output is wrong" becomes "the planner picked the wrong tool at step four", which is a bug you can fix.

## What carries over

None of this depends on the framework or the model. It's the same discipline whether the agents are summarising a document or deciding whether a page is ready to publish: make the work observable, and make something other than the producer sign off on it.
