Agent Loop -- createAgent / delegate / chain / consolidate / parallel
The agent layer runs a multi-step tool loop over an LLM: call the model, execute
tools it requests, feed results back, repeat until the model stops requesting
tools. The loop is built into complete() when you pass tools, but createAgent
gives you a stateful, reusable agent with persistent history and richer lifecycle
hooks.
When to reach for this
Section titled “When to reach for this”- You need a stateful agent that remembers conversation history across multiple
user turns (use
createAgent). - You want to compose agents: one agent delegates subtasks to another (
delegate), runs steps in sequence (chain), in parallel (parallel), or resolves disagreement between multiple agents (consolidate). - You need to observe agent events (tool calls, run completion) reactively
(
createObserver).
Main exports
Section titled “Main exports”| Export | What it does |
|---|---|
createAgent(opts) | Builds an AgentLoop with an optional pre-built LLMClient or a model string. Wires hooks from the engine. |
AgentLoop | The loop class. .complete(prompt) runs one conversation turn (tool loop included). .stream(prompt) streams events. |
delegate(name, description, agent) | Wraps an AgentLoop as an AgentTool so a parent agent can call it by name. The tool passes a task: string and returns the sub-agent’s reply. |
chain(steps, opts) | Sequential pipeline: each step’s output string becomes the next step’s input. Steps are either complete() call configs or plain async functions. |
parallel(tasks, opts) | Run multiple complete() calls simultaneously; returns all results. |
consolidate(opts) | Multi-agent debate: N agents answer in parallel over rounds, a judge LLM decides agreement, the loop ends early on consensus and produces a summary. |
createObserver(agent, event, reactor) | Subscribe to an agent lifecycle event; reactor is a plain async function or itself an agent config. |
ConversationHistory | Stores and replays the agent’s message history. Importable/exportable as a snapshot for persistence. |
ContextRegistry | Layered system-prompt builder. The agent loop populates it; you can write custom layers (e.g. facts, user profile). |
Type-only exports: AgentLoopConfig, AgentTool, AgentStreamEvent,
AgentRunReport, HistorySnapshot, ContextLayer, and related.
Minimal examples
Section titled “Minimal examples”Stateful agent (multi-turn)
Section titled “Stateful agent (multi-turn)”import { createEngine, createAgent, defineTool } from '@combycode/llm-sdk';
createEngine({ catalog: 'defaults', apiKeys: { anthropic: process.env.ANTHROPIC_API_KEY! },});
const agent = createAgent({ model: 'anthropic/claude-haiku-4.5', system: 'You are a helpful assistant.', tools: [ defineTool({ name: 'get_time', description: 'Return the current UTC time.', params: {}, execute: () => new Date().toISOString(), }), ],});
const r1 = await agent.complete('What time is it?');console.log(r1.text);
const r2 = await agent.complete('Add one hour to that time.');console.log(r2.text); // agent remembers r1's contextDelegate — agent as a tool
Section titled “Delegate — agent as a tool”import { createAgent, delegate, complete } from '@combycode/llm-sdk';
const researcher = createAgent({ model: 'anthropic/claude-haiku-4.5', system: 'You are a research specialist. Answer factual questions concisely.', apiKey: process.env.ANTHROPIC_API_KEY,});
const { text } = await complete({ model: 'anthropic/claude-haiku-4.5', apiKey: process.env.ANTHROPIC_API_KEY, prompt: 'Summarize the key facts about the Eiffel Tower.', tools: [delegate('research', 'Look up factual information on a topic.', researcher)],});console.log(text);Chain — sequential pipeline
Section titled “Chain — sequential pipeline”import { chain } from '@combycode/llm-sdk';
const pipeline = chain([ { model: 'anthropic/claude-haiku-4.5', apiKey: process.env.ANTHROPIC_API_KEY, name: 'summarize', prompt: (input) => `Summarize this in one sentence: ${input}`, maxTokens: 80, }, { model: 'anthropic/claude-haiku-4.5', apiKey: process.env.ANTHROPIC_API_KEY, name: 'translate', prompt: (input) => `Translate to French: ${input}`, maxTokens: 80, },]);
const result = await pipeline('The sky is blue because of Rayleigh scattering of sunlight.');console.log(result);Consolidate — multi-agent debate
Section titled “Consolidate — multi-agent debate”import { consolidate } from '@combycode/llm-sdk';
const result = await consolidate({ agents: [ { name: 'Analyst A', model: 'anthropic/claude-haiku-4.5', system: 'You are a financial analyst.' }, { name: 'Analyst B', model: 'openai/gpt-5.4-nano', system: 'You are a risk analyst.' }, ], task: 'Should a startup invest in GPU hardware or rent cloud compute?', judge: { model: 'anthropic/claude-opus-4.8' }, rounds: 3, onRound: ({ round, agreed }) => console.log(`Round ${round}: agreed=${agreed}`),});console.log(result.summary);Bounding the tool loop with maxSteps
Section titled “Bounding the tool loop with maxSteps”By default the loop allows up to 16 tool-followup rounds per complete() /
stream() call. If the model keeps requesting tools beyond that limit the loop
stops before the next LLM call and returns with:
AgentRunReport.reason === 'max_steps'CompletionResponse.finishReason === 'length'CompletionResponse.textset to"stopped: reached maxSteps (<N>)"
The cap exists to prevent runaway cost and latency when a model or tool enters a pathological loop.
Configuring the cap
Section titled “Configuring the cap”Pass maxSteps in AgentLoopConfig (or the createAgent options):
const agent = createAgent({ model: 'anthropic/claude-haiku-4.5', apiKey: process.env.ANTHROPIC_API_KEY!, tools: [...], maxSteps: 32, // raise the limit});Values <= 0 are ignored and the default (16) applies. There is no way to
fully disable the cap; set a very large number (e.g. 10_000) if you genuinely
need unbounded execution.
Detecting the cap in callers
Section titled “Detecting the cap in callers”const res = await agent.complete('...');if (res.finishReason === 'length') { // check whether it is a maxSteps stop, not a token-length truncation const report = agent.lastReport; if (report?.reason === 'max_steps') { console.warn('tool loop capped after', report.stepCount, 'steps'); }}The same reason is delivered in the onRunComplete hook payload.
Errors vs bounded stops
Section titled “Errors vs bounded stops”A failed LLM call during a run (auth, rate limit, context overflow, provider error) throws out
of agent.complete() / agent.stream() — the original LLMError (with kind/status) — matching a
no-tools complete(). It never silently returns empty text. For partial results + a status, catch the
error and read agent.lastReport — an AgentRunReport carrying reason, the per-step and per-tool
reports, and usage:
try { await agent.complete('…');} catch (err) { const report = agent.lastReport; // reason, steps, toolCalls, usage}Bounded stops are not errors: max_steps returns finishReason: 'length' (see
above) and a model refusal returns normally — both are inspectable via finishReason / report.reason.
finishReason is an open union, so always write a default branch. Providers keep inventing
terminal states, and the union is open (CONSTITUTION R1) precisely so that a new one is not a
breaking change for every consumer — including consumers of providers that changed nothing. Values
beyond the documented set reach you rather than being flattened: Google’s TOO_MANY_TOOL_CALLS
arrives as too_many_tool_calls instead of stop, and OpenAI’s incomplete_details.reason
distinguishes max_messages (a message cap, not a token cap) and steered (the turn was
superseded and a successor response follows automatically) instead of reporting both as length.
The raw provider value stays on response.raw.
And the answer does not depend on how the turn was fetched: the streamed and buffered paths share one table per provider. They did not, once — the stream carried a copy holding a single entry, so a streamed Google SAFETY block read as a clean stop with no content.
Recovering from a model failure (reflectAndRetry)
Section titled “Recovering from a model failure (reflectAndRetry)”Some failures are the model’s, not the network’s: a malformed tool call, a truncated call, a hallucinated tool name. Resending the identical request would never fix them, so the network retry layer correctly leaves them alone — and the run used to end there.
A tool call whose arguments do not parse is never executed. It used to fall back to {}, which
is a valid call rather than a failed one, so a stream cut at {"path": "/et ran the tool with no
arguments at all. The call is now marked, answered with an error result so the history stays valid,
and the step finishes as malformed_tool_call on every provider — previously only Google
reported that reason, so the same truncation elsewhere looked like a clean turn and nothing below
ever fired.
Interrupted turns are repaired
Section titled “Interrupted turns are repaired”A turn can end between “the model asked for a tool” and “the tool answered”: an early break out of
stream(), an output guardrail tripping, continueOnError: false, a pending approval, a caller who
stopped. Anthropic and OpenAI both reject a history containing a tool call with no result, so the
NEXT run died on send — one turn away from the cause, with an error naming neither.
Every run now repairs the history before its first request: an unanswered call gets a synthetic
result saying the tool never ran, and onWarning fires with unanswered_tool_calls_repaired. The
result is not a pretend success — the model can see the work did not happen. History truncation is
pair-aware for the same reason: a count-based cut that would keep a result whose call it removed
drops the orphan too, so keepLast is a ceiling rather than an exact count.
reflectAndRetry gives the model a bounded number of second chances, telling it what went wrong:
const agent = createAgent({ client, tools: [...], reflectAndRetry: { maxRetries: 2 }, // OFF unless configured — a retry costs a real request});On a triggering finish reason the loop injects structured guidance naming the attempt number and explicitly instructing the model not to repeat the same call (without that, models tend to re-emit identical arguments and burn the whole budget on one mistake), then retries the step. The failed assistant turn is not appended, so the model does not learn from its own broken output.
- Default trigger:
malformed_tool_call. Override withonFinishReasons— a content filter is usually a decision rather than a mistake, so it is not retried by default. - Failures are counted consecutively: a success clears the streak, so an agent that recovers and stumbles again much later gets a fresh budget rather than an inherited one.
- When the budget is spent the loop throws, naming the escape hatch. Set
throwIfExceeded: falseto get the unusable response back instead.
This exists because finishReason now tells the truth about these turns. Google’s
MALFORMED_FUNCTION_CALL was previously unmapped and read as a clean stop with no content — a
failure that looked like a successful empty answer.
Which text is the answer?
Section titled “Which text is the answer?”Codex-family models narrate before answering. Those parts are tagged
phase: 'commentary' versus 'final_answer', and response.text concatenates everything — so
an agent’s output used to include its own thinking-out-loud.
AgentLoop derives its answer with finalAnswerText(), which drops commentary. response.text
and contentText() are unchanged, so callers who want the narration still get it:
import { finalAnswerText } from '@combycode/llm-sdk';
const answer = finalAnswerText(res.content); // commentary excludedconsole.log(res.text); // everything, as beforeStreaming carries the phase too, on the agent event itself — so you can separate narration from the answer live, which is when it matters for a UI:
for await (const ev of agent.stream('…')) { if (ev.type !== 'text') continue; if (ev.phase === 'commentary') renderThinking(ev.text); else appendToReply(ev.text);}finalAnswerText() cannot do this job: it takes a finished message’s content, not deltas. Use it
on the assembled message; use ev.phase on the stream.
AssistantPhase is an open union, and finalAnswerText excludes only what is explicitly
'commentary' rather than keeping only 'final_answer' — the day a provider adds a third phase, an
allow-list would silently drop the answer.
Tool-name collisions
Section titled “Tool-name collisions”Tools are indexed by function name (or builtin type), so registering two under the same key means one silently replaces the other and the model never sees it. That surfaces much later as “the model called the wrong tool”, with nothing in the logs pointing at the cause.
Collisions are now reported. Default 'warn' keeps last-write-wins (changing it would break apps
that rely on a deliberate override) but emits an onWarning with code tool_name_collision naming
which tool lost. toolNameCollisionPolicy: 'error' throws at construction instead, before the model
is ever called.
Watching tool arguments form (tool_call_delta)
Section titled “Watching tool arguments form (tool_call_delta)”A streamed run now forwards the model’s tool-call arguments as they arrive:
for await (const ev of agent.stream('find the Q3 report')) { if (ev.type === 'tool_call_delta') { // Raw JSON TEXT, usually not parseable on its own — for rendering. process.stdout.write(ev.arguments); } if (ev.type === 'tool_call_start') { // The complete, PARSED arguments. This is the event to act on. console.log(ev.toolName, ev.arguments); }}The fragments used to reach the loop and die there — accumulated into the call’s
argument buffer and dropped — so a UI had no way to show a long argument list
forming, and tool_call_start only fires once it is complete. They are now
forwarded as well as accumulated: a second reader, not a handover, because the
loop still needs the whole string to parse at the end of the call.
argumentsis a raw JSON fragment. A partial one is usually not valid JSON, so render it rather than act on it.callIdis the accumulator’s id, not the event’s — several providers omit the id on later fragments, and the forwarded event has to carry something a consumer can group by.stepis the step the fragments belong to, so they can be correlated with the step that produced them.- Absent on providers that stream a call whole (most of them) and on every step that calls no tools, so a consumer that does not want them needs no change.
Bounding a model call (modelTimeout)
Section titled “Bounding a model call (modelTimeout)”const agent = new AgentLoop({ client, tools, modelTimeout: 30_000 });toolTimeout already bounded the tool half of a step. The model half was bounded
only by whatever the client was configured with — so on a long run one slow step
could hold the whole run open past any deadline the caller thought they had set.
Applied per step, not per run: a nine-step run with modelTimeout: 30_000
allows each step thirty seconds, not the run. A run-wide budget is a different
thing and you already have it — an AbortSignal you control.
A per-call ExecuteOptions.timeout still wins, so one complete() can ask for
longer. There is no new error type: the timeout surfaces as the
LLMError{ kind: 'timeout' } the network layer already raises, which is what code
catching timeouts already matches on. (Upstream names a ModelTimeoutError; a
second class for a condition we already report would mean every consumer has to
learn both.)
Checking tool arguments (validateToolArguments)
Section titled “Checking tool arguments (validateToolArguments)”const agent = new AgentLoop({ client, tools, validateToolArguments: true });Checks a tool call’s arguments against that tool’s own parameters schema before
running it. On a failure the tool is not executed and the errors go back to the
model as the tool’s result:
Invalid arguments for "lookup": $.city: expected string, got number.Call the tool again with arguments matching its schema.That shape is deliberate on both counts. It is a result, not an exception, because the model asked for something its own schema forbids — a thing it can fix on the next step — and ending the run would discard every step before it over a mistake the model usually corrects when told. And it says what to do: a bare validator message reads as an internal error, which models answer by apologising rather than by re-calling the tool.
The bound is maxSteps, the loop’s existing one, rather than a second retry budget
to tune that would give the same answer. Each refusal emits onWarning with code
tool_arguments_invalid, so a model that never gets it right is visible instead of
quietly eating the step budget.
Off by default for the same honest reason as structured.validate: the bundled
validator reads the common JSON Schema keywords, not all of Draft 2020-12, so on by
default it would refuse calls that are valid under a schema it cannot fully read.
Where a provider’s own strict mode is available that is the better guarantee — this
is for the models and surfaces where it is not, and for schemas strict mode cannot
express. A builtin tool ({ type: 'web_search' }) has no parameters and is left
alone.
Backup models (fallbackClients)
Section titled “Backup models (fallbackClients)”route() falls over between models for a one-shot complete(). A run is where it
matters more: a rate limit on step 7 of a nine-step run threw away six steps of
work and every tool call they paid for, and the only recourse was to start again.
const agent = new AgentLoop({ client: primary, // openai/gpt-5.6-sol fallbackClients: [backup], // anthropic/claude-haiku-4.5 tools: [...],});Four rules, and the third is the one that cannot be got wrong:
- Each client is tried once per step. Retrying one client is the network engine’s job; doing it here too would retry a single failure twice over, at two layers.
- Only a failure another model could survive moves on —
rate_limit,server_error,model_not_found,timeout,network,quota_exceeded,unsupported. An auth failure, a malformed request, a content filter or a prompt that is simply too long is the same request failing the same way everywhere, so it propagates immediately instead of being offered to every backup in turn. Override the set withfallbackOn. - A streamed step stops being able to fall over at its first event. A streamed turn can fail after several chunks, and the consumer has already rendered half an answer; a backup would start a different one mid-sentence and the two would be spliced into one turn. So the error reaches the caller.
- Every step starts from the primary again. A rate limit is transient, and a run that fell over once should not spend the rest of its life on the backup.
When every client fails, the last one’s error is what you get: the provider’s own message is the useful half, and a wrapper saying “all models failed” buries it.
Each hand-off emits onWarning with code model_fallback and
details: { from, to, kind }, per step — a silent fallback is a latency and cost
change nobody can see.
What gets attributed to whom. agent.model still reports the primary: it is
read before any request is made, and a caller asking what model an agent is
configured with means the primary. But each step’s history entry and span name
whoever actually served, and that is not cosmetic — provenance is model-bound, so
a stateful continuation (previous_response_id, previous_interaction_id) is only
valid against the model that issued the state. Recording the primary on a turn the
backup produced would have the next step offer the backup’s server state to the
primary.
Backups are supplied to AgentLoop.restore() the same way the client is — a
snapshot carries neither, because both are live objects, and a restore that
silently lost them would resume a run less resilient than the one it continues.
const resumed = AgentLoop.restore(snapshot, { client: primary, tools, fallbackClients: [backup],});Server-state continuation
Section titled “Server-state continuation”On a stateful API (OpenAI Responses, Google Interactions) the loop automatically continues by id
(previous_response_id / previous_interaction_id) between tool rounds instead of resending the whole
transcript — gated by the catalog’s retention TTL (openai/xai 30d, google 72h) and model binding. No
configuration needed; it falls back to full history when the stored turn is too old or the provider
isn’t stateful.
Naming an agent for telemetry
Section titled “Naming an agent for telemetry”An unnamed agent exports as a bare invoke_agent carrying only its generated id, and that id changes
per process — so a trace cannot say which of your agents ran, and two runs of the same agent cannot be
compared. Three optional fields fix that:
const agent = createAgent({ model: 'anthropic/claude-haiku-4.5', label: 'briefing', // names the span: `invoke_agent briefing` source: 'customer', // which part of YOUR system this belongs to attributes: { 'app.tenant': 'acme' }, // anything the two fixed fields do not cover});They change no behaviour and cost nothing when telemetry is off — they travel with the agent’s spans and nothing else. See Observability / Telemetry for how each is exported.