Skip to content

Agent Loop -- createAgent / delegate / chain / consolidate / parallel

The agent layer runs a multi-step tool loop over an LLM: call the model, execute tools it requests, feed results back, repeat until the model stops requesting tools. The loop is built into complete() when you pass tools, but createAgent gives you a stateful, reusable agent with persistent history and richer lifecycle hooks.

  • You need a stateful agent that remembers conversation history across multiple user turns (use createAgent).
  • You want to compose agents: one agent delegates subtasks to another (delegate), runs steps in sequence (chain), in parallel (parallel), or resolves disagreement between multiple agents (consolidate).
  • You need to observe agent events (tool calls, run completion) reactively (createObserver).
ExportWhat it does
createAgent(opts)Builds an AgentLoop with an optional pre-built LLMClient or a model string. Wires hooks from the engine.
AgentLoopThe loop class. .complete(prompt) runs one conversation turn (tool loop included). .stream(prompt) streams events.
delegate(name, description, agent)Wraps an AgentLoop as an AgentTool so a parent agent can call it by name. The tool passes a task: string and returns the sub-agent’s reply.
chain(steps, opts)Sequential pipeline: each step’s output string becomes the next step’s input. Steps are either complete() call configs or plain async functions.
parallel(tasks, opts)Run multiple complete() calls simultaneously; returns all results.
consolidate(opts)Multi-agent debate: N agents answer in parallel over rounds, a judge LLM decides agreement, the loop ends early on consensus and produces a summary.
createObserver(agent, event, reactor)Subscribe to an agent lifecycle event; reactor is a plain async function or itself an agent config.
ConversationHistoryStores and replays the agent’s message history. Importable/exportable as a snapshot for persistence.
ContextRegistryLayered system-prompt builder. The agent loop populates it; you can write custom layers (e.g. facts, user profile).

Type-only exports: AgentLoopConfig, AgentTool, AgentStreamEvent, AgentRunReport, HistorySnapshot, ContextLayer, and related.

import { createEngine, createAgent, defineTool } from '@combycode/llm-sdk';
createEngine({
catalog: 'defaults',
apiKeys: { anthropic: process.env.ANTHROPIC_API_KEY! },
});
const agent = createAgent({
model: 'anthropic/claude-haiku-4.5',
system: 'You are a helpful assistant.',
tools: [
defineTool({
name: 'get_time',
description: 'Return the current UTC time.',
params: {},
execute: () => new Date().toISOString(),
}),
],
});
const r1 = await agent.complete('What time is it?');
console.log(r1.text);
const r2 = await agent.complete('Add one hour to that time.');
console.log(r2.text); // agent remembers r1's context
import { createAgent, delegate, complete } from '@combycode/llm-sdk';
const researcher = createAgent({
model: 'anthropic/claude-haiku-4.5',
system: 'You are a research specialist. Answer factual questions concisely.',
apiKey: process.env.ANTHROPIC_API_KEY,
});
const { text } = await complete({
model: 'anthropic/claude-haiku-4.5',
apiKey: process.env.ANTHROPIC_API_KEY,
prompt: 'Summarize the key facts about the Eiffel Tower.',
tools: [delegate('research', 'Look up factual information on a topic.', researcher)],
});
console.log(text);
import { chain } from '@combycode/llm-sdk';
const pipeline = chain([
{
model: 'anthropic/claude-haiku-4.5',
apiKey: process.env.ANTHROPIC_API_KEY,
name: 'summarize',
prompt: (input) => `Summarize this in one sentence: ${input}`,
maxTokens: 80,
},
{
model: 'anthropic/claude-haiku-4.5',
apiKey: process.env.ANTHROPIC_API_KEY,
name: 'translate',
prompt: (input) => `Translate to French: ${input}`,
maxTokens: 80,
},
]);
const result = await pipeline('The sky is blue because of Rayleigh scattering of sunlight.');
console.log(result);
import { consolidate } from '@combycode/llm-sdk';
const result = await consolidate({
agents: [
{ name: 'Analyst A', model: 'anthropic/claude-haiku-4.5', system: 'You are a financial analyst.' },
{ name: 'Analyst B', model: 'openai/gpt-5.4-nano', system: 'You are a risk analyst.' },
],
task: 'Should a startup invest in GPU hardware or rent cloud compute?',
judge: { model: 'anthropic/claude-opus-4.8' },
rounds: 3,
onRound: ({ round, agreed }) => console.log(`Round ${round}: agreed=${agreed}`),
});
console.log(result.summary);

By default the loop allows up to 16 tool-followup rounds per complete() / stream() call. If the model keeps requesting tools beyond that limit the loop stops before the next LLM call and returns with:

  • AgentRunReport.reason === 'max_steps'
  • CompletionResponse.finishReason === 'length'
  • CompletionResponse.text set to "stopped: reached maxSteps (<N>)"

The cap exists to prevent runaway cost and latency when a model or tool enters a pathological loop.

Pass maxSteps in AgentLoopConfig (or the createAgent options):

const agent = createAgent({
model: 'anthropic/claude-haiku-4.5',
apiKey: process.env.ANTHROPIC_API_KEY!,
tools: [...],
maxSteps: 32, // raise the limit
});

Values <= 0 are ignored and the default (16) applies. There is no way to fully disable the cap; set a very large number (e.g. 10_000) if you genuinely need unbounded execution.

const res = await agent.complete('...');
if (res.finishReason === 'length') {
// check whether it is a maxSteps stop, not a token-length truncation
const report = agent.lastReport;
if (report?.reason === 'max_steps') {
console.warn('tool loop capped after', report.stepCount, 'steps');
}
}

The same reason is delivered in the onRunComplete hook payload.

A failed LLM call during a run (auth, rate limit, context overflow, provider error) throws out of agent.complete() / agent.stream() — the original LLMError (with kind/status) — matching a no-tools complete(). It never silently returns empty text. For partial results + a status, catch the error and read agent.lastReport — an AgentRunReport carrying reason, the per-step and per-tool reports, and usage:

try {
await agent.complete('…');
} catch (err) {
const report = agent.lastReport; // reason, steps, toolCalls, usage
}

Bounded stops are not errors: max_steps returns finishReason: 'length' (see above) and a model refusal returns normally — both are inspectable via finishReason / report.reason.

finishReason is an open union, so always write a default branch. Providers keep inventing terminal states, and the union is open (CONSTITUTION R1) precisely so that a new one is not a breaking change for every consumer — including consumers of providers that changed nothing. Values beyond the documented set reach you rather than being flattened: Google’s TOO_MANY_TOOL_CALLS arrives as too_many_tool_calls instead of stop, and OpenAI’s incomplete_details.reason distinguishes max_messages (a message cap, not a token cap) and steered (the turn was superseded and a successor response follows automatically) instead of reporting both as length. The raw provider value stays on response.raw.

And the answer does not depend on how the turn was fetched: the streamed and buffered paths share one table per provider. They did not, once — the stream carried a copy holding a single entry, so a streamed Google SAFETY block read as a clean stop with no content.

Recovering from a model failure (reflectAndRetry)

Section titled “Recovering from a model failure (reflectAndRetry)”

Some failures are the model’s, not the network’s: a malformed tool call, a truncated call, a hallucinated tool name. Resending the identical request would never fix them, so the network retry layer correctly leaves them alone — and the run used to end there.

A tool call whose arguments do not parse is never executed. It used to fall back to {}, which is a valid call rather than a failed one, so a stream cut at {"path": "/et ran the tool with no arguments at all. The call is now marked, answered with an error result so the history stays valid, and the step finishes as malformed_tool_call on every provider — previously only Google reported that reason, so the same truncation elsewhere looked like a clean turn and nothing below ever fired.

A turn can end between “the model asked for a tool” and “the tool answered”: an early break out of stream(), an output guardrail tripping, continueOnError: false, a pending approval, a caller who stopped. Anthropic and OpenAI both reject a history containing a tool call with no result, so the NEXT run died on send — one turn away from the cause, with an error naming neither.

Every run now repairs the history before its first request: an unanswered call gets a synthetic result saying the tool never ran, and onWarning fires with unanswered_tool_calls_repaired. The result is not a pretend success — the model can see the work did not happen. History truncation is pair-aware for the same reason: a count-based cut that would keep a result whose call it removed drops the orphan too, so keepLast is a ceiling rather than an exact count.

reflectAndRetry gives the model a bounded number of second chances, telling it what went wrong:

const agent = createAgent({
client,
tools: [...],
reflectAndRetry: { maxRetries: 2 }, // OFF unless configured — a retry costs a real request
});

On a triggering finish reason the loop injects structured guidance naming the attempt number and explicitly instructing the model not to repeat the same call (without that, models tend to re-emit identical arguments and burn the whole budget on one mistake), then retries the step. The failed assistant turn is not appended, so the model does not learn from its own broken output.

  • Default trigger: malformed_tool_call. Override with onFinishReasons — a content filter is usually a decision rather than a mistake, so it is not retried by default.
  • Failures are counted consecutively: a success clears the streak, so an agent that recovers and stumbles again much later gets a fresh budget rather than an inherited one.
  • When the budget is spent the loop throws, naming the escape hatch. Set throwIfExceeded: false to get the unusable response back instead.

This exists because finishReason now tells the truth about these turns. Google’s MALFORMED_FUNCTION_CALL was previously unmapped and read as a clean stop with no content — a failure that looked like a successful empty answer.

Codex-family models narrate before answering. Those parts are tagged phase: 'commentary' versus 'final_answer', and response.text concatenates everything — so an agent’s output used to include its own thinking-out-loud.

AgentLoop derives its answer with finalAnswerText(), which drops commentary. response.text and contentText() are unchanged, so callers who want the narration still get it:

import { finalAnswerText } from '@combycode/llm-sdk';
const answer = finalAnswerText(res.content); // commentary excluded
console.log(res.text); // everything, as before

Streaming carries the phase too, on the agent event itself — so you can separate narration from the answer live, which is when it matters for a UI:

for await (const ev of agent.stream('…')) {
if (ev.type !== 'text') continue;
if (ev.phase === 'commentary') renderThinking(ev.text);
else appendToReply(ev.text);
}

finalAnswerText() cannot do this job: it takes a finished message’s content, not deltas. Use it on the assembled message; use ev.phase on the stream.

AssistantPhase is an open union, and finalAnswerText excludes only what is explicitly 'commentary' rather than keeping only 'final_answer' — the day a provider adds a third phase, an allow-list would silently drop the answer.

Tools are indexed by function name (or builtin type), so registering two under the same key means one silently replaces the other and the model never sees it. That surfaces much later as “the model called the wrong tool”, with nothing in the logs pointing at the cause.

Collisions are now reported. Default 'warn' keeps last-write-wins (changing it would break apps that rely on a deliberate override) but emits an onWarning with code tool_name_collision naming which tool lost. toolNameCollisionPolicy: 'error' throws at construction instead, before the model is ever called.

Watching tool arguments form (tool_call_delta)

Section titled “Watching tool arguments form (tool_call_delta)”

A streamed run now forwards the model’s tool-call arguments as they arrive:

for await (const ev of agent.stream('find the Q3 report')) {
if (ev.type === 'tool_call_delta') {
// Raw JSON TEXT, usually not parseable on its own — for rendering.
process.stdout.write(ev.arguments);
}
if (ev.type === 'tool_call_start') {
// The complete, PARSED arguments. This is the event to act on.
console.log(ev.toolName, ev.arguments);
}
}

The fragments used to reach the loop and die there — accumulated into the call’s argument buffer and dropped — so a UI had no way to show a long argument list forming, and tool_call_start only fires once it is complete. They are now forwarded as well as accumulated: a second reader, not a handover, because the loop still needs the whole string to parse at the end of the call.

  • arguments is a raw JSON fragment. A partial one is usually not valid JSON, so render it rather than act on it.
  • callId is the accumulator’s id, not the event’s — several providers omit the id on later fragments, and the forwarded event has to carry something a consumer can group by.
  • step is the step the fragments belong to, so they can be correlated with the step that produced them.
  • Absent on providers that stream a call whole (most of them) and on every step that calls no tools, so a consumer that does not want them needs no change.
const agent = new AgentLoop({ client, tools, modelTimeout: 30_000 });

toolTimeout already bounded the tool half of a step. The model half was bounded only by whatever the client was configured with — so on a long run one slow step could hold the whole run open past any deadline the caller thought they had set.

Applied per step, not per run: a nine-step run with modelTimeout: 30_000 allows each step thirty seconds, not the run. A run-wide budget is a different thing and you already have it — an AbortSignal you control.

A per-call ExecuteOptions.timeout still wins, so one complete() can ask for longer. There is no new error type: the timeout surfaces as the LLMError{ kind: 'timeout' } the network layer already raises, which is what code catching timeouts already matches on. (Upstream names a ModelTimeoutError; a second class for a condition we already report would mean every consumer has to learn both.)

Checking tool arguments (validateToolArguments)

Section titled “Checking tool arguments (validateToolArguments)”
const agent = new AgentLoop({ client, tools, validateToolArguments: true });

Checks a tool call’s arguments against that tool’s own parameters schema before running it. On a failure the tool is not executed and the errors go back to the model as the tool’s result:

Invalid arguments for "lookup": $.city: expected string, got number.
Call the tool again with arguments matching its schema.

That shape is deliberate on both counts. It is a result, not an exception, because the model asked for something its own schema forbids — a thing it can fix on the next step — and ending the run would discard every step before it over a mistake the model usually corrects when told. And it says what to do: a bare validator message reads as an internal error, which models answer by apologising rather than by re-calling the tool.

The bound is maxSteps, the loop’s existing one, rather than a second retry budget to tune that would give the same answer. Each refusal emits onWarning with code tool_arguments_invalid, so a model that never gets it right is visible instead of quietly eating the step budget.

Off by default for the same honest reason as structured.validate: the bundled validator reads the common JSON Schema keywords, not all of Draft 2020-12, so on by default it would refuse calls that are valid under a schema it cannot fully read. Where a provider’s own strict mode is available that is the better guarantee — this is for the models and surfaces where it is not, and for schemas strict mode cannot express. A builtin tool ({ type: 'web_search' }) has no parameters and is left alone.

route() falls over between models for a one-shot complete(). A run is where it matters more: a rate limit on step 7 of a nine-step run threw away six steps of work and every tool call they paid for, and the only recourse was to start again.

const agent = new AgentLoop({
client: primary, // openai/gpt-5.6-sol
fallbackClients: [backup], // anthropic/claude-haiku-4.5
tools: [...],
});

Four rules, and the third is the one that cannot be got wrong:

  1. Each client is tried once per step. Retrying one client is the network engine’s job; doing it here too would retry a single failure twice over, at two layers.
  2. Only a failure another model could survive moves on — rate_limit, server_error, model_not_found, timeout, network, quota_exceeded, unsupported. An auth failure, a malformed request, a content filter or a prompt that is simply too long is the same request failing the same way everywhere, so it propagates immediately instead of being offered to every backup in turn. Override the set with fallbackOn.
  3. A streamed step stops being able to fall over at its first event. A streamed turn can fail after several chunks, and the consumer has already rendered half an answer; a backup would start a different one mid-sentence and the two would be spliced into one turn. So the error reaches the caller.
  4. Every step starts from the primary again. A rate limit is transient, and a run that fell over once should not spend the rest of its life on the backup.

When every client fails, the last one’s error is what you get: the provider’s own message is the useful half, and a wrapper saying “all models failed” buries it.

Each hand-off emits onWarning with code model_fallback and details: { from, to, kind }, per step — a silent fallback is a latency and cost change nobody can see.

What gets attributed to whom. agent.model still reports the primary: it is read before any request is made, and a caller asking what model an agent is configured with means the primary. But each step’s history entry and span name whoever actually served, and that is not cosmetic — provenance is model-bound, so a stateful continuation (previous_response_id, previous_interaction_id) is only valid against the model that issued the state. Recording the primary on a turn the backup produced would have the next step offer the backup’s server state to the primary.

Backups are supplied to AgentLoop.restore() the same way the client is — a snapshot carries neither, because both are live objects, and a restore that silently lost them would resume a run less resilient than the one it continues.

const resumed = AgentLoop.restore(snapshot, {
client: primary,
tools,
fallbackClients: [backup],
});

On a stateful API (OpenAI Responses, Google Interactions) the loop automatically continues by id (previous_response_id / previous_interaction_id) between tool rounds instead of resending the whole transcript — gated by the catalog’s retention TTL (openai/xai 30d, google 72h) and model binding. No configuration needed; it falls back to full history when the stored turn is too old or the provider isn’t stateful.

An unnamed agent exports as a bare invoke_agent carrying only its generated id, and that id changes per process — so a trace cannot say which of your agents ran, and two runs of the same agent cannot be compared. Three optional fields fix that:

const agent = createAgent({
model: 'anthropic/claude-haiku-4.5',
label: 'briefing', // names the span: `invoke_agent briefing`
source: 'customer', // which part of YOUR system this belongs to
attributes: { 'app.tenant': 'acme' }, // anything the two fixed fields do not cover
});

They change no behaviour and cost nothing when telemetry is off — they travel with the agent’s spans and nothing else. See Observability / Telemetry for how each is exported.