Skip to content

Tools -- defineTool

defineTool is the ergonomic builder for function tools. It infers TypeScript types from a compact params spec so you get typed args in execute without writing a JSON schema by hand.

  • You want to give the model a callable function (weather lookup, database query, API call, file read, etc.).
  • You want TypeScript inference on the tool’s argument types.

For built-in server-side tools (web search, code interpreter) pass them as plain objects — { type: 'web_search' } — directly in tools: [...]; no defineTool needed for those.

ExportWhat it does
defineTool(input)Build an AgentTool from a name, description, param spec, and execute function.
AgentTool (type)The shape expected by complete(), createAgent(), and delegate().
ParamSpec (type)Allowed param spec values: 'string', 'number', 'boolean', 'string[]', 'number[]', or an inline schema object.
import { complete, defineTool } from '@combycode/llm-sdk';
const getWeather = defineTool({
name: 'get_weather',
description: 'Get the current weather for a city.',
params: {
city: 'string',
unit: { type: 'string', enum: ['celsius', 'fahrenheit'] as const },
},
optional: ['unit'],
execute: ({ city, unit }) => {
// Return value is a string (or ContentPart[]) handed back to the model.
return `It is sunny in ${city} (${unit ?? 'celsius'}).`;
},
});
const { text } = await complete({
model: 'anthropic/claude-haiku-4.5',
apiKey: process.env.ANTHROPIC_API_KEY,
prompt: 'What is the weather in Paris?',
tools: [getWeather],
maxTokens: 128,
});
console.log(text);

parameters takes a plain JSON Schema or any Standard Schema (~standard) — a Zod/Valibot/ArkType schema passes straight in, so the shape is described once instead of twice. Same for outputSchema, where the schema’s OUTPUT side is used. The conversion happens once at the request boundary, and it is a protocol rather than a dependency. See Standard Schema.

Strict mode makes a provider constrain the tool name and argument shape while generating them, instead of leaving you to validate afterwards.

It is opt-in (strict: true on a FunctionTool) everywhere except the OpenAI Responses API, where it has long been the default. Measured against both providers it makes no difference to argument quality — 40 of 40 calls conformed with it and without it, including prompts written to pull away from the schema — while a schema the provider dislikes is rejected with a 400 rather than degraded. Its one real effect is that Anthropic refuses to call a tool that was never declared (10/10 undeclared without it, 0/10 with it), which only matters if something puts an undeclared tool in front of the model.

When you do ask for it, the schema must satisfy that provider’s rules, and the two constrain different things:

OpenAIAnthropic
optional properties (not in required)rejectedfine
maximum, minimum, multipleOf, maxItems, exclusive*finerejected
additionalProperties: true (an open object)rejectedrejected
{ type: 'object' } with no properties keyrejectedfine
more than 20 strict tools in one requestfinerejected
more than 24 optional parameters across all strict schemasfinerejected
”too complex to compile” (no published formula)—rejected

On the Responses API, where strict is the default, it is requested only for schemas that satisfy OpenAI’s rules — so a tool with an optional parameter simply runs without it rather than failing. strictSupport(schema, 'openai' | 'anthropic') is exported if you want to ask the question yourself; it returns { ok, reason }, and reason names the property or keyword responsible.

Anthropic’s last three rows are why strict is not defaulted on there. Two are aggregates over the whole request, so no per-schema check can see them, and the third has no published formula at all: 24 optional parameters spread over four tools compiles, the same 24 in one tool does not. Twelve ordinary tools with five optional parameters each already exceed the 24 limit. Non-strict tools count toward none of the limits.

One consequence worth knowing: a generic “router” tool can never be strict. If a parameter must accept any shape ({ type: 'object', additionalProperties: true }), that is the opposite of what strict means, and both providers refuse it.

A tool taking no arguments is unaffected: properties: {} is present but empty, which both providers accept.

Keys listed in optional are optional in the inferred execute args too, so unit above is string | undefined and the ?? 'celsius' is load-bearing. Anthropic keeps that working under strict: asked not to specify a unit it omitted the argument 10/10, asked for fahrenheit it supplied it 10/10, and never invented the second optional one.

complete() runs the full loop until the model stops requesting tools:

import { complete, defineTool } from '@combycode/llm-sdk';
const getUserCity = defineTool({
name: 'get_user_city',
description: "Get the user's current city.",
params: {},
execute: () => 'Paris',
});
const getWeather = defineTool({
name: 'get_weather',
description: 'Get the weather for a city.',
params: { city: 'string' },
execute: ({ city }) => `sunny in ${city}`,
});
const { text } = await complete({
model: process.env.LLM_MODEL!,
apiKey: process.env.LLM_API_KEY,
prompt: 'What is the weather where I am?',
tools: [getUserCity, getWeather],
maxTokens: 512,
});
console.log(text);

execute receives a second ToolExecutionContext argument with run trace ids and call metadata. Useful for logging, correlation, or accessing the agent’s conversation history.

ctx.trace carries three ids:

  • sessionId — the agent id (the ConversationHistory id, same as loop.id)
  • requestId — the run id for this specific .complete() / .stream() invocation
  • callId — this tool call’s id (same as ctx.callId)
import { defineTool } from '@combycode/llm-sdk';
import type { ToolExecutionContext } from '@combycode/llm-sdk';
const loggedTool = defineTool({
name: 'read_db',
description: 'Read a row from the database.',
params: { id: 'string' },
execute: async ({ id }, ctx: ToolExecutionContext) => {
console.log(
`Tool call ${ctx.callId} | agent ${ctx.trace?.sessionId} | run ${ctx.trace?.requestId}`,
);
return `row data for ${id}`;
},
});

Returning an image (or a PDF, or audio) from a tool

Section titled “Returning an image (or a PDF, or audio) from a tool”

execute may return ContentPart[] instead of a string, and media in it is sent to the model as media. A screenshot tool, a chart renderer, a “fetch this page as a PDF” tool: the model sees the picture rather than a description of one.

const screenshot = defineTool({
name: 'take_screenshot',
description: 'Take a screenshot of the current screen.',
params: {},
execute: async () => [
{ type: 'text', text: 'screenshot taken' },
{ type: 'image', source: { type: 'base64', mimeType: 'image/png', data: await grabPng() } },
],
});

Each API has its own place for this and they disagree about where, so the SDK splits the result into its text half and its media half and puts each where that provider takes it:

APImedia travels as
Anthropic Messagesblocks inside tool_result.content
OpenAI Responsesitems inside function_call_output.output
Google generateContentfunctionResponse.parts[].inlineData
OpenAI Chat Completionsits own user message, right after the tool results — the API has no slot
Google Interactionsits own user_input item, for the same reason

Verified live on 2026-09-30 against claude-haiku-4.5, gpt-5.4-nano (Responses and Completions), gemini-3.1-flash-lite (generateContent and Interactions) and grok-4.3: a tool returned a solid-colour square, and every model named the colour — which it can only do by decoding the image.

A string result is unchanged on every backend, so this costs nothing when a tool returns text.

Two edges worth knowing:

  • Anthropic has no tool_result block for audio or video, so those render as an [unsupported: …] note. Google’s functionResponse takes inline bytes only (fileData there is documented Vertex-only), so a URL-sourced image leaves an [image omitted: …] line in the result text. Both say so rather than dropping the part.
  • A content-part result containing only text now sends the text. It used to send the part wrapper as JSON — [{"type":"text","text":"a"}] to a model that only wanted a.

Attaching out-of-band data — customDataExtractor

Section titled “Attaching out-of-band data — customDataExtractor”

An AgentTool may declare an optional customDataExtractor(result, args, context) that runs after a successful execute. Its return value is attached to that tool call’s ToolCallReport.customData — for your own telemetry, routing, or audit. The model never sees it (it is not part of the tool result). A throwing extractor is swallowed, so this convenience can never break the tool result.

const lookup: AgentTool = {
definition: { name: 'lookup', description: 'Look up a record', parameters: { id: { type: 'string' } } },
execute: async ({ id }) => fetchRecord(id),
// model never sees this — it lands on the ToolCallReport.
customDataExtractor: (result, args, ctx) => ({ bytes: String(result).length, callId: ctx.callId }),
};

Server-side tools the provider runs are passed as plain objects in tools: [...] (no defineTool): { type: 'web_search' }, { type: 'web_fetch' }, { type: 'code_interpreter' }, { type: 'image_generation' }, { type: 'file_search' }, and { type: 'mcp' }. Provider-specific configuration goes in params, forwarded verbatim (e.g. { type: 'web_fetch', params: { allowed_domains: ['docs.example'], max_content_tokens: 4096 } }).

Programmatic tool calling (OpenAI Responses, gpt-5.6 family). Add { type: 'programmatic_tool_calling' } to let the model write JS that orchestrates your tool calls. Function tools can then declare who may invoke them and the shape they return: allowedCallers?: ('direct' | 'programmatic')[] and outputSchema? on a FunctionTool. Both are emitted only on the OpenAI Responses path (other providers ignore them). Model-gated — of the gpt-5 / o3 / o4 / codex models tested, only gpt-5.6-luna / -sol / -terra accept the builtin; the rest reject it by name.

The model’s program arrives as a program_call content part and its return value as program_result. The tool calls the program makes are ordinary tool_call parts — you execute them exactly as before — each carrying caller: { type: 'program', callerId } pointing back at the program that made it:

const res = await client.complete(messages, {
tools: [
{ type: 'programmatic_tool_calling' },
{ ...getWeather, allowedCallers: ['programmatic'] },
],
});
for (const part of res.content) {
if (part.type === 'program_call') console.log('model wrote:', part.code);
if (part.type === 'program_result') console.log('program returned:', part.result);
}
for (const call of res.toolCalls) {
console.log(call.name, call.caller?.type ?? 'direct'); // -> get_weather program
}

The program suspends at each await, so its calls still arrive one turn at a time; answer them the usual way and the program resumes.

Keep the program_call part in your history and send it back. Dropping it is not just a lost audit trail — the model re-emits the program and runs the whole thing again from the start. The adapter also re-sends the provider items the program is bound to, which the API requires.

allowedCallers is enforced locally as well as by the provider: a tool without it is direct-only, so model-written code cannot reach a tool that never opted in. A violation denies that single call (an error result to the model, plus an onWarning with code tool_caller_not_allowed) instead of ending the run.

Sources a hosted search cited are on response.citations (Citation[] — { url, title?, text? }), unified across the four ways providers report them: Anthropic on the text block (the only one that also gives the cited passage, as text), Google in groundingMetadata, OpenAI Responses and Chat as url_citation annotations, xAI as bare top-level URLs. It is distinct from builtinToolCalls, which records what the model invoked — a turn can run three searches and cite one page. Optional, so read it as response.citations ?? []; through an agent run the sources accumulate across every step and are deduped by URL. Google’s Interactions surface is not mapped yet and always reports none.

Note that Google reports each source as a vertexaisearch.cloud.google.com/grounding-api-redirect/… URL rather than the page itself — that is what the provider returns, and it redirects to the real source. The other four report the page URL directly.

Streaming reports them as they arrive. stream() yields a citation event per source, and the same sources land on the streamed final response’s citations, so the two call styles agree:

for await (const ev of llm.stream(messages, { tools: [{ type: 'web_search' }] })) {
if (ev.type === 'text') process.stdout.write(ev.text);
if (ev.type === 'citation') footnotes.push(ev.citation);
}

A citation arrives when the model cites it, which is not when the search ran — providers search early and cite while writing, so citation events interleave with text. Raw events are passed through exactly as the provider sent them, repeats included (Google resends its grounding chunks); deduplication by URL happens where the final response is assembled, so a consumer rendering live footnotes still sees everything that arrived.

Files a hosted tool produces (e.g. code-execution charts or data files) are surfaced uniformly on response.files (FileOutput[] — { id?, name?, mimeType?, data?, url?, ref?, source? }), independent of generated media. You don’t fetch per-provider — retrieveFile(file) / streamFile(file) resolve every shape (id via the provider’s files API, inline base64 data from Google/xAI, or a url). See the Code execution guide and Retrieving output files.

Which models support which builtin is in the catalog: capabilities.builtinTools, catalog.supportsBuiltinTool(provider, model, tool), or select('code_interpreter'). Coverage: web_search on all providers; code_interpreter on all except OpenRouter; web_fetch on Anthropic (web_fetch_20260318) and Google (urlContext) — OpenAI’s web_search already opens pages, and xAI / OpenRouter expose no separate fetch tool.

Seeing what ran. Provider-run builtins surface a durable trail on response.builtinToolCalls and, while streaming, { type: 'builtin_tool_start' } / { type: 'builtin_tool_end' } events as each runs. Each entry carries what the tool ran:

interface BuiltinToolCall {
tool: string; // 'web_search' | 'web_fetch' | 'code_interpreter' | 'shell'
id?: string;
callId?: string; // shell: the model's call id, to address an answer to
environment?: string; // shell: 'local' (it is asking YOU) or 'container_reference'
code?: string; // code_interpreter: the code the model executed; shell: the commands
output?: string; // code_interpreter: the code's stdout / logs
query?: string; // web_search: the query the model searched for
url?: string; // web_search: a page opened/read; web_fetch: the URL fetched
sources?: string[]; // web_search: the URLs the search drew on
results?: Array<{ // web_search: the results, when asked for
imageUrl?: string;
sourceWebsiteUrl?: string;
thumbnailUrl?: string;
caption?: string;
[key: string]: unknown; // whatever else the provider sent
}>;
}

The payload is normalized across providers and present on both complete() and streamed responses (the builtin_tool_end event carries the same fields). These are informational — unlike tool_call_* (a function call the client must execute), the provider runs these itself. Use them to show a ”🔎 Searching: ” / “⚙️ Running code” panel with the actual code and output.

Progress while one runs. A streamed turn also emits { type: 'builtin_tool_delta' } between the start and the end, carrying code (a fragment of what the model is about to run) or output (a fragment of what it printed). Both are fragments to append, like text; the complete values arrive again on builtin_tool_end, so a consumer that only wants the result can ignore them. Currently emitted by OpenAI’s shell tool — the only hosted tool that reports progress rather than just a result.

Shell commands (shell) — the builtin that may not run at all

Section titled “Shell commands (shell) — the builtin that may not run at all”

shell is the exception to everything above: it is only provider-run when you give it a container.

params.environmentWhat happens
{ type: 'container_auto' }OpenAI provisions a container, runs the commands, and streams stdout/stderr back. A normal hosted tool call. The response reports environment: 'container_reference' — container_auto was the request, not the answer.
{ type: 'local' }, or omittedThe model only asks. Nothing runs, the turn ends, and it is waiting on you.
const res = await llm.complete('Run `echo one` and summarise.', {
tools: [{ type: 'shell', params: { environment: { type: 'container_auto' } } }],
});
// res.builtinToolCalls[0] = { tool: 'shell', code: 'echo one',
// output: 'one
', environment: 'container_reference', callId: '…' }

A local call ends the turn with finishReason: 'stop' and text: ''. That looks exactly like a model with nothing to say, so the commands are reported on builtinToolCalls[].code and an onWarning with code shell_awaiting_caller says what happened and how to answer it:

engine.hooks.on('onWarning', (w) => {
if (w.code === 'shell_awaiting_caller') console.warn(w.message);
});

To answer one, run the commands yourself and send the result back addressed to the call’s callId.

Two wire details worth knowing, because they shape what you receive: a container run arrives as two output items (the commands, then the output, linked by call_id) and is reported as one tool call with one start and one end; and the per-command exit codes are not folded into output (that would mean inventing text inside program output) — output is stdout and stderr in the order the commands wrote them.

xAI supports shell too, but narrower: environment is required there (a 422 names the missing field) and only local is accepted, so a shell call on xAI is always a request for you to run something.

Reusing Anthropic’s code container, and loading skills (providerOptions.container)

Section titled “Reusing Anthropic’s code container, and loading skills (providerOptions.container)”

Anthropic’s code execution runs in a container. By default you get a cold one per turn and never see it. providerOptions.container makes it yours to reuse, and lets you load skills into it. GA — measured 2026-10-02 with no beta header.

const first = await llm.complete('Use python to compute 2+2.', {
tools: [{ type: 'code_interpreter' }],
providerOptions: {
container: { skills: [{ type: 'anthropic', skillId: 'xlsx', version: 'latest' }] },
},
});
// first.container = { id: 'container_…', expiresAt: '2026-10-02T12:55:18Z',
// skills: [{ type: 'anthropic', skillId: 'xlsx', version: '20260914' }] }
// Reuse it while it lives (about five minutes):
await llm.complete('Now compute 3+3.', {
tools: [{ type: 'code_interpreter' }],
providerOptions: { container: { id: first.container!.id } },
});

Four things worth knowing:

  • The version you get back is the one that RAN. Asking for 'latest' returns '20260914', so record response.container.skills, not what you requested.
  • Skills are validated by the provider. An unknown skill is a 400 Unknown Anthropic skill, not a silent ignore — a typo fails loudly. GET /v1/skills lists the built-in ones (xlsx, pptx, pdf, …) and is GA as well.
  • A malformed skill ref is refused before the request leaves, with the field named. Sending it would make the provider complain about a key you never wrote.
  • No container means no code ran. A turn that executes nothing reports none, because none was created — that is the normal answer, not a problem.

When streaming, the container arrives on the done event and on the streamed response, so you keep the id either way. (It rides the terminal frame because that is where the provider sends it: message_start.container is null on every streamed turn.)

Generating an image mid-conversation (image_generation)

Section titled “Generating an image mid-conversation (image_generation)”
const res = await llm.complete('Draw a red circle on white.', {
tools: [{ type: 'image_generation', params: { action: 'generate' } }],
});
const image = res.media[0]; // { type: 'image_output', mimeType, ... }

Supported on openai and xai — measured live on 2026-10-01 through this library (gpt-5.6-sol, gpt-5.4-nano, grok-4.6, grok-4.5, grok-4.3 all returned an image on response.media). catalog.supportsBuiltinTool(provider, model, 'image_generation') now says so; it used to answer false for both providers, including OpenAI where the tool has always worked, so a caller gating on the catalog refused itself.

params is a verbatim passthrough spread beside type, with ImageGenerationToolParams for editor help:

  • xAI takes action: 'auto' | 'generate' | 'edit'. It is validated, not inert — action: 'paint' is a 400 naming the three. edit with nothing to edit returns no image at all: a 200 with a text-only answer.
  • OpenAI takes output_format, quality, size, background.

On mimeType. OpenAI reports output_format and reports it accurately (ask for jpeg, get JPEG bytes). xAI reports none and returns JPEG. So the parser prefers the declared format, falls back to the image’s own magic bytes, and only then to PNG — an xAI image is labeled image/jpeg rather than mislabeled image/png. That matters past cosmetics: a file written from a wrong mimeType gets the wrong extension, and a strict validator downstream (Google Veo compares the declared mime against the bytes) answers 400. A provider that does declare a format is believed even if the bytes disagree — second-guessing it would only move the question to which of two wrong answers to trust.

Two halves, and only one of them is a parameter you set:

import type { WebSearchToolParams } from '@combycode/llm-sdk';
const search: { type: 'web_search'; params: WebSearchToolParams } = {
type: 'web_search',
params: {
search_content_types: ['image', 'text'],
image_settings: { max_results: 3, caption: true },
},
};
const { response } = await complete({ model: 'openai/gpt-5.4-nano', apiKey, prompt: '…', tools: [search] });
for (const call of response.builtinToolCalls ?? []) {
for (const r of call.results ?? []) console.log(r.imageUrl, r.caption);
}

The other half is include: ['web_search_call.results'] on the request, and the adapter adds it for you whenever search_content_types contains 'image'. That matters because the results are otherwise simply absent: measured 2026-09-30, the same request with the include returns results and without it returns none — a search that found images and a response that does not contain them, which reads as “no images found” rather than as a missing parameter.

external_web_access: false runs the search cache-only, fetching no new external content.

Restricting what web_fetch may fetch (Anthropic)

Section titled “Restricting what web_fetch may fetch (Anthropic)”

url_sources decides which URLs are eligible — use the exported WebFetchToolParams for editor help. Each key is a tagged variant: user_input is all or none; the two tool filters add only and except, whose tool_names must name tools declared in the same request.

import type { WebFetchToolParams } from '@combycode/llm-sdk';
const fetchTool: { type: 'web_fetch'; params: WebFetchToolParams } = {
type: 'web_fetch',
params: {
url_sources: {
user_input: { type: 'all' }, // URLs the user pasted
server_tool_results: { type: 'only', tool_names: ['web_search'] },
client_tool_results: { type: 'none' }, // nothing from your own tools
},
},
};

Worth setting deliberately: left unset, the fetchable set is whatever the server defaults to, and “any URL any tool result mentioned” is a wider reach than most callers intend.

OpenAI’s hosted MCP tool lets the model call a remote MCP server that OpenAI connects to. Identify the server with exactly one of three targets (use the exported McpToolParams type for editor help):

import type { McpToolParams } from '@combycode/llm-sdk';
// 1. Public server — OpenAI dials the URL directly.
{ type: 'mcp', params: { server_label: 'docs', server_url: 'https://mcp.example/sse' } }
// 2. Managed connector (Gmail, Drive, …) — needs `authorization`.
// DEPRECATED by OpenAI for models released after 1 September 2026; see below.
{ type: 'mcp', params: {
server_label: 'gmail',
connector_id: 'connector_gmail',
authorization: oauthToken,
} }
// 3. Secure MCP Tunnel — reach a private/local server (behind NAT/firewall, no
// public URL) through an outbound tunnel registered under a tunnel id.
{ type: 'mcp', params: { server_label: 'local', tunnel_id: `tunnel_${id}` } }

Optional params: authorization, headers, require_approval, allowed_tools, server_description. OpenAI enforces the “exactly one target” rule server-side — measured 2026-10-01, every pairing is refused by name (Mutually exclusive parameters: 'tools[0]'. Ensure you are only providing one of: 'server_url' or 'connector_id', and likewise for each other pair).

connector_id is deprecated, and still works. OpenAI documents it as deprecated for models released after 1 September 2026, in favour of server_url or tunnel_id. It is still sent and still honoured: measured 2026-10-01 on gpt-5.6-sol, connector_id with authorization answers 200. (Without authorization it answers Must specify 'authorization' parameter with 'connector_id' — that is the field’s own requirement, not the deprecation biting.) Nothing is removed here, because a field a provider still honours is not ours to withdraw; prefer server_url or tunnel_id for new code.

tunnel_id is pattern-validated: ^tunnel_[a-z0-9]{32}$. A malformed one is refused with that pattern quoted, and a well-formed one reaches the point of dialling the tunnel — so it is a working target, not a typed-only field.

This is the provider-hosted MCP path. For connecting the SDK itself to MCP servers as a client, see MCP (Model Context Protocol).

A large tool block is paid for on every request. lazy: true registers a tool without declaring it: the model finds it with a built-in tool_search and runs it through call_tool. Measured over 308 tools, that is identical correctness at -72% / -97% cost per task, for one extra round trip — and more expensive than declaring everything below roughly a hundred tools.

import { connectMcp, defineTool } from '@combycode/llm-sdk';
await connectMcp({ url: 'https://mcp.deepwiki.com/mcp' }, { lazy: true }); // whole server
defineTool({ name: 'rare_thing', description: 'Rarely needed.', params: {}, lazy: true, execute: () => 'ok' });

Full guide, including when NOT to use it: Lazy tools.