LLM Client -- complete / stream
The LLM layer is the core of the SDK. It provides a single, normalized API to
every provider. complete() is the one-shot helper for most use cases;
createLLM() gives you a reusable client for streaming, multi-turn conversations,
and fine-grained control.
When to reach for this
Section titled “When to reach for this”- You want to send a prompt and get text back (use
complete()). - You need a streaming reply (use
createLLM().stream()). - You are managing a multi-turn conversation with explicit message arrays.
- You want server-state round-trips (OpenAI/xAI Responses API — state held on the server side so only the new turn is sent each round).
Main exports
Section titled “Main exports”| Export | What it does |
|---|---|
complete(opts) | One-shot helper. Sends a prompt, runs the tool loop if tools are supplied, returns { text, response, parsed?, retrieveFile, streamFile }. The fastest path for most tasks. |
createLLM(opts) | Builds a reusable LLMClient bound to one provider/model. |
LLMClient | Low-level client class with .complete(), .stream(), .retrieveFile(), .streamFile(), .assistantMessage(), .destroy(). |
select(query) | Pick the best matching model from the catalog by capability query ('type:chat; vision; cheap'). Returns a provider/slug string. |
selectModels(query) | Same query syntax as select, but returns the full ranked ModelInfo[] list instead of just the first provider/slug string. |
listModels() | Return the curated catalog (pricing + capabilities). |
listModelsLive(opts) | Live-discovery fetch of model ids from the provider API. |
route(opts) | Send to a primary model with client-side (or OpenRouter native) fallback. |
Type-only exports: CompleteOptions, CompleteResult, Message, ContentPart,
Role, CompletionResponse, Usage, FinishReason, StreamEvent, NormalizedRequest,
RetrievedFile, FileStream.
Hosted-tool output files (code-execution charts/CSVs) surface on
response.files; fetch their bytes withretrieveFile/streamFile— see Retrieving output files.
Provider adapter exports: AnthropicAdapter, OpenAIResponsesAdapter,
GoogleAdapter, XAIAdapter, OpenRouterAdapter, and their batch/file/media
variants (used when building custom wiring; most users never touch these).
Minimal examples
Section titled “Minimal examples”One-shot completion
Section titled “One-shot completion”import { complete } from '@combycode/llm-sdk';
const { text } = await complete({ model: 'anthropic/claude-haiku-4.5', apiKey: process.env.ANTHROPIC_API_KEY, prompt: 'Say hello in one word.',});console.log(text);Streaming
Section titled “Streaming”import { createLLM } from '@combycode/llm-sdk';
const llm = createLLM({ model: 'openai/gpt-5.4-nano', apiKey: process.env.OPENAI_API_KEY,});
for await (const ev of llm.stream('Count to 5.')) { if (ev.type === 'text') process.stdout.write(ev.text);}Structured output (typed error + opt-in repair)
Section titled “Structured output (typed error + opt-in repair)”structuredComplete(input, schema, options) returns the parsed object typed as T. If the model’s
final output can’t be parsed it throws a typed InvalidFinalOutputError (extends AgentRunError,
carries reason: 'invalid_final_output' and the model’s rawText) — not a bare SyntaxError — so you
can differentiate and inspect. Pass structured.repairAttempts to have it re-prompt with the parse
error before giving up (default 0).
import { createLLM, InvalidFinalOutputError } from '@combycode/llm-sdk';
const llm = createLLM({ model: 'openai/gpt-5.4-nano', apiKey: process.env.OPENAI_API_KEY });const schema = { type: 'object', properties: { city: { type: 'string' }, tempC: { type: 'number' } } };
try { const weather = await llm.structuredComplete<{ city: string; tempC: number }>( 'Weather in Paris as JSON.', schema, { structured: { schema, repairAttempts: 1 } }, // retry once on a parse failure ); console.log(weather.city, weather.tempC);} catch (e) { if (e instanceof InvalidFinalOutputError) console.error('bad output:', e.rawText);}Checking the result yourself — structured.validate
Section titled “Checking the result yourself — structured.validate”The provider enforced the schema, so this is off by default. Turn it on where that
enforcement is weaker than the schema: a surface with no strict mode, a model that
ignores the schema under load, a required the provider treats as advisory.
await llm.structuredComplete('Weather in Paris as JSON.', schema, { structured: { schema, validate: true, repairAttempts: 2 },});A validation failure raises the same InvalidFinalOutputError as a parse failure,
so one repairAttempts budget covers both — and the re-prompt carries the
errors, which is what makes the retry better than a re-roll. A value that parsed
and was wrong is precisely the case re-prompting helps with; a separate error type
would have left the budget covering malformed JSON and not that.
It is opt-in for an honest reason, not caution: the bundled validator covers the
common JSON Schema keywords and not all of Draft 2020-12 (no allOf/anyOf, no
formats). On by default it would reject values that are valid under a schema it
cannot fully read, and disagree with the provider that had just enforced it. Every
error is reported, each with its path, because one error per round trip is a round
trip per mistake.
A Standard Schema validates through its own validate whether or not this flag
is set — it carries refinements the provider never saw, so there is nothing to
trust it with.
Standard Schema — pass the schema you already have
Section titled “Standard Schema — pass the schema you already have”Anywhere this library takes a JSON Schema it also takes a Standard Schema: any
object carrying the ~standard property, which Zod, Valibot, ArkType, Effect
Schema and others all expose. That means a tool’s parameters, a tool’s
outputSchema, and structured.schema / the schema argument of
structuredComplete.
import { createLLM, defineTool } from '@combycode/llm-sdk';import { z } from 'zod'; // or valibot, arktype, effect/Schema, ...
const llm = createLLM({ model: 'openai/gpt-5.4-nano', apiKey: process.env.OPENAI_API_KEY });const Weather = z.object({ city: z.string(), tempC: z.number().min(-90).max(60) });
// As a structured-output schema -- `tempC` is range-checked here, which no JSON// Schema the provider saw could have told it to do.const weather = await llm.structuredComplete('Weather in Paris as JSON.', Weather);
// ...and as a tool's parameters.const lookup = defineTool({ name: 'lookup', description: 'Look up a city', parameters: z.object({ city: z.string() }), execute: async ({ city }: { city: string }) => `sunny in ${city}`,});It is a protocol, not a dependency: the types are declared structurally and nothing is installed, so the library stays zero-dependency and a schema library that does not exist yet already works.
Two things happen, and the second is the one worth knowing about:
- Conversion. The schema is converted to JSON Schema once, at the request
boundary, via its own
~standard.jsonSchema. Everything downstream — the wire specs, the provider adapters, snapshots — sees plain JSON Schema and never learns the protocol exists. - Validation. For a Standard Schema the parsed result is also run through
~standard.validate, and the value it returns is what you get. A schema carries semantics JSON Schema cannot express (refinements, branded types, cross-field rules), so the provider never enforced them — they are checked here or nowhere. Andvalidatemay transform (coercions, defaults): handing back the parsed object instead would return something that looks right and skipped the schema’s work.
A validation failure throws the same InvalidFinalOutputError as a parse failure,
on purpose: both mean “the output did not match the requested schema”, so
structured.repairAttempts re-prompts for a value that parsed and was wrong —
which is the case a repair actually helps with.
Two things are refused rather than worked around:
- A Standard Schema with no
~standard.jsonSchema(an older schema library). There is nothing to put on the wire, and sending the request without a schema would leave the model unconstrained while you believed it was constrained — the failure you are least likely to notice, because the answer usually looks about right anyway. Use the library’s Standard JSON Schema adapter, or pass a plain JSON Schema. - An asynchronous
validate. Schemas are applied while parsing a response, in synchronous code; a Promise is a truthy object with noissues, so awaited nowhere it would have passed as a valid result and been handed back in place of your data.
A plain JSON Schema behaves exactly as before, including not being re-validated locally: the provider already enforced it, and a second check with a zero-dep validator would mostly surface places where we and the provider disagree.
isStandardSchema, isStandardSchemaWithJson, toJsonSchema and
validateStandardSchema are exported for callers building their own layer on top.
Finish reasons — and the two non-obvious ones
Section titled “Finish reasons — and the two non-obvious ones”response.finishReason is unified across providers: 'stop' | 'tool_use' | 'length' | 'content_filter' | 'error' | 'pending'.
'pending'— not terminal. The provider accepted the request but has not produced a completion, so the response carries no content. It comes from Google Interactionsqueuedand OpenAI Responsesqueued/in_progress(background mode). Treat it as “poll/retry”, never as a result. Before 1.8.0 these fell through to'stop', which reported a clean finish for an empty response.'error'— the provider reported a failure inside a 200 response (OpenAI Responsesstatus: 'failed', Google Interactionsstatus: 'failed'), so there is no exception to catch. When set,response.errorcarries{ code?, message?, misalignment? }— e.g. OpenAI’sdata_residency_mismatch. A code the provider sends as a number is read as its decimal string, so a numeric code is a code rather than an absent one.
const { response } = await complete({ model: 'openai/gpt-5.4-nano', apiKey, prompt: '…' });if (response.finishReason === 'pending') { // nothing ran yet — poll again, do not treat response.text as an answer} else if (response.finishReason === 'error') { console.error(response.error?.code, response.error?.message);}A safety block that explains itself
Section titled “A safety block that explains itself”OpenAI’s misalignment_policy_violation (2026-09) comes with error.misalignment, and it is the
one error body worth reading past message:
| Field | What it is |
|---|---|
detailedExplanation | why this particular turn looked wrong |
errorType | a classification — potentially_unintended_data_transfer, …_data_access, …_destructive_activity, other. Open: the provider says clients must accept more, so it is typed as a string |
steer.message | a continuation the provider suggests sending instead |
steer is the part that changes what an agent can do. Without it the only thing the run learns is
that it was stopped:
if (response.error?.misalignment) { const { detailedExplanation, steer } = response.error.misalignment; console.warn(`blocked: ${detailedExplanation}`); if (steer) { // A path forward the provider itself offered — worth surfacing to the user // before deciding whether to retry. console.warn(`suggested: ${steer.message}`); }}Absent unless the provider sent one, and never an empty object: misalignment being present means
a safety system actually explained itself.
Anthropic’s
refusalstop reason maps to'content_filter'(a safety decline is a block, not a clean finish), andmodel_context_window_exceededmaps to'length'.
Sampling parameters
Section titled “Sampling parameters”No sampling parameter is universal — not even temperature. The SDK emits each one only where
the provider actually accepts it, because sending one blindly is a hard 400, not a no-op:
| Option | Honoured by | Dropped for |
|---|---|---|
temperature / topP | Everywhere except Anthropic from the Opus 4.8 generation onward | Anthropic models on wire era messages@4.7 — claude-opus-4.8 and every Claude 5.x — which reject them (400 `temperature` is deprecated for this model). Measured 2026-10-01; claude-sonnet-4.6 and claude-haiku-4.5 still accept them |
topK | Anthropic, on models up to Opus 4.6 — behaviourally verified. Also sent to Google + xAI, which accept it but showed no effect when measured | OpenAI (no top-k); Anthropic models after Opus 4.6, which reject it (400 top_k is deprecated) |
seed | OpenAI chat-completions, Google (both surfaces), xAI (chat + responses), OpenRouter chat | Anthropic, OpenAI Responses (both reject it) |
presencePenalty / frequencyPenalty ([-2, 2]) | OpenAI/xAI chat-completions, OpenRouter, Google (generateContent + Interactions) | OpenAI/xAI Responses, Anthropic |
stop | Anthropic, Google, xAI, OpenAI chat | OpenAI Responses |
You pass them the same way regardless; where a provider can’t take one it is left out of the request rather than forwarded and rejected.
Anthropic refuses temperature and topP in the same request on the models that take either
(400 `temperature` and `top_p` cannot both be specified for this model. Please use only one.).
Set both and topP is dropped, temperature is sent, and you get an onWarning saying so — a
request that works beats one that fails, and the alternative is a 400 for a combination that is
perfectly ordinary elsewhere.
Every drop is reported. Each one reaches you as onWarning with code request_adjusted, naming
what was left out and why. Silence would be worse than the 400 it replaces: a caller who sets
temperature: 0 and gets default sampling has no way to find out, and topK behaved exactly that
way between the Opus 4.7 release and this fix.
Accepted is not the same as honoured. A
200only proves the field was not rejected. We testedtopKbehaviourally (top_k: 1must force greedy decoding): only Anthropic actually applies it — Google and xAI accept it and ignore it on the models we measured.seedis best-effort everywhere that takes it; determinism is never guaranteed.
await complete({ model: 'google/gemini-2.5-flash', apiKey, prompt: '…', topK: 40, seed: 42 });await complete({ model: 'openai/gpt-5.4-nano', apiKey, prompt: '…', presencePenalty: 0.6, frequencyPenalty: 0.3 });Reasoning (thinking)
Section titled “Reasoning (thinking)”thinking turns on a model’s reasoning and maps to each provider’s own control:
-
mode: 'auto' | 'on' | 'off'— enable/disable reasoning. -
mode: 'between_tools'— reason only BETWEEN tool calls. Anthropic-only and model-gated; see below. -
effort: 'low' | 'medium' | 'high' | 'max' | 'xhigh'— intensity, mapped per provider, never passed through: Anthropicbudget_tokensbelow 4.6 andoutput_config.efforton 4.6+, OpenAI and xAIreasoning.efforton Responses andreasoning_efforton chat-completions, GooglethinkingBudgeton 2.5 /thinkingLevelon 3.x.maxmeans the most this model will do, so it lands on the top rung of each provider’s own ladder —xhighon OpenAI and xAI,highon Google, whose ladder ends there. Namexhighdirectly when you want that rung specifically rather than “whatever the maximum is”; a model that does not take it answers 400 naming the value, which beats being quietly served a different amount of thinking than you asked for.Measured 2026-09-30: xAI honours the effort on the grok-4.3 line and up on both surfaces (grok-4.6 Responses low 449 → xhigh 3066 reasoning tokens, chat-completions low 927 → xhigh 7043), while the whole grok-4.20 line answers
400 "does not support parameter reasoningEffort"and the SDK therefore omits the field for it.grok-4.20-multi-agentaccepts it but reads it as an agent count, so it is not treated as an effort control. -
visibility: 'full' (default) | 'summary' | 'hidden'— how much reasoning comes back: Anthropicenabled.display, OpenAI Responsessummary, GoogleincludeThoughts. Best-effort — a provider without a middle state degradessummarytofull. -
context: 'auto' | 'current_turn' | 'all_turns'— cross-turn reasoning persistence (OpenAI Responses). Omitted, the model decides: thegpt-5.6family defaults toall_turns, earlier models tocurrent_turn.
Anthropic has two incompatible request shapes and the SDK picks per model — you do not configure
this. Claude 4.6 and later take thinking: {type:'adaptive'} and reject budget_tokens with a 400;
everything below 4.6 has no adaptive mode and requires the budget. An unrecognised model id gets
adaptive, since that is the shape Anthropic is moving to.
await complete({ model: 'anthropic/claude-haiku-4.5', apiKey, prompt: '…', thinking: { mode: 'auto', effort: 'high', visibility: 'hidden' } });between_tools is accepted by almost nothing, and the library checks before sending. Measured
2026-09-29 against every active Anthropic chat model: exactly one takes it — claude-sonnet-5.5 —
and the other twelve answer 400 "thinking.type.between_tools" is not supported for this model,
claude-opus-5.5 included. A deliberately invalid thinking type is refused everywhere, so the field
is read rather than tolerated.
Ask for it on a model the catalog does not record as accepting it and the mode is dropped, the
request goes out with the model’s ordinary reasoning, and you get an onWarning with code
request_adjusted naming the model that does take it. That is what Anthropic’s own fallback
middleware does with this value when it hops to another model — a request that works beats a 400.
Note this gate is the mirror of reasoning.canDisable: that one stops a request only on an explicit
false, because almost every model can disable reasoning. This one sends only on an explicit
true, because almost none accepts it. The cost is that a newly-released model that takes it needs
a catalog entry before callers can use it — and until then they get a warning, not silence.
(OpenAI’s Responses-only execution mode standard/pro is providerOptions.reasoningMode — see below.)
Changing the effort for the REST of a stored conversation
Section titled “Changing the effort for the REST of a stored conversation”thinking.effort applies to the request it is on. On a conversation the server is
holding there is a second thing you may want: from here on, think this hard. That
is a configuration_update content part.
await llm.complete( [ { role: 'user', content: [ { type: 'configuration_update', reasoning: { effort: 'none' } }, { type: 'text', text: 'Just give me the number.' }, ], }, ], { providerOptions: { openai: { conversation: conversationId } } },);// Every later turn on this conversation inherits `effort: 'none'` until another// update replaces it.The part is emitted as its own top-level item, before the message it travels with: the API applies an update to subsequent responses, so one placed after the message it was meant to govern governs the next one instead — a change that takes effect a turn late, with nothing reporting it.
Why this is not the same as the option. Measured on gpt-5.6-luna on
2026-10-01, three runs per arm, setting the effort in turn 1 and naming nothing in
turn 2:
| turn 1 set the effort via | turn 2 reasoning tokens |
|---|---|
configuration_update | 0, 0, 0 |
thinking.effort | 244, 189, 172 |
| nothing at all | 155, 129, 198 |
The option does not persist and the item does. There is no other way to say it.
Support is narrow. gpt-5.6-sol and gpt-5.6-luna accept the item;
gpt-5.4, gpt-5.4-nano and gpt-5.5 answer
400 The 'configuration_update' item type is not supported with this model. Every
non-OpenAI provider ignores the part.
effort takes the unified ladder plus none and minimal, which OpenAI’s item
accepts and the unified ladder does not yet carry — none being the value that
demonstrates the feature at all. max is mapped to xhigh, the same as
everywhere else, so the word means one thing across the whole surface. minimal
is model-dependent even within OpenAI: gpt-5.6-luna takes it, gpt-5.6-sol
answers 400 naming the values it does take.
One shape note, because the official SDK’s types disagree: reasoning and
reasoning.effort are both required, and effort: null is refused. Measured;
openai-ts types all three as optional or nullable.
Provider-specific options (providerOptions)
Section titled “Provider-specific options (providerOptions)”providerOptions is a passthrough for provider features that have no unified equivalent. Each adapter
reads the keys it understands and ignores the rest:
- Anthropic —
userProfileId→ theanthropic-user-profile-idheader (identifies the end user a request acts on behalf of; needs the account-leveluser-profilesbeta). - Anthropic —
workspaceId→ theanthropic-workspace-idheader (selects the Workspace, e.g.wrkspc_011CZ…). See Workspaces below. - Google generateContent —
translationConfig→generationConfig.translationConfig({ targetLanguageCode }; Gemini Developer API). - Google generateContent —
cachedContent→ top-levelcachedContent, an explicit context-cache resource (cachedContents/…). Moved off Interactions in 1.8.0: google 2.13 removedcached_contentfrom the Interactions request model and that endpoint now rejects it outright (400 Unknown parameter 'cached_content'), so sending it there was a hard failure. It remains valid ongenerateContent, which is where the passthrough now lives. - OpenAI Responses —
reasoningMode: 'standard' | 'pro'→reasoning.mode(chat-completions rejects it, so it’s not a unifiedthinkingknob). - OpenAI (Responses + chat) —
moderationPolicy→moderation.policy({ input?: { mode: 'score'|'block' }, output?: {…} }) for server-side moderation blocking. The unifiedmoderationoption stays report-only; use this (ormoderationGuardrailat the agent layer) to block. - OpenAI (Responses + chat, gpt-5.6+) —
promptCacheOptions→prompt_cache_options(typed asPromptCacheOptions:{ mode?: 'implicit'|'explicit', ttl?: '30m', prewarm?: boolean }). Note: OpenAI caches implicitly by default, so the unifiedcacheconfig already caches on OpenAI with no config — this is for manual control only.gpt-5.6+is not advisory: an older model refuses the whole object with400 prompt_cache_options is not supported on this model. See Warming the cache before you need it.
await complete({ model: 'anthropic/claude-haiku-4.5', apiKey, prompt: '…', providerOptions: { userProfileId: 'usr_42' } });How a model reads a video (Google)
Section titled “How a model reads a video (Google)”Two ways, and on a long video the difference is cost, not style:
processing | What happens |
|---|---|
'agentic' | the model navigates the video itself, seeking to what it needs |
'static' | a fixed frame rate, every extracted frame placed in the context window |
{ type: 'static', fps, startOffset, endOffset } | static, with the sampling spelled out |
The object form is the one worth reaching for: fps trades detail against tokens, and the offsets
are how a question about 30 seconds of a two-hour recording costs what 30 seconds should. Offsets
are seconds with an s suffix, as Google writes them.
await complete({ model: 'google/gemini-3.1-flash-lite', apiKey, messages: [ { role: 'user', content: [ { type: 'video', source: { type: 'url', url: 'https://www.youtube.com/watch?v=…' }, providerOptions: { processing: { type: 'static', fps: 1, startOffset: '5s', endOffset: '20s' } }, }, { type: 'text', text: 'What happens in this clip?' }, ], }, ],});The two Google surfaces take different shapes, and one cannot express the sampling.
generateContent has Part.mediaProcessing, an enum with exactly two values — so the object form
is honoured on Interactions and reduced to plain STATIC there. A window of a long video is a
request only Interactions can carry.
Measured 2026-09-30.
mediaProcessingis refused unless the same part carries a video mime type (400 mime_type must be set when media_processing is specified, and with a genericapplication/octet-streamit is400 media_processing can only be set on video parts). Oururlandfilesources carry no mime type at all, so one is supplied for a video that asked for processing — and only then, leaving every request that did not ask byte-identical.
'agentic'is gated per model: ongemini-3.1-flash-liteit is400 Agentic video processing is not enabled for this model. That gate is not encoded here — a hard-coded model list would go stale, and the provider’s own message already says exactly what is wrong.
Images Anthropic would otherwise shrink without telling you
Section titled “Images Anthropic would otherwise shrink without telling you”An image larger than the model’s maximum is downsized by default, silently. The model reasons over dimensions you did not choose, the answer comes back looking normal, and nothing in the response says the detail you were asking about was resampled away.
Measured 2026-09-30 — a 4000x4000 image sent to claude-haiku-4.5:
image dimensions 4000x4000 exceed the maximum image size of a model named on this request and would be downsized to 1092x1092; scale the image to at most 1092x1092 or set the image’s
oversized_imagesetting to"downsize"
1092x1092 is 7% of the pixels that were sent. For a screenshot of small text, or a scan someone is asking you to read, that is the difference between an answer and a guess.
oversized_image: 'error' turns the silent shrink into that refusal, per image:
await complete({ model: 'anthropic/claude-haiku-4.5', apiKey, messages: [ { role: 'user', content: [ { type: 'image', source: { type: 'base64', mimeType: 'image/png', data }, // Refuse rather than resample. The 400 names the dimensions and the // largest that would fit, so you can scale it deliberately. providerOptions: { transformations: { oversized_image: 'error' } }, }, { type: 'text', text: 'What does the error message in this screenshot say?' }, ], }, ],});Per image, not per request: one oversized screenshot in a long conversation should not change how every other image in it is handled. Omitted entirely when unset, so the server default stands.
A separate, higher limit exists above this one: a dimension over 8000px is refused outright whatever
oversized_imagesays.
Warming the cache before you need it
Section titled “Warming the cache before you need it”prewarm: true writes the prompt cache and generates nothing — it overrides generate to
false. Use it when a long shared prefix is about to be hit by several requests and you would rather
pay the cache write once, up front, than make the first real request slow.
const PREFIX = '…a few thousand tokens of shared context…';
// 1. warm it. Comes back completed and empty — that is success, not a failure.const warm = await complete({ model: 'openai/gpt-5.6-terra', apiKey, prompt: PREFIX, providerOptions: { promptCacheOptions: { prewarm: true, ttl: '30m' } },});console.log(warm.response.finishReason); // 'stop'console.log(warm.response.text); // ''
// 2. the real requests read itconst answer = await complete({ model: 'openai/gpt-5.6-terra', apiKey, prompt: `${PREFIX}
Q: …` });console.log(answer.response.usage.cachedTokens); // most of the prefixA prewarm response has an empty output[], which this library reports as an ordinary empty result
— finishReason: 'stop', no content, no error. Do not read the blank text as a broken request.
It is not free: the prewarm pays for the input tokens it writes (usage.cacheWriteTokens), so it
is worth it only when the prefix will actually be reused.
Measured 2026-09-30 on
gpt-5.6-terra: the prewarm call returned 0 output items and 0 cached tokens, and the next call on the same 4177-token prompt read 4174 of them from cache. The prompt carried a per-run nonce, so that hit can only have come from the prewarm. Note also that caching needs a prompt long enough to qualify — an earlier 613-token attempt cached nothing.
Workspaces
Section titled “Workspaces”Anthropic accounts spend, rate limits and retention against a Workspace, and
anthropic-workspace-id is what selects one. A credential scoped to a single Workspace may omit
it. A credential that can act on several and omits it does not fail — it charges the default
Workspace. That is the failure worth designing against: silent, and first visible on a bill.
So the header is sent on every Anthropic request this library makes, not only completions. Each surface takes it where that surface is configured:
| Surface | Where |
|---|---|
| completions | providerOptions.workspaceId per request, or new AnthropicAdapter({ apiKey, workspaceId }) as a client-wide default — the request wins |
| files | new AnthropicFileAdapter({ apiKey, workspaceId }) |
| batches | new AnthropicBatchAdapter({ apiKey, workspaceId }) — on submit and on every poll |
| token counting | new AnthropicCountApi(apiKey, fetch, baseURL, workspaceId) |
| model listing | listModelsLive({ provider: 'anthropic', apiKey, workspaceId }) |
| retrieving a file a turn produced | filled in from the client’s adapter — llm.retrieveFile(f) uses the Workspace the turn was billed to |
await complete({ model: 'anthropic/claude-haiku-4.5', apiKey, prompt: '…', providerOptions: { workspaceId: 'wrkspc_011CZkZaBF1tNoB5wlCeusgy' },});Omitted entirely when unset. “No Workspace named” and “the Workspace named is the empty string” are different requests, and only the first one means what an unconfigured client means.
Data residency (OpenAI)
Section titled “Data residency (OpenAI)”OpenAI serves the same API from four hosts — api.openai.com and
{us,eu,ae}.api.openai.com — and a project provisioned for one region must call that
region’s host. dataResidency names the region instead of making you write the URL:
import { createEngine } from '@combycode/llm-sdk';import type { OpenAIDataResidency } from '@combycode/llm-sdk';
const region: OpenAIDataResidency = 'eu';const engine = createEngine({ apiKeys: { openai: process.env.OPENAI_API_KEY ?? '' } });const llm = engine.createClient({ model: 'openai/gpt-5.4-nano', dataResidency: region });// -> every request goes to https://eu.api.openai.com/v1/...dataResidency | Host |
|---|---|
| (unset) | api.openai.com |
'global' | api.openai.com |
'us' | us.api.openai.com |
'eu' | eu.api.openai.com |
'ae' | ae.api.openai.com |
The wrong region fails loudly rather than leaking. Measured 2026-10-01 from an
unrestricted project, us. answered Attempted to access resource with incorrect regional hostname. Please make your request to api.openai.com and eu. answered
This endpoint is only accessible by projects with geography restrictions enabled.
Both replies name the host that was actually reached.
Three things are refused instead of being resolved for you:
dataResidencytogether withbaseURL— two different answers to “which host”. Picking a winner would silently discard a configuration you wrote, and you could not tell which one survived.- A region that is not one of the four —
'EU'would otherwise fall through to the default host and send EU-resident data to the global endpoint, which is the single outcome this option exists to prevent. dataResidencyon any other provider — none of them has regional hosts. Accepting it quietly would let you believe traffic was pinned to a region when the option did nothing at all. UsebaseURLif a provider offers a regional endpoint of its own.
What cache: 'auto' actually does per provider
Section titled “What cache: 'auto' actually does per provider”cache: 'auto' is one option over three quite different mechanisms, and usage.cachedTokens
reports what the provider says it reused. What you should expect differs sharply:
| Provider | Mechanism | Do you get a hit? |
|---|---|---|
| Anthropic | explicit cache_control breakpoints we set for you | Deterministic above the model’s minimum (~1024 tokens) |
| OpenAI | implicit, always on | Reliable on a repeated long prefix; promptCacheOptions for manual control |
| implicit, best-effort | Only on a large prefix, and not guaranteed even then |
Google deserves the warning. Measured on 2026-08-09 with an identical repeated request:
- A ~5,000-token prefix produced no cache hit at all on
gemini-3.6-flashorgemini-2.5-flash— neither assystemInstructionnor as leading content. Placement is not the issue; size is. - At ~15,000–40,000 tokens
gemini-3.6-flashreported hits every time (e.g. 40,010 prompt tokens → 32,737 cached). gemini-2.5-flashhit at 10k and 20k but missed at 15k and 30k in the same run.
So Google implicit caching is genuinely best-effort: a miss is not a bug, and no
cost model should assume the hit. Treat usage.cachedTokens as an observation after the fact.
When you need a guaranteed, billable cache on Google, create a cachedContents resource and pass
its name through providerOptions.cachedContent — that is explicit and deterministic.
Asking WHY the cache missed (cacheDiagnostics)
Section titled “Asking WHY the cache missed (cacheDiagnostics)”usage.cachedTokens says how much was reused. It does not say what broke the prefix, and on a
long system prompt that is the only question worth asking. Anthropic and OpenAI both answer it,
under different names; the unified option asks, and response.cacheDiagnostics carries the reply.
import { createLLM } from '@combycode/llm-sdk';
const llm = createLLM({ model: 'anthropic/claude-haiku-4.5' });const longPolicyText = 'Rule: answer in one word. '.repeat(900);
const first = await llm.complete('Summarise the policy.', { system: longPolicyText, cache: { system: true }, cacheDiagnostics: {}, // opt in, nothing to compare yet});
const second = await llm.complete('And the exceptions?', { system: longPolicyText, cache: { system: true }, cacheDiagnostics: { compareWith: first.id },});
second.cacheDiagnostics;// -> { status: 'miss', reason: 'system_changed', missedTokens: 9197, raw: {...} }// or undefined on Anthropic when the prefix WAS reused — see below.Opt in on every request in the chain, not only the one you are asking about. Measured on
2026-09-29: Anthropic keeps the prompt fingerprint only for requests that themselves sent
cacheDiagnostics. Comparing against an ordinary response returns comparison_not_found even
though the id is perfectly valid — which reads exactly like a broken feature. That is why the
first call above passes cacheDiagnostics: {} with nothing to compare.
A hit is not reported the same way, and this is the part to design around.
| Anthropic | OpenAI | |
|---|---|---|
| Prefix reused | cacheDiagnostics is absent | status: 'hit' |
| Prefix broken | status: 'miss' + reason + missedTokens | same, with a finer reason set and reusableTokens |
Unknown compareWith | status: 'comparison_not_found' (HTTP 200) | same (HTTP 200) |
| Nothing to diagnose | absent | status: 'unavailable' |
Anthropic has no hit variant: a request whose prefix was reused returns the same body an
undiagnosed request returns. Nothing here turns that silence into status: 'hit', because that
would publish our inference as the provider’s answer — read usage.cachedTokens for whether the
cache was used, and this for why it was not.
Two more measured facts worth knowing before you build on it:
- OpenAI gates it to
gpt-5.6and later. Every earlier model answersunavailableto an otherwise identical request, so testing on a mini/nano model shows a feature that appears dead. - It is a Responses-only feature on OpenAI. Chat Completions takes a
prompt_cache_optionstoo, but that one carriesmodeandttland nothing else — there is no field to name a comparison against, so the request is reported as adjusted rather than sent. reasonkeeps each provider’s own word.system_changed(Anthropic) andinput_changed(OpenAI) are not the same claim, so neither is translated into the other; bothstatusandreasonare open unions (R1) and an unrecognised value reaches you unchanged.
Requesting it where no field exists — Google, xAI, Chat Completions — is reported as an
onWarning with code request_adjusted rather than dropped in silence.
Streaming reports it too. Both providers send the diagnosis in the stream (Anthropic on
message_start, before a token is generated; OpenAI in the response envelope), so it arrives as a
cache_diagnostics stream event and is also collected onto the streamed final response —
stream() and complete() answer the same question.
Multi-turn with server-state
Section titled “Multi-turn with server-state”import { createLLM, type Message } from '@combycode/llm-sdk';
const llm = createLLM({ model: 'openai/gpt-5.4-nano', apiKey: process.env.OPENAI_API_KEY });
const messages: Message[] = [{ role: 'user', content: 'Remember the number 42.' }];const r1 = await llm.complete(messages);messages.push(llm.assistantMessage(r1)); // stamps server response id when availablemessages.push({ role: 'user', content: 'What number did I ask you to remember?' });const r2 = await llm.complete(messages);console.log(r2.text);assistantMessage() carries more than text (Google Interactions)
Section titled “assistantMessage() carries more than text (Google Interactions)”llm.assistantMessage(response) is not a convenience wrapper around the text — it stamps the
turn’s provenance, and on Google Interactions that provenance is load-bearing. Build history
with it rather than by hand, or the next turn goes out missing state the provider expects back.
A Gemini Interactions turn returns a thought step carrying nothing but a signature. Measured
2026-09-29 on gemini-3.1-flash-lite: echoing that step on the next request is accepted, and
echoing it with the signature corrupted is refused 400 Corrupted thought signature — so the
server reads it rather than tolerating it. The library keeps it on response.signatures, copies
it to message.origin.signatures, and the adapter sends it back in the position it arrived in.
const first = await llm.complete('Think, then say OK.');const history = [ { role: 'user', content: 'Think, then say OK.' }, llm.assistantMessage(first), // carries origin.signatures { role: 'user', content: 'Now say DONE.' },] satisfies Message[];await llm.complete(history); // the signed step rides alongsignatures is opaque and provider-bound: nothing here reads it, and the adapter sends it only
when origin.provider matches its own. A streamed turn keeps it too — the signature arrives as its
own delta mid-stream and rides out on the terminal done event onto response.signatures.
When you continue server-side instead (previousResponseId, or the default stateful behaviour),
the transcript is not resent at all and the provider already holds the state, so nothing is echoed.
A failed interaction now says why. Interaction.errors[] — Google’s diagnostic faults — is
lifted onto response.error when the interaction failed, where it used to arrive as
finishReason: 'error' and nothing else: an empty answer, no exception to catch, and no way to
tell a content refusal from a platform fault. On a completed interaction the field stays on
response.raw, because Google documents it as diagnostics rather than as a cause and reporting a
successful call as failed would be worse than saying nothing.
Capability-based model selection
Section titled “Capability-based model selection”import { createEngine, select, complete } from '@combycode/llm-sdk';
createEngine({ catalog: 'defaults', apiKeys: { anthropic: process.env.ANTHROPIC_API_KEY! },});
// Pick the cheapest model that supports vision.const model = select('type:chat; vision; cheap');const { text } = await complete({ model: model!, prompt: 'Describe the scene.' });console.log(text);Pre-flight cost estimate + budget guard
Section titled “Pre-flight cost estimate + budget guard”import { estimate, complete, BudgetExceededError } from '@combycode/llm-sdk';
// Estimate without sending anything.const est = await estimate({ model: 'anthropic/claude-haiku-4.5', prompt: 'Write a detailed essay on the history of computing.', maxTokens: 2000,});console.log(`Expected cost: $${est.cost.expected.toFixed(6)}`);
// Or use the inline guard on complete():try { const { text } = await complete({ model: 'anthropic/claude-haiku-4.5', apiKey: process.env.ANTHROPIC_API_KEY, prompt: 'Write a detailed essay on the history of computing.', maxTokens: 2000, maxCostUsd: 0.001, // throw before sending if estimated cost exceeds this }); console.log(text);} catch (err) { if (err instanceof BudgetExceededError) { console.error('Request would exceed budget, not sent.'); }}