Skip to content

Realtime (Live) Sessions

createRealtime() opens a persistent WebSocket session to a provider’s live API, normalizing the two very different provider protocols (OpenAI’s typed event stream, Google’s turn-based bidirectional stream) onto one event model.

Beta. Both underlying provider APIs are in beta. Expect breaking changes from providers independent of this SDK.

ProviderModel exampleNotes
openaigpt-realtime-2Full duplex, text + audio
googlegemini-3.1-flash-liveTurn-based bidirectional, audio-native

Passing any other provider throws immediately with a clear message.

createRealtime(opts: CreateRealtimeOptions): RealtimeSession

CreateRealtimeOptions:

FieldTypeRequiredNotes
modelstringyesBare (gpt-realtime-2) or namespaced (openai/gpt-realtime-2)
providerProviderNamewhen model is bareIgnored when model is namespaced
apiKeystringnoFalls back to engine.apiKeys[provider]
modalitiesRealtimeModality[]no'text' and/or 'audio'. Default ['text']
audioAudioOptionsno{ voice?, format? } for audio output
voicestringnoDeprecated; use audio.voice
instructionsstringnoSystem-level instructions for the session
translation{ targetLanguageCode?, echoTargetLanguage? }noGoogle. Turns the session into a live translator — see below
affectiveDialogbooleannoGoogle. Detect the speaker’s emotion and adapt the reply
inputTranscription{ mode?: 'VERBATIM' | 'SMART' }noGoogle. How to transcribe what the session HEARS
engineEngineHandlenoDefaults to the registered engine

Returns a RealtimeSession immediately (synchronous). The underlying WebSocket connection opens asynchronously; listen for the 'open' event before sending.

gemini-3.5-live-translate is a translator, and until now there was no way to tell it what to translate into — so it could be connected to and had nothing to do. translation is what makes it usable:

const session = createRealtime({
model: 'google/gemini-3.5-live-translate',
translation: { targetLanguageCode: 'es', echoTargetLanguage: false },
modalities: ['audio'],
});

echoTargetLanguage decides what happens when the target language is already being spoken: true parrots it back, false stays quiet. In a two-way conversation that is the difference between hearing yourself repeated and not, so false is usually what you want — and it is sent whenever you mention it, rather than being dropped for being falsy.

Measured 2026-09-30 by streaming the same English sentence into two sessions, which is the only way to check this: a bogus target language also returns setupComplete, so “the server accepted it” proves nothing. With no translationConfig the output transcription came back empty; with targetLanguageCode: 'es' it came back “Buenos días. La reunión se ha”.

interface RealtimeSession {
send(input: RealtimeInput, opts?: { turnComplete?: boolean }): void;
on<E extends RealtimeEventType>(type: E, cb: (e: ...) => void): () => void;
close(): void;
}

RealtimeInput:

interface RealtimeInput {
text?: string;
audio?: Uint8Array; // raw audio bytes (provider-specific encoding, e.g. PCM16)
}

send() defaults to turnComplete: true — commits the turn and requests a response. Pass turnComplete: false to stream a single turn across multiple send() calls (useful for chunked audio).

on() returns an unsubscribe function. Call it to stop receiving events of that type.

type RealtimeEvent =
| { type: 'open' }
| { type: 'text'; delta: string }
| { type: 'audio'; chunk: Uint8Array; mimeType: string; sampleRate?: number }
| { type: 'turnComplete' }
| { type: 'usage'; usage: Usage }
| { type: 'error'; error: Error }
| { type: 'close' };
EventWhen
openWebSocket connected and session ready
textText delta from the model (stream chunks)
audioAudio chunk from the model
turnCompleteModel finished a response turn
usageToken usage reported (fires once per turn; wired into cost pipeline)
errorTransport or protocol error
closeSocket closed (normal or abnormal)

Usage events are automatically forwarded to the onCompletion hook so the CostCollector tracks and prices realtime calls alongside regular completions.

import { createEngine, createRealtime } from '@combycode/llm-sdk';
createEngine({ apiKeys: { openai: process.env.OPENAI_API_KEY! } });
const session = createRealtime({
model: 'openai/gpt-realtime-2',
modalities: ['text'],
instructions: 'You are a helpful assistant.',
});
const unsubText = session.on('text', (e) => {
process.stdout.write(e.delta);
});
session.on('turnComplete', () => {
console.log('\n[turn complete]');
session.close();
});
session.on('open', () => {
session.send({ text: 'Hello! Tell me a short joke.' });
});
session.on('error', (e) => console.error('Realtime error:', e.error));
session.on('close', () => unsubText());
import { createEngine, createRealtime } from '@combycode/llm-sdk';
createEngine({ apiKeys: { openai: process.env.OPENAI_API_KEY! } });
const session = createRealtime({
model: 'openai/gpt-realtime-2',
modalities: ['audio', 'text'],
audio: { voice: 'alloy' },
instructions: 'Respond concisely.',
});
// Collect audio chunks.
const audioChunks: Uint8Array[] = [];
session.on('audio', (e) => {
audioChunks.push(e.chunk);
});
session.on('turnComplete', () => {
console.log(`Got ${audioChunks.length} audio chunks.`);
// Combine and play / write to file as needed.
session.close();
});
session.on('open', () => {
// Send pre-encoded PCM16 audio or a text prompt.
session.send({ text: 'Say "Hello, world!" in a friendly tone.' });
});

In addition to the onCompletion event (for cost tracking), the network layer emits three realtime-specific hooks on engine.hooks:

HookWhen
onRealtimeOpenSocket connected (provider, model, url)
onRealtimeFrameEach frame direction/size (metadata only, no payload)
onRealtimeCloseSocket closed (code, reason)
onRealtimeErrorTransport error
engine.hooks.on('onRealtimeOpen', (ctx) => {
console.log(`Realtime connected: ${ctx.provider}/${ctx.model}`);
});