WikifitaGitHub live67e8de5
pesquisa · openai-research/voice-and-tasks

Voz e Tasks no ChatGPT Desktop

Ditado (/transcribe), live voice thread-scoped (WebRTC realtime) e tasks do Work criadas pelo Chat (/wham)

Baixar raw

Voz e Tasks no ChatGPT Desktop

Part of openai-research.

Dictate (one-shot transcription)

Renderer grava áudio (webm) → IPC → main → POST /backend-api/transcribe (multipart/form-data: arquivo codex.webm + language) com Attach-Auth + Attach-Integrity-State{text}. Requer login ("Sign in to ChatGPT to use dictation" quando o app-server não tem token). Histórico local em transcription-history.jsonl.

Live voice (voice chat) — real capture 2026-08-26 (chat submode!)

The voice chat of the ChatGPT mode (the 🎙 button in the composer) is handled by the embedded app-server — even when used from the chat submode (a chat-mode realtime conversation produced a realtime thread in state_5).

  • Thread created: threads.thread_source = "realtime_voice", title = the literal marker <realtime_delegation>\n<input>…</input>… (the delegation input embedded in the title), e.g. thread 01a04027-… with model: gpt-5.6-terra, reasoning_effort: low, tokens_used: 67,230.
  • Wire: the app-server opens the live realtime websocket: realtime websocket: wss://api.openai.com/v1/live/rtc_u2_<token>realtime websocket connected ws_url=… status=101 Switching Protocols; events flow through codex_api::endpoint::realtime_websocket::protocol ("realtime websocket parsed event" — the app-server ignores unsupported text frames, some events are parsed; reflected into codex_core::realtime_conversation "received realtime conversation event"). Audio/voice over the live socket; the transcribed text becomes a TurnInput (submission with the user text) → the agent turn runs in that thread (observed: "Ah, eu tô bem. Acho que é isso. Você pode encerrar a conversa" → turn with gpt-5.6-terra, reasoning_effort=low).
  • UX: the turn shows "Worked for 7s" + "End realtime voice call"; sidebar thread hover shows execution-location local ("Runs on your computer") with "Voice chat" — because it is an app-server thread (works Local).
  • No /f/conversation for the live turns; the classic HTTP path appears only around the session (warm-up settings/voices?spoken_language=en-US&voice_mode=advanced, prepare, attestation_challenge).
  • GET /backend-api/celsius/ws/user is the conversation/presence websocket of the chat web backend (getConversationWebSocketUrl(){websocket_url}; main initiator y; ~30min reconnect cadence) — NOT the voice channel.
  • POST /backend-api/transcribe (dictate, all submodes): multipart boundary=----codex-transcribe-<uuid>, main process initiator y (stack captured), result text lands in the composer where the user edits before sending (/f/conversation/prepare follows).
  • Capture note: the live WS is created by the Rust app-server — the renderer/main WebSocket hooks (006) do not see it; the frames are in the 005 stdio dump + logs_2.sqlite (event names, not payloads).

Live voice (thread-scoped, same in chat and work submodes + Codex mode)

Three-layer architecture:

  1. Renderer: AudioWorklet (realtime-buffered-audio-worklet) captura/playa; surface própria (realtime-voice-launch-surface, realtime-voice-stage-layout)
  2. App-server nativo (protocolo local v2 — app-server-protocol/src/protocol/v2/realtime.rs): ThreadRealtimeStartParams (thread_id, client_managed_handoffs, delegation_ack_filler, codex_response_handoff_mode, BEM channel prefixes…)
  3. Wire: WebRTC com a Realtime API da OpenAI — POST /realtime/calls (SDP offer) + sideband (spawn_webrtc_sideband_input_task em core/src/realtime_conversation.rs)

Notificações locais thread/realtime/*: started, itemAdded, transcript/delta, transcript/done, outputAudio/delta, sdp, error, closed. O transcript da voz alimenta o mesmo pipeline de turn (handoff para o Codex com codex_response_handoff_mode).

Work tasks created from the chat submode (wham)

The chat submode fires Work tasks passing the conversation as context:

POST /wham/tasks  (nova task)
{
  "new_task": {"branch": "<ref>", "environment_id": "...", "run_environment_in_qa_mode": false},
  "metadata": {"model_slug": "gpt-5.5"},
  "input_items": [
    {"type": "prior_conversation", "conversation": [...], "diff": "..."},
    {"type": "message", "role": "user", "content": [{"content_type": "text", "text": "<prompt>"}]}
  ]
}

POST /wham/tasks  (follow-up)
{"follow_up": {"task_id": "...", "turn_id": "...", "environment_mode": "ask|code"}, "input_items": [...]}

Ciclo: GET /wham/tasks/list, /wham/tasks/{id}, /wham/tasks/{id}/turns, /turns/{turn_id}/logs, /turns/{turn_id}/pr; cancel/archive/recover. Ambientes: /wham/environments, worktree snapshots (/wham/worktree_snapshots/upload_url). The UI "Continuing from …" reply is the prior_conversation in action.

Realtime voice internals (WebRTC + protocol types)

The voice chat is WebRTC end-to-end: the renderer runs the RTCPeerConnection:

  • RTCPeerConnection created without any iceServers (0 hits in the asar) — no STUN/TURN config in the client (candidates stay host; the rendezvous is the OpenAI realtime service via the app-server call). This is the "server on the other side" picture: the app-server fronts the call.
  • Media: microphone stream (getWebRtcStream), hidden <audio> element (autoplay, muted configurable), data channel c.createDataChannel(<var>) (channel name is a variable — not a literal), onmessage → string → the realtime event callback; onclose/onerror → connection failure handling.
  • Startup-gate reason codes: CODEX_REALTIME_VOICE_STARTUP_MEDIA_FAILURE_REASON_* {MEDIA_APIS_UNAVAILABLE, PEER_CONNECTION_INITIALIZATION_FAILED, DATA_CHANNEL_INITIALIZATION_FAILED…} and startUnavailable "Voice chat is not ready ({reason})" for {session_already_starting, configuration_not_ready, microphone_unavailable, webrtc_unavailable, voice_disabled}.
  • Protocol thread/realtime/start (app-server v2): {threadId, clientManageHandoffs?, delegationAckFiller?, flushTranscriptTailOnSessionEnd, codexResponsesAsItems?, codexResponseItemPrefix?, codexResponseHandoffMode?, includeStartupContext, initialItems, outputModality, realtimeSessionId?, voice?, prompt?, realtimeStart/EndInstructions?, transport: {type: 'webrtc', sdp}} — the JSON body the renderer/main sends (exact field list from the asar; the source types in codex-rs/app-server-protocol/src/protocol/v2/realtime.rs):
    • ThreadRealtimeStartParams, ThreadRealtimeStartTransport { SdpOffer } | CallId, ThreadRealtimeAudioChunk {data, sample_rate, num_channels, samples_per_channel?, item_id?}, ThreadRealtimeInitialItem {role, text} (dev note: "SDP offer generated by a WebRTC RTCPeerConnection after configuring audio and the realtime events data channel")
    • Append/stop: thread/realtime/appendAudio|appendText|appendSpeech|stop; list: thread/realtime/listVoicesRealtimeVoicesList
    • Notifications: thread/realtime/{started, itemAdded, transcript/delta, transcript/done, outputAudio/delta, sdp, error, closed}
  • App-server wire: connects as a client to the live realtime websocket wss://api.openai.com/v1/live/rtc_u2_<token> (status 101; codex_api::endpoint::realtime_websocket — "parsed event", "ignored unsupported text frame"; codex_core::realtime_conversation "received realtime conversation event") — this is the server-side channel of the call (the RTC rendezvous/session events).
  • Turn path: the transcribed user text lands as a TurnInput submission (thread realtime_voice), the agent turn runs with the thread model (observed gpt-5.6-terra, reasoning_effort=low, thread tokens_used: 67,230this is Codex-backend processing → codex quota). UI: "Worked for Ns" + "End realtime voice call"; sidebar shows "Work · Voice chat" (execution-local).

WebSocket inventory (what actually uses a WebSocket)

ChannelWhoPurposeCaptured by
wss://api.openai.com/v1/live/rtc_u2_…app-server (Rust)the realtime call channel (events for the voice session)005-stdio + logs_2.sqlite (event names only)
GET /celsius/ws/user → {websocket_url} then new WebSocket(url)main (gets URL) / renderer (connects)conversation presence websocket of the chat backend (the getConversationWebSocketUrl API + reconnect-capable connect() class) — reconnection cadence ~30min observed001 (URL fetch) / 006 (renderer ws)
remote_control_controller_websocket / remote_control_client_websocketrenderer/mainremote control sessions (device-key enrollment, codex.remote_control.enroll, scope remote_control_controller_websocket)
WebRTC data channel (voice)rendererrealtime events (NOT a websocket)not by 006

Correction on the earlier capture note: voice chat is WebRTC (with the app-server live wss://api.openai.com/v1/live channel), NOT a renderer websocket — 006 does not (and cannot) see the voice frames; 005 sees the Rust connection. 006 stays for renderer websockets that exist (conversation presence, remote-control).

Other chat submode backend services mapped

  • /wham/usage, /wham/usage/thread_usage/query — usage por thread
  • /wham/rate-limit-reset-credits[(/consume)] — créditos de reset de limite
  • /wham/remote/control/* — pareamento de clientes de controle remoto
  • /wham/shared_threads[/upload_urls] — threads compartilhadas
  • /ecosystem/*, /gizmos/* — plugins/widgets/projetos (gizmos = Projects)
  • /celsius/ws/user — websocket de usuário
  • /automations/* (Scheduled), /files/*, /projects/*

TypeScript types (voice + tasks)

// ── dictate ──────────────────────────────────────────────────────────────
export interface TranscribeMultipart {
  boundary: `----codex-transcribe-${string}`   // multipart/form-data
  file: Blob                                      // 'codex.webm'
  language: string
}
export interface TranscribeResponse { text: string }

// ── realtime start (app-server v2) ───────────────────────────────────────
export type RealtimeVoice = string | null          // voice id (listVoices)
export type RealtimeOutputModality = 'text' | 'audio' | ('text' | 'audio')[] | unknown
export interface ThreadRealtimeInitialItem { role: 'user' | 'assistant'; text: string }
export interface ThreadRealtimeStartParams {
  thread_id: string
  client_managed_handoffs?: boolean | null
  delegation_ack_filler?: boolean | null
  flush_transcript_tail_on_session_end?: boolean | null
  codex_responses_as_items?: boolean | null
  codex_response_item_prefix?: string | null
  codex_response_handoff_mode?: string | null       // 'thinking' | … (Frameless Bidi)
  include_startup_context?: boolean
  initial_items?: ThreadRealtimeInitialItem[]
  output_modality?: RealtimeOutputModality
  realtime_session_id?: string | null
  voice?: RealtimeVoice
  prompt?: string | null
  realtime_start_instructions?: string | null
  realtime_end_instructions?: string | null
  transport?: { type: 'webrtc'; sdp: string } | { type: 'call'; call_id: string }
}
export interface ThreadRealtimeAudioChunk {
  data: string
  sample_rate: number
  num_channels: number
  samples_per_channel?: number
  item_id?: string
}
// notifications thread/realtime/* (the union):
export type RealtimeNotification =
  | { type: 'started'; thread_id: string; realtime_session_id?: string; version: 'v1' | 'v2' | 'v3' | string }
  | { type: 'itemAdded'; thread_id: string; item: unknown }
  | { type: 'transcript/delta'; thread_id: string; role: string; delta: string }
  | { type: 'transcript/done'; thread_id: string; role: string; text: string }
  | { type: 'outputAudio/delta'; thread_id: string; audio: ThreadRealtimeAudioChunk }
  | { type: 'sdp' | 'error' | 'closed'; thread_id: string; [k: string]: unknown }

// ── WebRTC (renderer) ────────────────────────────────────────────────────
export interface RealtimeStartupReasons {
  session_already_starting?: string
  configuration_not_ready?: string
  microphone_unavailable?: string
  webrtc_unavailable?: string
  voice_disabled?: string
}
export interface RealtimeErrorCodes {
  CODEX_REALTIME_VOICE_STARTUP_MEDIA_FAILURE_REASON_MEDIA_APIS_UNAVAILABLE: string
  CODEX_REALTIME_VOICE_STARTUP_MEDIA_FAILURE_REASON_PEER_CONNECTION_INITIALIZATION_FAILED: string
  CODEX_REALTIME_VOICE_STARTUP_MEDIA_FAILURE_REASON_DATA_CHANNEL_INITIALIZATION_FAILED: string
}

// ── wham tasks ───────────────────────────────────────────────────────────
export interface WhamTaskCreateBody {
  new_task: { branch: string; environment_id: string; run_environment_in_qa_mode: boolean }
  metadata: { model_slug: string }                  // 'gpt-5.5'
  input_items: Array<
    | { type: 'prior_conversation'; conversation: unknown[]; diff: string | null }
    | { type: 'message'; role: 'user'; content: Array<{ content_type: 'text'; text: string }> }
  >
}
export interface WhamTaskFollowUpBody {
  follow_up: { task_id: string; turn_id: string; environment_mode: 'ask' | 'code' }
  input_items: WhamTaskCreateBody['input_items']
}
export interface WhamUsageResponse {
  user_id: string
  account_id: string
  email: string
  plan_type: 'plus' | 'free' | string
  rate_limit: {
    allowed: boolean
    limit_reached: boolean
    primary_window: { used_percent: number; limit_window_seconds: number; reset_after_seconds: number; reset_at: number }
    secondary_window: { used_percent: number; limit_window_seconds: number; reset_after_seconds: number; reset_at: number }
  }
  credits: { has_credits: boolean; unlimited: boolean; overage_limit_reached: boolean; balance: string
             approx_local_messages: [number, number]; approx_cloud_messages: [number, number] }
  spend_control: { reached: boolean; individual_limit: unknown | null }
  rate_limit_reached_type: string | null
  rate_limit_reset_credits: { available_count: number; applicable_available_count: number }
}

Cross-references