optionOS optionOS
← All updates

Speak, mark, and follow the agent: Dictation context now reaches Cockpit as one flow

Optionalist Dictation now does more than turn audio into text. It keeps the target app, captured visuals, drawn regions, pending recordings, and the context delivered to the agent visible inside one session. Cockpit t…

optionOS

Optionalist Dictation now does more than turn audio into text. It keeps the target app, captured visuals, drawn regions, pending recordings, and the context delivered to the agent visible inside one session. Cockpit then shows what that packet became on the agent side through one card.

This release connects recording history, the live flow, annotated-image previews, the capture stack, and the Cockpit agent card. The first visual opens the complete nested journey; the remaining visuals let you inspect each screen directly.

Even this What's New was produced by speaking

Five raw voice recordings took 14 minutes 54 seconds in total. The recordings, marked visuals, and final instruction were delivered as one context packet to a gpt-5.6-sol · xhigh · fast session. Enter region 3 to move through recording history and live Dictation, then continue to Cockpit or the Codex delivery.

Five starred voice recordings and the single AI prompt assembled from those spoken recordings Opening recordings — starred voice conversations that begin the promptContinuation recordings — the other starred conversations collected for the same taskAssembled prompt — one agent packet containing the raw speech and visual context
·
1 / 3

Opening recordings — starred voice conversations that begin the prompt

Opening recordings — starred voice conversations that begin the prompt
Five recordings → one prompt → Dictation → Cockpit or Codex. Click through the nested visual path.

Seven screens and 63 marked regions share one tour; the hover door also names the real app's hover behavior.

Find recordings by duration, tag, relationship, or text

Dictation Browser lists every raw conversation with its duration. It keeps Command-D or context-menu tags, relationships between conversations pasted together, and the full selected transcript on one surface. Select a word, natural pause, or waveform position to play the audio from that point.

Dictation Browser recording list, durations and tags, related conversations, audio playback, and search suggestions Conversation list — shows all Dictation recordings in one placeDuration — shows the length of each raw conversationTags — user labels added with Command-D or the context menuRelationship mark — records the link between conversations pasted togetherSelected conversation — shows the human's transcript in detailRelationship strip — shows other conversations connected to this recordingTransformation trace — exposes a phrase removed by the custom cleanup algorithmAudio playback bar — plays from a word, pause, or waveform positionNatural pause — shows a silent area in light grayTotal speech duration — combined speaking time of the panel's recordingsSearch — finds recordings by text, tag, and capability expressionSuggestions — offers matching filters as the query changes
·
1 / 12

Conversation list — shows all Dictation recordings in one place

Conversation list — shows all Dictation recordings in one place
Search, suggestions, duration, tags, relationships, and playback live on the same recording surface.

Add visuals and marked regions while you keep speaking

The live flow shows the target app and places marks such as O1.R1 and captured-image references inline. Switch between Conversation, Events, and Context to inspect plain speech, structured events sent to the agent, or the human-readable context separately.

Live Dictation flow with transformed speech, visual references, Quick Actions, and the captured-item stack Live speech flow — shows the session state while recording continuesTarget app — shows the context where the conversation startedPipeline stages — hovering opens the transformed-conversation panelTransformed conversation — keeps speech, annotations, and visual context togetherInline O1.R1 reference — the first area marked while speakingOther rectangle references — connect drawn image regions to the conversationCaptured-image reference — the O1 name and capture identityView selector — switches between Conversation, Events, and ContextEvents — shows the structured context sent to the agentConversation — shows plain human speech without annotation tokensDelivery target — shows that this conversation is headed to CodexQuick Actions — opens capture actions by click or shortcutScreen capture — adds a new screen to the annotation flowCaptured-item stack — gathers items attached to the conversationSource app and title — shows that the capture came from CockpitRepresentation — shows that the item will be sent as an Image and can be changedApproximate magnitude — tokens when known, otherwise kilobytes
·
1 / 17

Live speech flow — shows the session state while recording continues

Live speech flow — shows the session state while recording continues
Speech, target app, screen captures, and rectangle references stay in one uninterrupted session.

Enter the visual in the conversation, then inspect the screen inside it

Hovering the live-flow summary opens the annotated-image preview. Select the Cockpit visual inside that preview to enter the next layer. The image-within-image relation preserves the conversation's real context tree instead of flattening it into a gallery.

Annotated Cockpit image and source context opened from the live speech-flow hover preview Live-flow summary — hovering opens its detailed previewAnnotated Cockpit preview — shows the referenced image with its marked regionsFiles and Image selector — opens the content's source file or visual viewSource context — shows that the capture came from Cockpit and optionOS
·
1 / 4

Live-flow summary — hovering opens its detailed preview

Live-flow summary — hovering opens its detailed preview
The hover preview opens the annotated Cockpit screen together with its source context.

Follow the agent's goal, messages, and file changes from one card

The new compact Cockpit card combines terminal source, agent goal, first and last human and agent messages, subagent conversations, and file states. Copy only the human messages or an editable context, fork the session, or start optionOS's own Compact action.

Cockpit terminal filters, expanded agent card, conversation and file-change summaries, and session actions Terminal filters — reveal open Claude, Codex, and Terminal sessionsExpanded agent card — gathers the session's essential information in one cardCard index — shows that this is agent 333 in the complete sequenceAgent identity — provider icon, session title, and primary informationGoal — appears on the card when the agent has set oneTerminal source — the terminal running the agent; Ghostty hereLive-stream control — starts or pauses tracking the agent's changesSession activity — messages, subagent conversations, and file changesUser messages — only the first and last human messageAgent messages — only the first and last agent messageSubagent conversations — messages exchanged with spawned subagentsEdited files — number of files edited by the agentUncommitted files — changed files that have not entered a commitCommitted changes — number of changes already committedConversation actions — experimental copy optionsCopy human messages onlyCopy editable context — messages and changed-file summariesLast meaningful activity — working area and elapsed timeFork — separates the session into a new terminal flowCompact — condenses the conversation with optionOS's own mechanism
·
1 / 20

Terminal filters — reveal open Claude, Codex, and Terminal sessions

Terminal filters — reveal open Claude, Codex, and Terminal sessions
Twenty marked regions expose the compact agent card from identity through Compact.

Keep speaking before you paste; recordings stay in one session

The captured-item stack grows as Quick Actions adds new screens. The red waveform shows that audio is being received, while the small progress indicator shows background transformation. Unpasted recordings wait and are delivered together as one session.

Quick Actions capture menu, growing item stack, live waveform, and pending recordings Captured items — the stack grows with new visuals while speech continuesSource item — shows that AGENTS.md was captured from VS CodeLive waveform — makes audio reception immediately visibleBackground transformation — shows that speech is still becoming textPending recordings — accumulates unpasted speech into one session
·
1 / 5

Captured items — the stack grows with new visuals while speech continues

Captured items — the stack grows with new visuals while speech continues
Speak, capture, and speak again; pending recordings stay together until delivery.

The result is not a manually typed prompt; it is context delivered through speech

The content arriving in the Ghostty Codex session appears as an optionOS transcript packet. There is no separately typed explanation in this example: speech, visuals, and marked regions were delivered together.

Open Codex session in Ghostty and an optionOS transcript packet delivered entirely through speech Codex session — the Ghostty tab that received the conversationTranscript packet — spoken content delivered without manually typed companion text
·
1 / 2

Codex session — the Ghostty tab that received the conversation

Codex session — the Ghostty tab that received the conversation
The Codex session and its speech-delivered transcript packet share one screen.