Speak, mark, and follow the agent: Dictation context now reaches Cockpit as one flow
Optionalist Dictation now does more than turn audio into text. It keeps the target app, captured visuals, drawn regions, pending recordings, and the context delivered to the agent visible inside one session. Cockpit t…
optionOS
Optionalist Dictation now does more than turn audio into text. It keeps the target app, captured visuals, drawn regions, pending recordings, and the context delivered to the agent visible inside one session. Cockpit then shows what that packet became on the agent side through one card.
This release connects recording history, the live flow, annotated-image previews, the capture stack, and the Cockpit agent card. The first visual opens the complete nested journey; the remaining visuals let you inspect each screen directly.
Even this What's New was produced by speaking
Five raw voice recordings took 14 minutes 54 seconds in total. The recordings, marked visuals, and final instruction were delivered as one context packet to a gpt-5.6-sol · xhigh · fast session. Enter region 3 to move through recording history and live Dictation, then continue to Cockpit or the Codex delivery.
Five recordings → one prompt → Dictation → Cockpit or Codex. Click through the nested visual path.
Seven screens and 63 marked regions share one tour; the hover door also names the real app's hover behavior.
Find recordings by duration, tag, relationship, or text
Dictation Browser lists every raw conversation with its duration. It keeps Command-D or context-menu tags, relationships between conversations pasted together, and the full selected transcript on one surface. Select a word, natural pause, or waveform position to play the audio from that point.
Search, suggestions, duration, tags, relationships, and playback live on the same recording surface.
Add visuals and marked regions while you keep speaking
The live flow shows the target app and places marks such as O1.R1 and captured-image references inline. Switch between Conversation, Events, and Context to inspect plain speech, structured events sent to the agent, or the human-readable context separately.
Speech, target app, screen captures, and rectangle references stay in one uninterrupted session.
Enter the visual in the conversation, then inspect the screen inside it
Hovering the live-flow summary opens the annotated-image preview. Select the Cockpit visual inside that preview to enter the next layer. The image-within-image relation preserves the conversation's real context tree instead of flattening it into a gallery.
The hover preview opens the annotated Cockpit screen together with its source context.
Follow the agent's goal, messages, and file changes from one card
The new compact Cockpit card combines terminal source, agent goal, first and last human and agent messages, subagent conversations, and file states. Copy only the human messages or an editable context, fork the session, or start optionOS's own Compact action.
Twenty marked regions expose the compact agent card from identity through Compact.
Keep speaking before you paste; recordings stay in one session
The captured-item stack grows as Quick Actions adds new screens. The red waveform shows that audio is being received, while the small progress indicator shows background transformation. Unpasted recordings wait and are delivered together as one session.
Speak, capture, and speak again; pending recordings stay together until delivery.
The result is not a manually typed prompt; it is context delivered through speech
The content arriving in the Ghostty Codex session appears as an optionOS transcript packet. There is no separately typed explanation in this example: speech, visuals, and marked regions were delivered together.
The Codex session and its speech-delivered transcript packet share one screen.