# Speak, mark, and follow the agent: Dictation context now reaches Cockpit as one flow

> Optionalist Dictation now does more than turn audio into text. It keeps the target app, captured visuals, drawn regions, pending recordings, and the context delivered to the agent visible inside one session. Cockpit t…

Canonical: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/
Published: 2026-07-19
Source: optionos-product:r19

Optionalist Dictation now does more than turn audio into text. It keeps the target app, captured visuals, drawn regions, pending recordings, and the context delivered to the agent visible inside one session. Cockpit then shows what that packet became on the agent side through one card.

This release connects recording history, the live flow, annotated-image previews, the capture stack, and the Cockpit agent card. The first visual opens the complete nested journey; the remaining visuals let you inspect each screen directly.

### Even this What's New was produced by speaking

Five raw voice recordings took **14 minutes 54 seconds** in total. The recordings, marked visuals, and final instruction were delivered as one context packet to a `gpt-5.6-sol · xhigh · fast` session. Enter region 3 to move through recording history and live Dictation, then continue to Cockpit or the Codex delivery.

![Five starred voice recordings and the single AI prompt assembled from those spoken recordings](https://optionos.app/whats-new/voice-context-recording-source.png)
*Five recordings → one prompt → Dictation → Cockpit or Codex. Click through the nested visual path.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-workflow

### Find recordings by duration, tag, relationship, or text

Dictation Browser lists every raw conversation with its duration. It keeps Command-D or context-menu tags, relationships between conversations pasted together, and the full selected transcript on one surface. Select a word, natural pause, or waveform position to play the audio from that point.

![Dictation Browser recording list, durations and tags, related conversations, audio playback, and search suggestions](https://optionos.app/whats-new/voice-context-dictation-browser.png)
*Search, suggestions, duration, tags, relationships, and playback live on the same recording surface.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-dictation-browser

### Add visuals and marked regions while you keep speaking

The live flow shows the target app and places marks such as O1.R1 and captured-image references inline. Switch between Conversation, Events, and Context to inspect plain speech, structured events sent to the agent, or the human-readable context separately.

![Live Dictation flow with transformed speech, visual references, Quick Actions, and the captured-item stack](https://optionos.app/whats-new/voice-context-live-dictation.png)
*Speech, target app, screen captures, and rectangle references stay in one uninterrupted session.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-live-dictation

### Enter the visual in the conversation, then inspect the screen inside it

Hovering the live-flow summary opens the annotated-image preview. Select the Cockpit visual inside that preview to enter the next layer. The image-within-image relation preserves the conversation's real context tree instead of flattening it into a gallery.

![Annotated Cockpit image and source context opened from the live speech-flow hover preview](https://optionos.app/whats-new/voice-context-hover-preview.png)
*The hover preview opens the annotated Cockpit screen together with its source context.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-hover-preview

### Follow the agent's goal, messages, and file changes from one card

The new compact Cockpit card combines terminal source, agent goal, first and last human and agent messages, subagent conversations, and file states. Copy only the human messages or an editable context, fork the session, or start optionOS's own Compact action.

![Cockpit terminal filters, expanded agent card, conversation and file-change summaries, and session actions](https://optionos.app/whats-new/voice-context-cockpit.png)
*Twenty marked regions expose the compact agent card from identity through Compact.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-cockpit

### Keep speaking before you paste; recordings stay in one session

The captured-item stack grows as Quick Actions adds new screens. The red waveform shows that audio is being received, while the small progress indicator shows background transformation. Unpasted recordings wait and are delivered together as one session.

![Quick Actions capture menu, growing item stack, live waveform, and pending recordings](https://optionos.app/whats-new/voice-context-capture-stack.png)
*Speak, capture, and speak again; pending recordings stay together until delivery.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-capture-stack

### The result is not a manually typed prompt; it is context delivered through speech

The content arriving in the Ghostty Codex session appears as an optionOS transcript packet. There is no separately typed explanation in this example: speech, visuals, and marked regions were delivered together.

![Open Codex session in Ghostty and an optionOS transcript packet delivered entirely through speech](https://optionos.app/whats-new/voice-context-codex-delivery.png)
*The Codex session and its speech-delivered transcript packet share one screen.*
Interactive scene: https://optionos.app/whats-new/en/2026-07-19-optionos-mac-voice-context-workflow/#voice-context-codex-delivery
