optionOS optionOS

Use speech, visuals, and marked regions as one task context

How do I use it?
Tags
  • VOICEVoice-driven production.
  • CAPTUREScreen, GIF, scrolling capture or clipboard signal.
  • ANNOTATIONMark, number, box or arrow tied to a speech moment.
  • WORKFLOWRepeatable working flow.

1 — Start recording and speak naturally. Do not try to make the first sentence perfect. The live waveform shows that audio is being received, and the small progress indicator shows that speech is becoming text in the background.

Quick Actions capture menu, growing item stack, live waveform, and pending recordings Captured items — the stack grows with new visuals while speech continuesSource item — shows that AGENTS.md was captured from VS CodeLive waveform — makes audio reception immediately visibleBackground transformation — shows that speech is still becoming textPending recordings — accumulates unpasted speech into one session
·
1 / 5

Captured items — the stack grows with new visuals while speech continues

Captured items — the stack grows with new visuals while speech continues
Speak, capture, and speak again; pending recordings stay together until delivery.

2 — When you say “this part,” capture the screen and mark the region. Choose Screen Capture from Quick Actions or use the capture shortcut. The rectangle enters the conversation as a clickable O1.R1-style reference, so you do not need to write a separate explanation list.

Live Dictation flow with transformed speech, visual references, Quick Actions, and the captured-item stack Live speech flow — shows the session state while recording continuesTarget app — shows the context where the conversation startedPipeline stages — hovering opens the transformed-conversation panelTransformed conversation — keeps speech, annotations, and visual context togetherInline O1.R1 reference — the first area marked while speakingOther rectangle references — connect drawn image regions to the conversationCaptured-image reference — the O1 name and capture identityView selector — switches between Conversation, Events, and ContextEvents — shows the structured context sent to the agentConversation — shows plain human speech without annotation tokensDelivery target — shows that this conversation is headed to CodexQuick Actions — opens capture actions by click or shortcutScreen capture — adds a new screen to the annotation flowCaptured-item stack — gathers items attached to the conversationSource app and title — shows that the capture came from CockpitRepresentation — shows that the item will be sent as an Image and can be changedApproximate magnitude — tokens when known, otherwise kilobytes
·
1 / 17

Live speech flow — shows the session state while recording continues

Live speech flow — shows the session state while recording continues
Speech, target app, screen captures, and rectangle references stay in one uninterrupted session.

3 — Check which app and visual the conversation belongs to. The live flow shows the target app. Conversation opens plain speech, Events exposes the structure sent to the agent, and Context opens the combined human-readable explanation. Each captured-item row carries its source app, representation, and approximate magnitude.

4 — You do not need to paste immediately. If you start another recording, the previous unpasted recording waits. When you deliver, pending recordings merge into one session instead of making you rebuild context by hand.

5 — Find an older conversation in Dictation Browser. Search by duration, user tag, relationship, text, or a suggested filter. Select a word, pause, or waveform position in the chosen recording to play audio from there.

Dictation Browser recording list, durations and tags, related conversations, audio playback, and search suggestions Conversation list — shows all Dictation recordings in one placeDuration — shows the length of each raw conversationTags — user labels added with Command-D or the context menuRelationship mark — records the link between conversations pasted togetherSelected conversation — shows the human's transcript in detailRelationship strip — shows other conversations connected to this recordingTransformation trace — exposes a phrase removed by the custom cleanup algorithmAudio playback bar — plays from a word, pause, or waveform positionNatural pause — shows a silent area in light grayTotal speech duration — combined speaking time of the panel's recordingsSearch — finds recordings by text, tag, and capability expressionSuggestions — offers matching filters as the query changes
·
1 / 12

Conversation list — shows all Dictation recordings in one place

Conversation list — shows all Dictation recordings in one place
Search, suggestions, duration, tags, relationships, and playback live on the same recording surface.

6 — Follow the context tree through nested visuals. The first tour moves from the assembled prompt into recording history, live Dictation, the hover preview, and Cockpit. Its other branch moves from the capture stack to the Codex delivery. Each screen opens from the context of the previous visual.

Five starred voice recordings and the single AI prompt assembled from those spoken recordings Opening recordings — starred voice conversations that begin the promptContinuation recordings — the other starred conversations collected for the same taskAssembled prompt — one agent packet containing the raw speech and visual context
·
1 / 3

Opening recordings — starred voice conversations that begin the prompt

Opening recordings — starred voice conversations that begin the prompt
Five recordings → one prompt → Dictation → Cockpit or Codex. Click through the nested visual path.

Seven screens and 63 marked regions share one tour; the hover door also names the real app's hover behavior.

7 — Watch the Cockpit card while the agent works. The card shows the goal, terminal, first and last human and agent messages, subagent conversations, and file states. Copy only human messages, take editable context, fork, or use the optionOS Compact action when needed.

Cockpit terminal filters, expanded agent card, conversation and file-change summaries, and session actions Terminal filters — reveal open Claude, Codex, and Terminal sessionsExpanded agent card — gathers the session's essential information in one cardCard index — shows that this is agent 333 in the complete sequenceAgent identity — provider icon, session title, and primary informationGoal — appears on the card when the agent has set oneTerminal source — the terminal running the agent; Ghostty hereLive-stream control — starts or pauses tracking the agent's changesSession activity — messages, subagent conversations, and file changesUser messages — only the first and last human messageAgent messages — only the first and last agent messageSubagent conversations — messages exchanged with spawned subagentsEdited files — number of files edited by the agentUncommitted files — changed files that have not entered a commitCommitted changes — number of changes already committedConversation actions — experimental copy optionsCopy human messages onlyCopy editable context — messages and changed-file summariesLast meaningful activity — working area and elapsed timeFork — separates the session into a new terminal flowCompact — condenses the conversation with optionOS's own mechanism
·
1 / 20

Terminal filters — reveal open Claude, Codex, and Terminal sessions

Terminal filters — reveal open Claude, Codex, and Terminal sessions
Twenty marked regions expose the compact agent card from identity through Compact.