optionOS optionOS

From ⌥A to close: everything a voice session collects while you talk

How do I use it?
Tags
  • VOICEVoice-driven production.
  • CAPTUREScreen, GIF, scrolling capture or clipboard signal.
  • ANNOTATIONMark, number, box or arrow tied to a speech moment.
  • WORKFLOWRepeatable working flow.

This guide was produced with the very method it describes: one 12-minute spoken session, rectangles drawn while talking, and captures taken along the way. Every visual below is interactive — hover or click the numbered targets and the explanation opens next to the target.

The 12-minute recording that produced this guide, shown in the transcript panel This guide's recordingTranscript contentDuration badge
This guide's recording
Everything on this page was narrated in a single voice session; the recording sits in the left list like any other conversation.
The recording that produced this guide sits in the left list.
Everything on this page was captured in a single 12-minute conversation; transcript, marks and visuals share one timeline.

The transcript view of the conversation that produced this guide; the content was narrated in a single voice session.

Start recording with ⌥A. The recording bar appears with the Capture Actions menu next to it; while you talk, everything is collected with timestamps.

optionOS recording bar and Capture Actions menu Recording barCapture Actions menuCapture ScreenCaptured content queueAdd NoteGet TextRecord GIF
Recording bar
While recording, everything you say is collected with timestamps.
This bar appears when you start recording with ⌥A.
While recording, every capture action lives in one menu, each with its shortcut.

The recording bar and the seven targets of the Capture Actions menu.

Capture a screen region when you need to show something visual; the capture lands in the conversation queue — hover for a preview, click to edit.

Preview and edit area of content captured into the conversation Captured content rowPreview areaEdit areaAgent info
Captured content row
Hover to open the preview.
Content captured during the conversation shows on this row.
Hover: preview. Click: edit. The bar below shows which agent and voice profile you are talking to.

Preview of captured content inside the conversation, the edit area and agent info.

In the editor the rectangle tool comes pre-selected: say "this part needs fixing", draw a rectangle, and the pause matches the drawing to that exact moment of speech.

Annotation toolbar: rectangle, number and undo Rectangle toolUndoNumber tool
Rectangle tool
Say 'this part needs fixing' and draw — the pause matches the rectangle to that exact spot in your speech.
The rectangle tool comes pre-selected; draw while saying 'this part'.
Draw rectangles, drop numbers if you prefer; remove a wrong mark by tapping it or with ⌘Z.

Editor toolbar: pre-selected rectangle tool, number tool and undo.

Drawing and image references in a transcript line Transcript lineRectangle referenceImage reference
Transcript line
Type, source and size appear as chips in the flow of speech.
Sources embed into your speech, right where you said them.
References embed into the flow of speech; a rectangle chip binds to the moment it was drawn, an image chip to the original file.

Type, source and drawing references embedded as chips in the transcript line.

Add Note attaches typed text to that moment of the conversation — use it sparingly, for proper nouns and words the transcriber gets wrong.

Add Note panel open Note panel
Note panel
Use it sparingly: proper nouns, mistranscribed words and text that isn't written anywhere. If the text can be copied, copying is enough.
Add Note attaches what you type to that moment of the conversation.
For special words only; Enter attaches the note.

The Add Note panel; typed text attaches to the speech timeline.

Use Get Text to pull text straight out of an image, and GIF Record to show anything that moves.

GIF recording panel over the selected region Recorded regionGIF panelStop / Cancel
Recorded region
The GIF watches the region you selected, not the whole screen.
The selected region is recorded.
The selected region is recorded; Stop finishes, Cancel discards.

The panel shown while a GIF records: duration, size, stop and cancel.

A short GIF of an agent working
You watch the GIF; the agent receives a frame digest of changed regions instead.

A real GIF captured during the session.

Changed-region frames of the GIF in the panel Copied GIF rowFrames badgeFrame digestA changed regionSecond changeMoved object
Copied GIF row
Every capture lands in the same queue; the GIF is no exception.
The copied GIF lands in the usual panel.
Only changed regions become frames — moving objects and changing text are captured one by one.

Changed-region frames produced from the GIF, with the Frames badge.

GIF preview and format selector GIF previewFormat selector
GIF preview
This area is for your own verification; the agent uses the frame digest.
You can watch the GIF yourself.
Watch the GIF to verify; switch what the agent receives from the selector.

GIF preview and the GIF/Frames format switch.

For time-based work, the Speech Context menu marks a start flag (press again to close the range) and Speech Source hands the agent the fully timestamped transcript — it finds the exact second you said "I clicked this and that happened".

Speech Context menu with start marking and speech source Speech ContextMark StartSpeech Source
Speech Context
The menu that turns your speech itself into data.
Time-based work lives under Speech Context.
Your speech is data; flags define ranges, the source carries exact time.

Speech Context menu: start marking and the timestamped speech source.

Scrolling Capture under Special Capture Scrolling Capture
Scrolling Capture
Long pages — landing pages, Twitter threads — become one tall image instead of two or three fragments.
Long content becomes a single image (⌥X).
A long page becomes one tall image instead of two or three fragments.

Scrolling Capture under Special Capture.

It is all one flow: start with ⌥A, talk, mark what you see, capture motion as GIF, text as text, special words as notes. When the session ends, the agent holds your speech, marks and visuals merged on a single timeline.