optionOS optionOS
← All updates

Watch your speech live, see its tags, step inside the image you captured

Speech is now visible end to end on the Dictation side: the live bar watches the speech and its context, the transformed panel shows the rectangles you drew and the images with their tags, and hovering a captured imag…

optionOS

Speech is now visible end to end on the Dictation side: the live bar watches the speech and its context, the transformed panel shows the rectangles you drew and the images with their tags, and hovering a captured image steps inside the content you annotated earlier. The recording bar and the Dictation panel were renewed in the same release.

From the live bar to four panels; from the image queue into annotated content

You watch the speech from the live bar; the context shows which app it was started in, and hovering the flow indicator opens the four panels. The panel holds the transformed form of your speech: rectangles you draw while speaking appear as tags in the transcript, images as references with the moment they were saved. You read the content in three views: Speech is your plain words, Events is what goes to the agent, Context is the human form. Capture actions open inside the conversation; the most used is Capture Screen — select an area with Shift and it lands tagged into your speech. Captured images drop into the queue: the app they were captured in, the form they will reach the agent in and their size sit on one row; press the form badge to send the image as text or as a picture, and when size cannot be counted as tokens it falls back to kilobytes. Hovering the row opens the final step: the content you annotated earlier stands in front of you, picture inside picture, with its rectangles and numbers.

Dictation live bar with the transformed conversation panel, view tabs and capture actions marked Speech watch barConversation contextFlow indicatorTransformed conversation panelRectangle tag inside the speechRectangle tag in the transcriptImage referenceView pickerEvents viewSpeech viewConversation receiverCapture actionsCapture ScreenCaptured image queueApp it was captured inForm sent to the agentSize indicator
·
Live bar

The bar watches your speech; the context shows which app it was started in. Hovering the flow indicator opens the four panels.

We watch the speech from here: context on the left, the flow indicator beside it. Hovering the flow opens four panels.
Live bar at 1–3; transformed speech and tags at 4–7; the final step walks from the image queue into the annotated content.

This card is two nested visuals: hovering the queue row steps into the preview of previously annotated Cockpit content.

The recording bar tracks context live; speeches paste into a single session

As you speak and capture, the context list grows; the system actively tracks the context and updates itself while you talk. The red waveform shows your speech is being heard; the small indicator says it is being transcribed in the background. While an old speech waits unpasted you can record a new one: they wait separately, and pasting merges them into a single session.

Recording bar with the context list, waveform, background transcription and pending speech row marked Growing context listActive context trackingWaveformBackground transcriptionPending speech row
·
Context list

Every capture and note joins the conversation context; the list updates itself while you keep talking.

As you speak and capture, this list grows; the system actively tracks the context.
Growing context list and live tracking at 1–2; waveform at 3; background transcription at 4; pending speech row at 5.

The Dictation panel: tag, relate, listen, search

Your conversations list in the panel with their durations and tags; tag with ⌘D or right click. Conversations merged by pasting count as related and their bond is recorded. Read your own conversation on the right; click anywhere in the text and the audio starts playing from that moment, light grey regions are where you stayed silent. For Whisper users our own algorithm cleans mistranscribed subtitle-like patterns; the panel names what was hidden and one click brings it back. Your total speech record sits in the corner of the panel; the list updates as you type a search, and typing guide then pressing Tab keeps every guide-tagged conversation in the panel.

Dictation panel with the conversation list, tags, audio playback, search and suggestions marked Conversation listConversation durationTagsRelated conversationsConversation detailChain it belongs toHidden patternAudio info barSilent areasTotal speech recordSearchSuggestions
·
Conversation list

Tag with ⌘D or right click. I tag conversations and hand them to agents as production material; these pages were written that way.

Your conversations list here; duration and tags sit on the row.
List, duration, tags and relations at 1–4; detail, hidden pattern and audio at 5–9; total record, search and suggestions at 10–12.