Every screen relates to the others; this page merges them into one structure. The flow below is not separate images: it is a single journey that tells what a person does from start to finish, with the screens linked to each other. Open your terminal, start the speech with ⌥A and walk the flow click by click.
11 screens in one flow: start the speech, watch, copy, record a GIF, finish, preview, merge, edit and paste to the agent.
The last step of every screen is the door to the next: click or hover the door region and you step into the picture inside the picture.
Small notes that live inside the flow: there is also a cursor follower — while speaking you see the recording state next to your cursor; no photo needed, you will run into it. Content is copied as it is pulled; the photos and clipboard update inside the speech as you copy, and the system sees it all as one session.
Your whole speech archive lives at ⌥H. The Dictation History panel collects your speeches on one searchable, tagged, starred and badged surface; since audio is saved every time, no speech is lost even if the app breaks. The second flow below walks that panel end to end.
From ⌥H to the audio bar: search, star-versus-tag, badges, multi-select, including, recording safety and listening from text.
The badge in the panel and the badge in the speech are the same badge; the system works coherently.
The trick for not losing yourself in a long speech: make it count. I have the AI enumerate the tasks — so my own speech is not forgotten. At the end it hands everything back in one pass. When I paste the output I also make the agent tell me what it understood; every round sharpens itself. The pattern I use:
> Enumerate every task as [n/m] and proceed; change the counter when the goal changes. At the end, hand everything back to me in one pass.