The content on this page did not begin as a line-by-line written prompt. I made five voice recordings, captured screens at the important moments, and marked 63 regions while speaking. The raw speaking time was 14 minutes 54 seconds in total. The AI session was gpt-5.6-sol · xhigh · fast.
The measured input cost in this example is time: 14:54 of speech. No monetary amount appears in the visuals, so I do not claim one.
Five recordings → one prompt → Dictation → Cockpit or Codex. Click through the nested visual path.
Seven screens and 63 marked regions share one tour; the hover door also names the real app's hover behavior.
What happens if the transcript is wrong? A twenty-minute conversation can contain a mistaken word, and my thinking can move from A to B to C. When later sentences correct an earlier mistake, the AI can read the complete conversation together with the marked visuals and desired outcome instead of treating one word as the whole instruction. I preserve intent rather than waiting for a perfect transcript.
I sometimes end the conversation with this sentence:
> I am not sure whether we should do it this way; choose the correct approach.
This does not delegate the desired result to the agent. I provide the outcome, the problems I noticed, and the options I considered; the AI chooses whether the technical method should be A, B, or C after reading the real system.
My working method:
1. Speak at the natural speed of thought.
2. Capture the screen and mark the region as soon as I say “this part.”
3. Add another recording before pasting when I remember something else.
4. Fix the desired outcome in the final sentence and let the system's reality determine the technical method.
5. Deliver the assembled packet to the agent.
6. Follow messages, subagents, and file changes in Cockpit.
The Codex session and its speech-delivered transcript packet share one screen.
This is not a “make the transcript perfect” method. The point is to keep speech, the marked visual, app context, and the final decision in one packet. This page is evidence of the method itself: speech was the source, the visuals lived inside the conversation, and the published content was produced from that context.