optionOS optionOS

How I work: I speak instead of typing and mark ambiguity on the visual

How do I work?
Tags
  • VOICEVoice-driven production.
  • AI-DEVELOPMENTAI development practice.
  • WORKFLOWRepeatable working flow.

The content on this page did not begin as a line-by-line written prompt. I made five voice recordings, captured screens at the important moments, and marked 63 regions while speaking. The raw speaking time was 14 minutes 54 seconds in total. The AI session was gpt-5.6-sol · xhigh · fast.

The measured input cost in this example is time: 14:54 of speech. No monetary amount appears in the visuals, so I do not claim one.

Five starred voice recordings and the single AI prompt assembled from those spoken recordings Opening recordings — starred voice conversations that begin the promptContinuation recordings — the other starred conversations collected for the same taskAssembled prompt — one agent packet containing the raw speech and visual context
·
1 / 3

Opening recordings — starred voice conversations that begin the prompt

Opening recordings — starred voice conversations that begin the prompt
Five recordings → one prompt → Dictation → Cockpit or Codex. Click through the nested visual path.

Seven screens and 63 marked regions share one tour; the hover door also names the real app's hover behavior.

What happens if the transcript is wrong? A twenty-minute conversation can contain a mistaken word, and my thinking can move from A to B to C. When later sentences correct an earlier mistake, the AI can read the complete conversation together with the marked visuals and desired outcome instead of treating one word as the whole instruction. I preserve intent rather than waiting for a perfect transcript.

I sometimes end the conversation with this sentence:

> I am not sure whether we should do it this way; choose the correct approach.

This does not delegate the desired result to the agent. I provide the outcome, the problems I noticed, and the options I considered; the AI chooses whether the technical method should be A, B, or C after reading the real system.

My working method:

1. Speak at the natural speed of thought. 2. Capture the screen and mark the region as soon as I say “this part.” 3. Add another recording before pasting when I remember something else. 4. Fix the desired outcome in the final sentence and let the system's reality determine the technical method. 5. Deliver the assembled packet to the agent. 6. Follow messages, subagents, and file changes in Cockpit.

Open Codex session in Ghostty and an optionOS transcript packet delivered entirely through speech Codex session — the Ghostty tab that received the conversationTranscript packet — spoken content delivered without manually typed companion text
·
1 / 2

Codex session — the Ghostty tab that received the conversation

Codex session — the Ghostty tab that received the conversation
The Codex session and its speech-delivered transcript packet share one screen.

This is not a “make the transcript perfect” method. The point is to keep speech, the marked visual, app context, and the final decision in one packet. This page is evidence of the method itself: speech was the source, the visuals lived inside the conversation, and the published content was produced from that context.