Voice captures my thinking. AI helps me articulate it. I remain responsible for the result.
For decades, the keyboard and mouse have defined how most of us interact with desktop computers. Even as computers became millions of times faster, our primary way of giving them information barely changed.
That leaves us with a basic bottleneck: ourselves.
In my previous article, “10 Bits per Second,” I wrote about the mismatch between the speed of computers and the speed of human thought. We cannot type at the speed of thought. We cannot even speak at the speed of thought.
But speaking is still much faster than typing.
The original prompt for this article was roughly 1,000 words. I spoke it in about six or seven minutes. At a typical keyboard speed of 40 to 50 words per minute, simply typing the same material would have taken around 20 to 25 minutes—and that assumes I already knew exactly what I wanted to say.
This does not mean voice made the entire article four times faster to produce. Writing is not just entering words. It includes reviewing the argument, checking facts, restructuring sections and rewriting sentences.
Voice reduced the most mechanical part: manually digitizing my thoughts.
That gave me more time for the parts that mattered.
This is not my first attempt at voice
My first experiments with voice input go back to the Classic Mac OS era. Apple was doing some remarkably advanced work at the time, including PlainTalk and Chinese-language dictation tools.
My Chinese typing was terrible, so speaking seemed like an obvious alternative. It kind of worked. That was the problem with early voice technology: it worked just well enough to make you believe it should be useful, but not well enough to trust.
I made a more serious attempt around 2011, while writing my MBA thesis. I used Dragon Dictate, then one of the leading consumer speech-recognition products. It helped enormously with writing long essays and was certainly faster than typing.
But whatever time I gained while speaking, I often lost while editing. I had to watch the transcription closely, correct recognition errors and restructure sentences that made sense when spoken but not when written.
Dictation gave me more words, but not necessarily a better document.
That was the missing piece—and it explains why the current generation feels different.
From transcription to understanding
OpenAI released Whisper in 2022, trained on 680,000 hours of multilingual audio and transcripts collected from the internet. It was an important step in making robust, multilingual transcription widely available.
There is an interesting story behind its development. OpenAI’s published paper describes an effort to create a robust, general-purpose speech-recognition system. Separately, The New York Times reported that OpenAI developed Whisper partly to transcribe YouTube videos, creating more conversational text for training GPT-4.
Whatever the original motivation, Whisper helped make high-quality transcription accessible. Since then, the technology has continued to improve, especially for widely used languages such as English and Mandarin.
But better transcription is only half the change.
Traditional dictation expected me to speak as if I were writing. I had to formulate each sentence, choose the right words, dictate punctuation and then correct the mistakes. The keyboard disappeared, but most of the cognitive work remained.
An LLM changes that.
I can now speak naturally. I can hesitate, repeat myself, wander onto a related subject and correct myself halfway through a sentence. The transcript may be messy, but the model can find the argument underneath it, remove repetition and help organize it into a sensible draft.
I do not consider writing my natural medium. Translating my thoughts into polished prose has never been my strength. Voice lets me express those thoughts in a form that feels natural, and AI helps me articulate them in writing.
That, to me, is the breakthrough.
Saving time is not the real objective
AI is often presented as a way to finish the same work faster: write an article in ten minutes, answer an email in ten seconds or produce a presentation without really making one.
That framing makes AI sound like a shortcut—a way to cheat, be lazy or avoid doing the work.
I see it differently.
The point is not merely to compress the time required to produce something. It is to reduce the tedious parts of the process so that we can spend more time on the interesting parts: examining an argument, finding weaknesses, considering other views and deciding what we actually believe.
The 15 or 20 minutes I saved by speaking the original prompt could have been used to publish the article sooner. But a better use of that time is to review the argument, check the history, remove unsupported claims and rewrite the weaker sections.
Readers rarely remember how efficiently you typed an article. They remember whether it contained an idea worth thinking about.
Grammar, tone and structure still matter because they help communicate the idea. But perfect grammar cannot rescue an empty argument. AI can help with expression, leaving us more time for meaning and judgment.
Five ways to make voice part of your workflow
1. Start with longer prompts
Typing is often easier for a short question or precise correction. Use voice when you need to explain a complicated situation, provide background, develop an argument or think through a problem.
A simple rule: if the prompt will take more than a few sentences, try speaking it.
2. Talk—don’t dictate
Do not carefully compose every sentence in your head before saying it. That preserves the bottleneck you are trying to remove.
Talk about the subject as if you were explaining it to a colleague. Include examples, relevant background and uncertainty. If you make a mistake, correct yourself aloud and continue.
You are providing context, not producing the final text.
3. Tell the AI what to do with the transcript
Give the AI a clear instruction after you finish speaking. For example:
- “Turn this into a concise email.”
- “Organize this into a well-argued article.”
- “Remove repetition but preserve my meaning and tone.”
- “Identify weaknesses or unsupported claims.”
- “Ask me about anything that is unclear.”
The transcript contains your thinking; the instruction defines the job.
4. Make good audio readily available
Use AirPods, a headset or a conferencing-quality microphone—and make sure it is always within reach.
At my desk, I use an old Poly USB conference speakerphone. When I am out, I use AirPods Pro. The best microphone is not necessarily the most expensive one. It is the one you can use immediately.
5. Remove the friction
Map transcription to a dedicated key or keyboard shortcut. Some people even use a foot pedal.
If starting voice input requires opening an application and searching through menus, you will keep typing. If it takes one press, you are much more likely to speak.
A division of labor that works for me
I do not think of AI as taking over my writing. I think of it as support.
In the past, I had assistants and colleagues who helped with research, groundwork and early drafts. They might suggest an idea, organize the material or find a clearer way to present something. I would review the work, change it and decide what I wanted to use.
AI plays a similar role in my current workflow.
I am a geek and a nerd—a computer scientist by training, with an MBA. Writing has never been my forte. I can often explain an idea more naturally than I can turn it into polished prose.
Voice lets me get the idea out. AI helps me organize and articulate it. I then review the result until it says what I actually mean.
For me, the division of labor is becoming clear:
Voice for expression. AI for articulation. Human judgment for the final result.
This is not a grand theory about how everyone should write. It is simply a workflow that has been tremendously useful for me.
I still do the reviewing, editing and deciding. I just spend less time manually typing everything out—and more time on the parts of the work I find interesting.
That feels like a pretty good trade.
Leave a Comment