← Resources

How much faster is dictation than typing, after corrections

Short answer: the famous figure is 3x, and it is a real measurement of copying short phrases into an iPhone. It is not a measurement of writing. Once you subtract the time spent fixing what the recognizer got wrong, and accept that thinking of the words does not get any faster, the honest number for composing on a Mac lands closer to 1.2x or 1.3x on a whole session. That is still worth having. It is just an hour saved in a nine-hour week of writing, not six hours. This page shows where the 3x came from, what it actually measured, and how to work out your own number instead of borrowing anyone else's. For the background on how the tools themselves work, start with the guide to local speech-to-text on Mac.

Where the 3x comes from

Almost every page that quotes a speed multiple for dictation is quoting the same study, usually without naming it. It is Sherry Ruan, Jacob Wobbrock, Kenny Liou, Andrew Ng and James Landay, published in the Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies in December 2017. Thirty-two participants, sixteen working in English and sixteen in Mandarin. For English the result was 152.86 words per minute by speech against 52.24 by keyboard, which is where 2.93x comes from. The preprint was published under the title "Speech Is 3x Faster than Typing", and that title has done more work in the world than the paper has.

Three details about the study change what the number means, and I have not found a single page on the first page of results that mentions any of them.

The keyboard was an iPhone. Specifically the built-in iOS keyboard on an iPhone 6 Plus, at 52.24 words per minute. The study never tested a desktop keyboard, so the ratio is speech against thumb-typing on a phone. If you are sitting at a Mac with a full keyboard, the comparison you are reading about was never run.

The task was copying, not writing. Participants transcribed 120 prewritten phrases averaging 28.3 characters, with no punctuation and almost no capitalization. The authors are direct about the limits of that: they chose transcription because composition introduces confounds such as thinking time and word choice, they note that using transcription to draw conclusions about general text entry "might be suspect", and they describe their results throughout as upper-bound performance.

And the counter-number is inside the same paper. Ruan and colleagues' own literature review cites earlier work by Karat and colleagues in which users composing text by speech managed 7.8 words per minute, and transcribing by speech managed 13.6, against 32.5 words per minute for a keyboard-and-mouse method. The paper that gave the internet "3x faster" contains a citation in which speech composition was four times slower. Recognizers have improved out of all recognition since 1999, so I would not carry that 7.8 forward as a current figure. The point is the size of the gap between copying and composing, which nobody has closed by making the recognizer better, because the recognizer was never the bottleneck.

What each study actually measured

The single most useful thing you can do with this literature is stop reading the headline numbers and start reading what was in front of the person being timed.

StudyWhat it measuredResultWhat it does not tell you
Ruan et al. 201732 people copying 120 short phrases into an iPhone 6 PlusSpeech 152.86 wpm, iOS keyboard 52.24 wpm, ratio 2.93Anything about desktop keyboards, punctuation, or composing your own words
Dhakal et al. 2018136 million keystrokes from 168,000 people transcribing given sentencesMean 51.56 wpm, fastest tenth above 78 wpmHow fast a general working population types. The sample self-selected on a typing-test site, mean age 24.5
Yuan et al. 20062,438 real telephone conversations, word-aligned196 wpm including silences, 236 wpm excluding themHow fast you dictate. Conversation is not dictation, and the gap is large
Koester 200323 experienced speech-recognition users doing real work tasksNet 16.7 wpm. 37% of task time spent correcting. 7 of 18 users were slower with speechHow a modern recognizer performs. Accuracy averaged 85%, well below current models
Blackley et al. 202010 physicians writing the same clinical notes both ways, randomized orderDocumentation time similar between methods. Dictated notes longer and rated higher qualityMuch, with n=10. But it is a direct like-for-like test and it found no clear speed win
Zhou et al. 2018217 clinical notes from 144 clinicians, errors counted at three stages7.4% error rate in the raw draft, 0.4% after review, 0.3% once signedHow long the review took. It quantifies the cleanup, not its duration

The two numbers everyone quotes have no source

Most articles on this topic build their ratio from two figures: people speak at about 150 words per minute, and people type at about 40. Both are folklore. I went looking for the primary source behind each and neither has one.

The 150 is attributed almost everywhere to the National Center for Voice and Speech, with no paper, dataset or method named, and the pages it was taken from are no longer live. It is also contradicted by the best actual measurement of conversational English: Jiahong Yuan, Mark Liberman and Christopher Cieri, presented at Interspeech 2006, aligned 2,438 telephone conversations and found 196 words per minute counting silence and 236 with silence removed. If 150 were a measured conversational rate it would be far too low.

The 40 is worse, because there is a good study available and nobody cites it. Vivek Dhakal, Anna Feit, Per Ola Kristensson and Antti Oulasvirta collected 136 million keystrokes from 168,000 people for CHI 2018 and measured a mean of 51.56 words per minute, with the fastest tenth above 78. But that sample was recruited through a typing-test website, had a mean age of 24.5, and 72% had taken a typing course. It is a room full of people who enjoy typing tests, transcribing sentences they were given. So it is an upper bound too, in the same way the speech figure is, and the correct move is to use both as upper bounds rather than pretending one is a hard floor and the other a hard ceiling.

This matters more than it looks. A ratio built from two unsourced numbers pointing in convenient directions is not evidence, and the direction of the error is not random. Take the folklore pair and you get 3.75x. Take the two best measurements of the same quantities and you get something closer to 3x before you have subtracted a single correction.

What correction actually costs

This is the part the ranking pages skip, and it is the whole argument. Every raw words-per-minute figure prices your cleanup pass at zero. Three separate measurements say it is not zero.

The Stanford study gives the floor. Before any correction, initial speech transcription ran at 178.92 words per minute; after users fixed their errors it settled at 152.86. The authors state the ratio as 1.17, so correction took about 15% of raw throughput. Hold onto how favorable those conditions were: a quiet lab, five-word phrases, no punctuation to speak, and a recognizer being asked the easiest question in the field.

Heike Koester's 2003 study gives the ceiling. She watched 23 experienced speech-recognition users, all with at least six months of practice, complete real word-processing tasks on video. Correcting recognition errors took 37% of total task time. Each individual correction ran from 11 to 42 seconds and averaged 23. Net text entry came out at 16.7 words per minute. And the finding that should give any confident writer pause: comparing the 18 users who did both conditions, speech helped 11 of them and made the other 7 slower.

A 2020 CHI study by Foley and colleagues isolates the mechanism cleanly. Trials that contained no correction averaged 141.1 words per minute. Trials that contained one averaged 54.6. A single correction episode roughly halves your throughput for that stretch, because you stop producing, switch task, locate the error, fix it, and restart. The cost is not the keystrokes, it is the context switch.

So the honest range for correction overhead runs from 15% in the friendliest lab conditions ever measured to 37% among experienced users doing real work. Where you land is set almost entirely by your error rate, which is the subject of its own page: word error rate explained covers how the figure is calculated and why a benchmark number is not the number you will get at your desk.

Do the arithmetic

Nobody publishes the calculation, so here it is. Net throughput is your raw speaking rate divided by one plus your correction overhead:

net wpm = raw wpm / (1 + correction seconds per minute of speech / 60)

Work an example with numbers you can defend. Say you dictate at 150 words per minute and your recognizer has a 5% word error rate, so 150 words of speech leaves you roughly 7 or 8 errors. At Koester's average of 23 seconds per fix that is nearly three minutes of correction per minute of speech, which is catastrophic, and it tells you immediately that 23 seconds is the wrong constant for modern conditions. Her users were correcting by voice command against an 85% accurate recognizer in 2003. Correcting a word by hand in a modern editor is a few seconds, not 23.

So take 4 seconds per fix, which is a plausible cost for noticing a wrong word and retyping it. Seven errors at 4 seconds is 28 seconds of correction per minute of speech, giving 150 / (1 + 0.47), or about 102 net words per minute. Against a 51.56 wpm typist that is 2x, not 3x. Push the error rate to 10% and you get roughly 74 net wpm, or 1.4x. Push it to 20% and dictation loses to a fast typist entirely.

The break-even is the useful output here. On these assumptions, dictation stops beating a median typist somewhere around a 30% word error rate, and stops beating a top-decile typist at around 12%. That is why accuracy is a speed feature rather than a comfort feature, and why the gap between a good model and a mediocre one shows up in your calendar rather than in your transcript.

Two honest caveats about this arithmetic. The 4 seconds is my estimate, not a measured constant, and it is the number the whole calculation swings on. And the model assumes you fix errors as you go; if you dictate the whole draft and clean up afterwards the total may be lower, because you avoid the context switch that Foley's data suggests is the expensive part.

Thinking time is the real ceiling

Even if correction were free, you would not get 3x, because typing is not the only thing you do when you write. John Gould and Stephen Boies published a study in Science in 1978 comparing writing, dictating and speaking letters. Their central finding has aged better than anything else in this literature: planning takes about two-thirds of composition time, and it takes that share regardless of which method you use. They also found that novices learned to dictate at the speed and quality of their writing within a few hours, and that experienced dictators were only 0 to 25% faster than novices, which is a strong hint that practice is not the missing variable either.

Put that together with an input speedup and the result is deflating. Suppose a writing session is 100 minutes, and following Gould, 67 of those are thinking and 33 are getting words down. Make the getting-words-down part three times faster and it takes 11 minutes instead of 33. Total session: 78 minutes. That is 1.28x, from a 3x input improvement. Apply the more realistic 2x from the section above and you get 84 minutes, or 1.19x.

This single frame reconciles everything else on this page. It explains why the lab measures 3x and the clinic measures nothing much. It explains why the writers who try dictation and report it life-changing are usually the ones with the words already in their head, and why the ones who bounce off it are usually composing as they go. And it explains why short messages show a bigger multiple than long-form writing does: a text message is nearly pure transcription, so it is close to the lab condition, while an essay is mostly planning.

What happens outside the lab

Clinical documentation is the one field that has measured this at scale on real work, because hospitals have both the budget and the motive. The results are much less flattering than the consumer literature, and the pages quoting 3x do not mention them.

Suzanne Blackley and colleagues ran the cleanest test in the International Journal of Medical Informatics in 2020. Ten physicians, each documenting simulated encounters twice, once dictated and once typed, in randomized order. Their finding: documentation time was similar between methods, with dictation slightly ahead, and their conclusion states plainly that whether dictation is objectively faster than typing remains unclear. The dictated notes were substantially longer, 320.6 words against 180.8, and were rated higher quality. Interestingly the typed notes carried more uncorrected errors, 2.9 per note against 1.5, mostly misspellings. Ten physicians is a small study and I would not hang a decision on it alone, but it is a direct like-for-like comparison and it did not find the multiple.

At larger scale, a 2025 cross-sectional study in JAMA Network Open looked at speech-recognition use across clinicians at a rural US health system and found each 1% rise in usage associated with a 0.25% rise in lines documented per hour. That is an association in one health system rather than a controlled speedup, and it should not be extrapolated to a headline multiple, but the direction and magnitude are consistent with everything above: a real gain, in the tens of percent, not a tripling.

The error side has its own good measurement. Li Zhou and colleagues examined 217 clinical notes from 144 clinicians across two health systems for JAMA Network Open in 2018 and counted errors at three stages. The raw speech-recognition draft had a 7.4% error rate. After a transcriptionist reviewed it, 0.4%. After the physician signed it, 0.3%. That is the clearest published picture of what cleanup buys: review removed roughly 94% of the errors. It is also the clearest picture of why draft speed and finished speed are two different measurements, since the draft in that study was never usable as it came out.

Where the time actually goes: dead air

There is one number in the Stanford study that nobody quotes and that I think is the most practically useful thing in it. Of the total time in a dictation trial, only 39.2% was spent actually speaking. Another 31.5% was lost to user delay, meaning the participant had activated the microphone and was not yet talking, or had finished talking and had not yet pressed done.

Nearly a third of the interaction was dead air. And this was people reading a phrase off a card. In real writing, where you genuinely stop to think, the share is going to be higher, not lower.

That is a design problem, not a speech problem, and it is the one place where how the app is built changes your actual throughput. An always-on recognizer is listening through every one of those pauses. It pays for your hesitation twice, once in latency while it decides you have stopped, and once in whatever it transcribes from the room while you were thinking. Systems that solve this with a silence timeout, as Apple's built-in dictation does by stopping after 30 seconds of quiet, solve it by punishing exactly the pause you needed.

Push-to-talk inverts the arrangement. You hold a key while you talk and release it when you are done, so the boundary is set by your hand rather than inferred from your silence. Thinking costs nothing, because the microphone is not open while you do it. The trade is a keypress and the discipline of remembering it. I have written up the full comparison separately in push-to-talk versus always-on dictation, including where always-on genuinely wins.

This is also the honest reason the throughput literature is dated in one specific respect. Every speed figure quoted on this page was produced against a server-based recognizer, and Ruan and colleagues note in their own limitations that embedded systems have lower latency and better reliability. Eight years on, on-device Mac dictation is ordinary and the round-trip is gone, but nobody has redone the arithmetic for local inference. If anything the published figures now understate the input side, which makes the thinking-time ceiling above matter more rather than less.

When dictation genuinely wins, and when it does not

The multiple is not a property of dictation. It is a property of the task you point it at, and the pattern across all the evidence above is consistent.

It is genuinely faster when:

  • You already know what you want to say. Replies, notes, briefs, anything closer to transcription than composition. This is the lab condition, and the lab condition is real work for a lot of people.
  • The output is narrative prose. Long unbroken stretches of ordinary sentences, where punctuation is predictable and structure is linear.
  • Your hands are the constraint. If typing hurts, the comparison stops being about speed. More on that below.
  • You are producing a rough first draft on purpose. Blackley's finding that dictated notes came out substantially longer is a feature if volume is what you want out of a first pass.

It is slower, or not worth it, when:

  • The exact characters matter. Code, formulae, identifiers, anything where you would spend longer speaking the punctuation than typing it. The punctuation command list shows how much you have to say out loud to get a well-formed paragraph.
  • The document is mostly structure. Templates, forms and tables are navigation and selection work, and dictation does not help with either.
  • You are composing as you go. Gould's two-thirds is the whole reason. If you are finding the words while speaking, you are paying for dead air.
  • You can be overheard, or the room is noisy. One costs you privacy, the other costs you accuracy, and accuracy costs you time.
  • You are editing rather than writing. Restructuring existing text is reading and deciding, not producing.

Measure your own number

Every page on this subject, this one included, is giving you an average from someone else's conditions. Your own figure takes about forty minutes to establish and it is the only one that should influence what you do. The trick is to measure finished words rather than produced words, because that is the quantity you actually get paid in.

  1. Pick one real task you do often. Not a typing test. A genuine email, note or section of a document, of the kind you write weekly.
  2. Do it typed, timed, to finished quality. Stop the clock when you would actually send it. Count the words.
  3. Do a comparable one dictated, timed the same way. Include the cleanup pass in the time. If you would not send it as it stands, the clock is still running.
  4. Divide finished words by total minutes for each. That ratio is your number, and it already contains your error rate, your thinking time, your accent, your room and your subject matter, none of which any study can give you.
  5. Repeat once a week for three weeks. Dictation has a learning curve that typing does not, so a single trial will understate it. Gould's finding suggests the curve is shorter than people expect, but it is not zero.

If your ratio comes out near 1.0, you are probably composing rather than transcribing, and the honest conclusion is that dictation is not your bottleneck. If it comes out above 1.5, you have found real time and the only remaining question is your error rate. And if the two numbers are close but dictating leaves you less tired at the end of the day, that is a legitimate reason to switch that has nothing to do with speed.

The other reason people switch, honestly assessed

A good share of the people who move to dictation are not chasing speed at all. They are trying to keep working with wrists, shoulders or a neck that have stopped cooperating, and for them a ratio of 1.0 would be a fine outcome. The evidence here is real, but it is more qualified than the marketing suggests, and it cuts both ways.

Birgit Juul-Kristensen and colleagues measured muscle activity directly, publishing in Ergonomics in 2004. Speech recognition reduced static muscle activity in the forearm, neck and to some extent the shoulder, with longer pauses in muscle activity in the forearm and shoulder. But it increased static activity in the voice muscles, and their own recommendation was to use speech recognition as a supplementary tool alongside a keyboard and mouse rather than as a replacement.

Two other findings belong here for balance. Elsbeth de Korte and Peter van Lingen, in Applied Ergonomics in 2006, found improved wrist, forearm, upper arm, shoulder and neck postures after a six-week training period, and also that productivity decreased for most subjects. And Olson and colleagues documented five patients with existing upper-limb strain injuries who developed muscle tension dysphonia shortly after taking up speech recognition, all of whom had normal voices in ordinary speech. Five retrospective cases is not a rate, and I am not going to present it as one, but "you can hurt your voice this way" is a real thing that exists and is absent from every enthusiastic article on the subject.

I should be equally clear about what is missing: I could not find a single randomized trial of dictation as a treatment or prevention for repetitive strain injury. The case rests on the mechanical evidence above, which is genuine, rather than on outcome trials, which do not appear to have been run. The dictation for RSI page goes into the practical setup in more detail.

How I approach this

I make Parakeety, a one-person Mac dictation app, so I had an obvious commercial reason to write "dictation is 3x faster" and stop there. I have argued the opposite, because the 3x claim does not survive contact with the paper it comes from, and because a reader who buys a dictation app expecting to triple their output will be disappointed by a tool that genuinely helps them. Every figure above is attributed to a named study with a date, and where I could not find a primary source, as with the ubiquitous 150 and 40 words per minute, I have said so instead of repeating it.

Three things I want to be straight about. The net-throughput arithmetic in this piece is a model, not a measurement, and it turns on a 4-second-per-correction estimate that is mine rather than published. Several of the studies are old, and 2003-era recognizers were far worse than current ones, which means the correction overhead I quote as a ceiling is almost certainly too pessimistic today. And I have not run a controlled trial of my own app against typing, so there is no Parakeety number in that table, because there is no study behind it. The self-test above is what I would do in place of trusting anyone's marketing, including mine.

FAQ

Am I actually going to write faster if I switch to dictation?
Probably yes, but by less than you have been told, and the gain depends on what you write. The widely quoted 3x comes from a lab study where people copied short prewritten phrases into an iPhone, which is close to the best case dictation will ever have. Real writing adds three costs that study did not measure: you have to think of the words, you have to punctuate them, and you have to fix what the recognizer got wrong. Thinking is the big one. Gould and Boies measured planning at roughly two-thirds of composition time back in 1978, and no input method touches that share. Work the arithmetic through and a 3x input speedup lands nearer 1.3x on a whole writing session. That is a real gain, it is just not the one on the box.
Is speech-to-text really 3x faster than typing?
The number is real and it is also narrower than it sounds. It comes from Ruan and colleagues at Stanford, published in 2017, who measured 152.86 words per minute for English speech against 52.24 for the keyboard, a ratio of 2.93. But the keyboard in that comparison was the iOS onscreen keyboard on an iPhone 6 Plus, and the study never tested a desktop keyboard at all. The task was copying 120 short phrases that averaged 28 characters, with no punctuation, which the authors themselves describe as upper-bound performance. So 3x is a sound measurement of dictating a text message on a phone. It was never a measurement of writing on a computer, and most pages quoting it do not say so.
Is dictation still faster once you account for editing?
Usually yes, but the margin narrows a lot and it can invert. The best case is the Stanford study itself, where correction cost about 15% of raw throughput: initial transcription ran at 178.92 words per minute and dropped to 152.86 after users fixed their errors. That was a quiet lab, five-word phrases and no punctuation. The worst measured case is Koester in 2003, who watched 23 experienced speech-recognition users do real work tasks and found 37% of total task time went on correcting errors, with each fix taking 11 to 42 seconds. In that study speech made 7 of the 18 compared users slower rather than faster. Your own figure sits somewhere between those two, and it is set mostly by your error rate.
How many words per minute can you dictate?
Measured dictation of short messages came out at about 153 words per minute in lab conditions, including the time spent correcting. That is well below how fast people actually talk: on 2,438 real telephone conversations, Yuan, Liberman and Cieri measured 196 words per minute counting the silences and 236 with the silences stripped out. The gap between 236 and 153 is the honest cost of dictating rather than chatting, and it is mostly hesitation and correction rather than slower speech. For composition rather than copying, the only figures available are much lower again. Karat and colleagues measured 7.8 words per minute for composing by speech in 1999, against 32.5 for keyboard and mouse, though recognizers have improved enormously since then.
When is typing still better than dictation?
Four situations, and none of them are about the recognizer being bad. Code, formulae and anything where the exact characters matter, because you spend longer speaking the punctuation than typing it. Heavily structured or templated documents, where most of the work is navigation and selection rather than producing prose. Anywhere you can be overheard or where the room is noisy, since both your privacy and your error rate suffer. And structural editing, which is reading, deciding and moving text, not generating it. There is also a fifth that has nothing to do with the task: if you are still finding the words while you speak, the dead air is the cost, and a keyboard lets you pause for free.

Try it

If the arithmetic above persuaded you of anything, it should be that the two levers on your real speed are your error rate and your dead air, not your speaking rate. Parakeety is built around exactly those two. It runs NVIDIA's Parakeet TDT 0.6B v3 on the Apple Neural Engine, which posts a 6.34% word error rate on the Open ASR Leaderboard, and that number is a benchmark on clean read audio rather than a promise about your desk, your accent or your room. And it is push-to-talk by design: hold the section key while you speak, release when you have finished, so every pause to think costs you nothing and nothing is captured while you are silent. Everything runs on your Mac, so there is no cloud round-trip in the loop either. It needs Apple Silicon and macOS 14 or later. There is a free 7-day trial with no card, which is long enough to run the self-test three times and get your own number before you decide.

Try Parakeety free →