Richer, and rarer.
What the published research actually says about asking people to speak their answers instead of typing them — including the part the vendors in this category leave out.
14 cited findings · 8 with DOIs · every number checked against the primary source
In one paragraph
Asking an open question out loud gets you two to three times as many words, in less than half the respondent’s time, on topics a text box does not surface. It also gets you those answers from fewer people: every study that measured completion found break-off roughly doubled. The gain is real, the cost is real, and most of the cost turns out to be the recorder rather than the act of speaking. This paper is the evidence for all three claims, and what follows from them.
A typed open-ended answer averages two words
Antoun, C., Couper, M.P. & Conrad, F.G. (2017) ran a randomised crossover on a probability panel of 1,390 people. Open-ended answers averaged ~2.1 words on smartphone, ~2.0 on PC. The device barely mattered — the phone and the desktop produced almost the same thing.
This is the number most survey design is quietly built around. Researchers ask an open question, get “good”, “fine”, “nothing”, and conclude that open questions do not work. What actually happened is that writing is expensive and nobody accounted for the price.
People do not give short answers because they have nothing to say. They give short answers because the format charged them too much to say it.
Two to three times the words, in half the time
Höhne, J.K., Gavras, K. & Claassen, J. (2024), N=1,001 across four topics, all p<.001: 22 words typed → 54–56 spoken. Revilla, M. & Couper, M.P. (2026) measured the cost to the respondent at the same time — 27 seconds spoken vs 61 seconds typed. Longer answers, less of their time. That combination is unusual enough to be worth stating twice.
It is not the same answer, only longer
The finding that changes how you read the rest: Gavras, K., Höhne, J.K., Blom, A.G. & Schoen, H. (2022), N=2,402, found the topic overlap between spoken and written answers: 30–70%. Speaking is not a more generous text box. It is a different instrument, and it collects things the text box never gets.
And now the part that gets left out
Every study above that also measured completion found the same thing, in the same direction: voice answers are rarer.
- Höhne, Gavras & Claassen (2024), N=1,001 — break-off 24% typing, 51% speaking
- Gavras et al. (2022), N=2,402 — dropout 13% typing, 45% speaking
- Revilla & Couper (2021) tried fixing it with clearer instructions. Item non-response stayed near 40%.
Any vendor telling you that switching to voice raises your response rate is describing something the literature does not contain. The honest claim is about depth per respondent, never headcount.
Most of that cost is the recorder, not the speaking
Revilla, M., Couper, M.P., Bosch, O.J. & Asensio, M. (2020) tested two implementations of voice input in the same study, on the same people, for the same task. One lost about 60% of answers. The other lost under 5%. ~60% failed to answer with one voice input · under 5% with the other.
A twelvefold difference between two versions of the same idea is not a fact about human beings. It is a fact about software. People will speak; they will not fight a bad record button — and the literature’s break-off figures were collected on instruments nobody designed for this.
That is the single most actionable finding here, and it is why the rest of this paper is about design rather than about persuasion.
Against video, voice wins decisively
Claassen, J., Lenzner, T., Höhne, J.K. & Ziller, C. (2026) asked 1,992 people, randomised, about the same survey with only the answer format changed: 67% would type · 47% would speak · 38% would appear on camera. χ²(9)=137.11, p<.001.
The reasons matter more than the ranking. Privacy is cited by 41% of those refusing video and 7.7% of those refusing voice. “Uncomfortableness” is 19% for video and essentially absent for voice. Offered a real camera in a live study, 202 of 302 people declined — 51% because they dislike seeing themselves, 48% because they were not camera-ready.
And the camera does not only cost you respondents. Shockley et al. (2021) manipulated camera use across 1,408 observations and found it made people measurably less likely to speak up — an effect strongest for women and for newer employees. It changes what the people who stay are willing to say.
Six design consequences
- Never voice-only. A text field under every record button costs you nothing and recovers the people the break-off figures are made of.
- Spend your effort on the recorder. One tap, no account, no app, no permissions maze. That is where the twelvefold difference lives.
- Say what it costs before they start, and be right. Song & Schwarz (2008) found the same instructions judged to take 15.1 minutes rather than 8.2 purely because they were harder to read. Perceived effort is doing the work here, not real effort.
- Make it shorter. Edwards et al. (2009), a Cochrane review of 481 randomised trials, puts shortening the questionnaire at an odds ratio of 1.64 — one of the few interventions that reliably works.
- Save every answer as it is given. If a half-finished response is worth nothing to you, you have thrown away the majority of what a harder format produces.
- Do not show the speaker their transcript. Reading it back converts a spoken answer into a written one, and you lose the thing you changed format to get.
Statistics in this category that did not survive checking
Several of these appear in competitors’ current marketing. One is not in the report its own publisher cites for it. They are listed so that nobody — including us — repeats them in good faith.
“Switching to voice raises your response rate.”
The opposite is measured. Every study finding voice answers richer also found them rarer — break-off 24% → 51% (Höhne et al. 2024), dropout 13% → 45% (Gavras et al. 2022). Better instructions did not fix it (Revilla & Couper 2021). Claim depth per respondent, never headcount.
“Voice answers are more honest.”
Schober et al. (2015) found text produced MORE disclosure of socially undesirable behaviour. The defensible claim is "less edited", which is what the researchers themselves say.
“Surveys longer than 12 minutes see a sharp drop-off.”
No primary source states 12. SurveyMonkey’s own published threshold is 7–8 minutes.
“Viewers retain 95% of a video versus 10% of text.”
A known marketing myth. Vendor blogs citing vendor blogs, no primary study.
“People spend 39% of a video call looking at themselves.”
Attributed to a paper whose abstract does not state it, on a women-only sample.
“The Hawthorne effect explains why people behave differently when observed.”
Paradis & Sutkin (2016): "Evidence of a Hawthorne Effect is scant, and amounts to little more than a good story." Say evaluation apprehension instead.
“A progress bar improves completion.”
Villar, Callegaro & Yang (2013), 32 experiments: constant bars have no significant effect (p=.365) and a bar that starts slow and speeds up makes break-off WORSE. Show honest time remaining, and never flatter the early questions.
“Low response rates make traditional surveys inaccurate.”
Groves (2006) found the correlation between nonresponse rate and nonresponse bias is only 0.33. The supportable argument is that they are expensive and slow, not that they are wrong.
Every source, and what it does not prove
A typed open-ended answer averages about two words — on a phone and on a desktop alike.
~2.1 words on smartphone, ~2.0 on PC
Antoun, C., Couper, M.P. & Conrad, F.G. (2017). Effects of Mobile versus PC Web on Survey Response Quality. Public Opinion Quarterly 81(S1), 280–306 · doi:10.1093/poq/nfw088
Randomised crossover on a probability panel, n=1,390.
The same question answered out loud produces two to three times as many words.
22 words typed → 54–56 spoken
Höhne, J.K., Gavras, K. & Claassen, J. (2024). Typing or Speaking? Comparing Text and Voice Answers to Open Questions on Sensitive Topics in Smartphone Surveys. Social Science Computer Review 42(4), 1066–1085 · doi:10.1177/08944393231160961
N=1,001, four topics, all p<.001.
Speaking does not produce a longer version of the same answer. It produces a different one.
topic overlap between spoken and written answers: 30–70%
Gavras, K., Höhne, J.K., Blom, A.G. & Schoen, H. (2022). Innovating the collection of open-ended answers: the linguistic and content characteristics of written and oral answers to political attitude questions. Journal of the Royal Statistical Society Series A 185(3), 872–890
N=2,402.
A spoken answer is longer and takes the respondent less than half the time.
27 seconds spoken vs 61 seconds typed
Revilla, M. & Couper, M.P. (2026). Comparing voice and text open-ended answers. Survey Research Methods 20(1) · doi:10.18148/srm/2026.v20i1.8456
Asked to answer the same survey, willingness falls with every step towards being recorded — and falls furthest at video.
67% would type · 47% would speak · 38% would appear on camera
Claassen, J., Lenzner, T., Höhne, J.K. & Ziller, C. (2026). A Survey Mode of the Future? Investigating Respondents' Willingness to Participate in Self-Administered Video-Based Web Surveys. methods, data, analyses 20(1), 317–346 · doi:10.12758/mda.2025.10
Randomised, n=1,992, χ²(9)=137.11, p<.001. Measures stated willingness, not observed behaviour.
Privacy is the top objection to video and barely registers for voice, from the same research group using the same instrument.
41% cite privacy for video · 7.7% for voice
Claassen et al. (2026); Lenzner, T. & Höhne, J.K. (2022) (2026). Reasons for refusing video and voice answers. methods, data, analyses; International Journal of Market Research 64(5), 594–610
Multiple mentions allowed. "Uncomfortableness" is 19% for video, absent for voice.
Offered the camera in a real study, two-thirds of people said no — and the reasons were about being seen, not about privacy.
202 of 302 declined · 51% dislike seeing themselves · 48% not camera-ready · 32% found it intrusive
Socratic Technologies (2024). Video vs. Text Responses. Socratic Technologies research note, 24 April 2024
A market-research vendor publishing against its own video product. Real experiment with a stated N, but not peer reviewed.
Turning the camera on makes people more tired and less likely to speak up. This is causal, not correlational.
1,408 observations across 103 people, camera use manipulated over four weeks
Shockley, K.M., Gabriel, A.S., Robertson, D., Rosen, C.C., Chawla, N., et al. (2021). The fatiguing effects of camera use in virtual meetings: A within-person field experiment. Journal of Applied Psychology 106(8), 1137–1155 · doi:10.1037/apl0000948
Effect was stronger for women and for newer employees. Video meetings, not surveys.
Seeing your own face while you talk lowers what you think of your own performance.
self-evaluation 4.99 with self-view vs 5.57 without, d=0.54
Shin, S.Y., Ulusoy, E., Earle, K., Bente, G. & Van Der Heide, B. (2022). Mirror, mirror on my screen: Focus on self-presentation on video conferencing platforms. Journal of Computer-Mediated Communication 28(1) · doi:10.1093/jcmc/zmac028
When the camera is genuinely optional, most people keep it off — and the reason they give first is embarrassment, not privacy.
55.1% reluctant · shame and anxiety 19.4% vs home privacy 8.4%
Gherheș, V., Șimon, S. & Para, I. (2021). Analysing Students' Reasons for Keeping Their Webcams On or Off. Sustainability 13(6), 3203
Single university, engineering students, education setting.
Making a questionnaire shorter is one of the few reliably effective ways to raise response.
odds ratio 1.64 across 481 randomised trials
Edwards, P.J., Roberts, I., Clarke, M.J., DiGuiseppi, C., Wentz, R., Kwan, I., Cooper, R., Felix, L.M. & Pratap, S. (2009). Methods to increase response to postal and electronic questionnaires. Cochrane Database of Systematic Reviews, Art. No. MR000008 · doi:10.1002/14651858.MR000008.pub4
95% CI 1.43–1.87. The electronic-questionnaire arm rests on only two trials.
How hard something looks changes how long people think it will take, and whether they agree to do it at all.
the same instructions judged to take 15.1 minutes in a hard-to-read font vs 8.2 in a clear one
Song, H. & Schwarz, N. (2008). If It's Hard to Read, It's Hard to Do. Psychological Science 19(10), 986–988 · doi:10.1111/j.1467-9280.2008.02189.x
Most of the friction people blame on speaking is actually the recorder. The same study found a twelvefold difference between two implementations of the same idea.
~60% failed to answer with one voice input · under 5% with the other
Revilla, M., Couper, M.P., Bosch, O.J. & Asensio, M. (2020). Testing the Use of Voice Input in a Smartphone Web Survey. Social Science Computer Review 38(2), 207–224
A spoken answer cannot easily be edited, so it arrives closer to what the person actually thought.
Höhne, J.K., Gavras, K. & Claassen, J. (2024). Typing or Speaking? (p. 1068). Social Science Computer Review 42(4)
Say "less edited", never "more honest": Schober et al. (2015, PLoS ONE 10(6):e0128337) found text produced MORE disclosure of socially undesirable behaviour than a voice interview.
Want this as an email you can forward?
We will send you a copy, once. No sequence, no newsletter, unsubscribe is a reply.