THE METHOD

A typed answer averages
two words.

That is not a failure of your question. It is the price of the format. Every screen here is designed against a specific, measured reason people give you less than they know.

“I spent years optimising survey response rates for around 300 universities. The lesson was never that people don’t want to help. It is that every second of hesitation costs you a person — and almost nobody designs against that.”

SPENCER PETERSON — FOUNDER

Why this exists

Survey tools are built for the person writing the survey. Logic jumps, piping, dashboards, quotas — every feature serves the researcher. Almost nothing in a typical form builder was designed for the person on the other end, standing at a bus stop being asked to type a paragraph into a box on a phone.

So they type two words. Antoun and colleagues measured it on a probability panel of 1,390 people: open-ended answers average about two words, and it barely matters whether they are on a phone or a desktop. Then the researcher concludes that open questions don’t work — when what actually happened is that writing is expensive and nobody accounted for the price.

People don’t give short answers because they have nothing to say. They give short answers because the format charged them too much to say it.

Ask the same question out loud and the answer comes back at fifty-four words instead of twenty-two — and it takes them 27 seconds instead of 61. Longer, and less of their time. That is the whole product in one sentence.

The part that surprised us: it is not the same answer, only longer. Gavras and colleagues found the topics people raise overlap by only 30–70% between writing and speaking. Voice is not a nicer text box. It is a different instrument, and it collects things the text box never gets.

THE FOUR FRICTIONS

What actually stops someone answering

Each is a real cost paid by the respondent. Each has a design response, and each response is in the product.

01

Not knowing what it costs

“How long is this?” — unanswered, that question is itself a reason to close the tab. Song and Schwarz showed that identical instructions in a harder-to-read font were judged to take 15 minutes rather than 8, and people were less willing to start.

WHAT WE DIDThe time is stated before the first question and counts down honestly as you go. Never a bar that flatters the early questions — Villar and colleagues found across 32 experiments that a progress bar which starts slow and speeds up makes drop-off worse, not better.

02

The cost of writing

Writing is slower than speaking and runs a second process alongside the thinking: spelling, grammar, whether that sentence reads badly. That monitoring is where nuance gets flattened into “it was fine”.

WHAT WE DIDAnswer out loud. And because a spoken answer cannot easily be edited, it arrives closer to what the person actually thought — the researchers’ own explanation for why the content differs.

03

Being seen

Ask for video and willingness falls off a cliff: 67% of people will type an answer, 47% will speak one, and 38% will appear on camera. Offered a real camera, two-thirds declined — 51% because they dislike seeing themselves and 48% because they weren’t camera-ready.

WHAT WE DIDThere is no camera. There will never be a camera. The argument is below, and it is the strongest thing on this page.

04

Fear of losing the work

If a survey might discard what you have already given, the rational move is not to start. Anything that saves only at the end asks for a leap of faith.

WHAT WE DIDEvery answer saves the moment it is given. Stop after two questions and those two answers are yours — a half-finished survey is half a survey, not nothing.

A COMMITMENT

We will never add video.

It is the most common feature request in this category and the one we will keep refusing — not because video is hard, but because of what asking for it does to the person answering.

The moment a camera is involved, a respondent acquires a second job. They check the room behind them. They notice their hair. They watch their own face in the corner while they talk, which is a thing almost nobody does comfortably: Shin and colleagues found that seeing your own video lowers what you think of your own performance. Then they say something slightly wrong, see themselves say it, and delete the take — or decide it isn’t worth it and you get nothing.

This is measurable, and it is causal rather than correlational. Shockley and colleagues manipulated camera use across four weeks and 1,408 observations: turning the camera on made people more tired and measurably less likely to speak up, with the effect strongest for women and for newer employees. A camera does not just cost you respondents. It changes what the ones who stay are willing to say.

And when the camera is genuinely optional, most people quietly keep it off. Gherheș and colleagues found 55% reluctant — and the reason people gave first was not privacy. It was embarrassment, at more than twice the rate.

Even our competitors say it

“Appearing on camera may feel intrusive or require extra preparation. Audio offers a useful middle ground by preserving vocal context without asking participants to be seen.”— Typeform, which owns VideoAsk

And we never show them the transcript

This one surprises people. We transcribe every recording — but the respondent never sees it. The moment somebody reads back what they just said, they start editing: “that came out wrong”, “I said um four times”, “let me do that again”. A transcript shown to the speaker converts a spoken answer back into a written one, and you lose precisely the thing you came for.

So they talk, and they move on. You get the pause before “well, honestly…”. They get to not care about their grammar.

WHAT IT IS WORTH

Every response is a call you didn’t take

$50

of your time, per response

A discovery call is not thirty minutes. It is the scheduling, the reschedule, the five minutes of “can you hear me”, the call, and the writing-up. Costed at what an hour of a founder’s or a researcher’s time is worth, one response is around $50 you did not spend.

Twelve responses is a week of calendar tetris that never happened — $600 of your time, and you still have the recordings. The app shows this figure as it climbs.

An estimate, and an incomplete one: it counts your hour only, and ignores the respondent’s time entirely. Put your own hourly rate against it if ours is wrong for you — the arithmetic is one response, one hour.

What we are careful not to claim

Plenty of tools here will tell you that switching to voice raises your response rate. We will not, because the published evidence says the opposite. In every study that measured it, voice answers were richer and rarer — break-off roughly doubled, and when researchers tried fixing it with better instructions, it did not work.

So the honest claim is about depth per respondent, not headcount. From the people who do answer, you get several times what a text box would have given you, on topics a text box would not have surfaced. Every question also has a text field under the record button, so nobody who would rather type is turned away.

Voice-first, never voice-only. The moment you force it, you have built the same trap video builds.

One more thing worth knowing, because it is the reason we obsess over the recorder itself: when Revilla and colleagues tested two implementations of voice input in the same study, one lost 60% of answers and the other lost under 5%. Most of what looks like reluctance to speak is a bad record button. That is a product problem, and it is ours.

CHECK OUR WORKING

Sources

Every figure on this page, with the paper it came from and what it does not prove.

  1. A typed open-ended answer averages about two words — on a phone and on a desktop alike.

    ~2.1 words on smartphone, ~2.0 on PC

    Antoun, C., Couper, M.P. & Conrad, F.G. (2017). Effects of Mobile versus PC Web on Survey Response Quality. Public Opinion Quarterly 81(S1), 280–306 · doi:10.1093/poq/nfw088

    Randomised crossover on a probability panel, n=1,390.

  2. The same question answered out loud produces two to three times as many words.

    22 words typed → 54–56 spoken

    Höhne, J.K., Gavras, K. & Claassen, J. (2024). Typing or Speaking? Comparing Text and Voice Answers to Open Questions on Sensitive Topics in Smartphone Surveys. Social Science Computer Review 42(4), 1066–1085 · doi:10.1177/08944393231160961

    N=1,001, four topics, all p<.001.

  3. Speaking does not produce a longer version of the same answer. It produces a different one.

    topic overlap between spoken and written answers: 30–70%

    Gavras, K., Höhne, J.K., Blom, A.G. & Schoen, H. (2022). Innovating the collection of open-ended answers: the linguistic and content characteristics of written and oral answers to political attitude questions. Journal of the Royal Statistical Society Series A 185(3), 872–890

    N=2,402.

  4. A spoken answer is longer and takes the respondent less than half the time.

    27 seconds spoken vs 61 seconds typed

    Revilla, M. & Couper, M.P. (2026). Comparing voice and text open-ended answers. Survey Research Methods 20(1) · doi:10.18148/srm/2026.v20i1.8456

  5. Asked to answer the same survey, willingness falls with every step towards being recorded — and falls furthest at video.

    67% would type · 47% would speak · 38% would appear on camera

    Claassen, J., Lenzner, T., Höhne, J.K. & Ziller, C. (2026). A Survey Mode of the Future? Investigating Respondents' Willingness to Participate in Self-Administered Video-Based Web Surveys. methods, data, analyses 20(1), 317–346 · doi:10.12758/mda.2025.10

    Randomised, n=1,992, χ²(9)=137.11, p<.001. Measures stated willingness, not observed behaviour.

  6. Privacy is the top objection to video and barely registers for voice, from the same research group using the same instrument.

    41% cite privacy for video · 7.7% for voice

    Claassen et al. (2026); Lenzner, T. & Höhne, J.K. (2022) (2026). Reasons for refusing video and voice answers. methods, data, analyses; International Journal of Market Research 64(5), 594–610

    Multiple mentions allowed. "Uncomfortableness" is 19% for video, absent for voice.

  7. Offered the camera in a real study, two-thirds of people said no — and the reasons were about being seen, not about privacy.

    202 of 302 declined · 51% dislike seeing themselves · 48% not camera-ready · 32% found it intrusive

    Socratic Technologies (2024). Video vs. Text Responses. Socratic Technologies research note, 24 April 2024

    A market-research vendor publishing against its own video product. Real experiment with a stated N, but not peer reviewed.

  8. Turning the camera on makes people more tired and less likely to speak up. This is causal, not correlational.

    1,408 observations across 103 people, camera use manipulated over four weeks

    Shockley, K.M., Gabriel, A.S., Robertson, D., Rosen, C.C., Chawla, N., et al. (2021). The fatiguing effects of camera use in virtual meetings: A within-person field experiment. Journal of Applied Psychology 106(8), 1137–1155 · doi:10.1037/apl0000948

    Effect was stronger for women and for newer employees. Video meetings, not surveys.

  9. Seeing your own face while you talk lowers what you think of your own performance.

    self-evaluation 4.99 with self-view vs 5.57 without, d=0.54

    Shin, S.Y., Ulusoy, E., Earle, K., Bente, G. & Van Der Heide, B. (2022). Mirror, mirror on my screen: Focus on self-presentation on video conferencing platforms. Journal of Computer-Mediated Communication 28(1) · doi:10.1093/jcmc/zmac028

  10. When the camera is genuinely optional, most people keep it off — and the reason they give first is embarrassment, not privacy.

    55.1% reluctant · shame and anxiety 19.4% vs home privacy 8.4%

    Gherheș, V., Șimon, S. & Para, I. (2021). Analysing Students' Reasons for Keeping Their Webcams On or Off. Sustainability 13(6), 3203

    Single university, engineering students, education setting.

  11. Making a questionnaire shorter is one of the few reliably effective ways to raise response.

    odds ratio 1.64 across 481 randomised trials

    Edwards, P.J., Roberts, I., Clarke, M.J., DiGuiseppi, C., Wentz, R., Kwan, I., Cooper, R., Felix, L.M. & Pratap, S. (2009). Methods to increase response to postal and electronic questionnaires. Cochrane Database of Systematic Reviews, Art. No. MR000008 · doi:10.1002/14651858.MR000008.pub4

    95% CI 1.43–1.87. The electronic-questionnaire arm rests on only two trials.

  12. How hard something looks changes how long people think it will take, and whether they agree to do it at all.

    the same instructions judged to take 15.1 minutes in a hard-to-read font vs 8.2 in a clear one

    Song, H. & Schwarz, N. (2008). If It's Hard to Read, It's Hard to Do. Psychological Science 19(10), 986–988 · doi:10.1111/j.1467-9280.2008.02189.x

  13. Most of the friction people blame on speaking is actually the recorder. The same study found a twelvefold difference between two implementations of the same idea.

    ~60% failed to answer with one voice input · under 5% with the other

    Revilla, M., Couper, M.P., Bosch, O.J. & Asensio, M. (2020). Testing the Use of Voice Input in a Smartphone Web Survey. Social Science Computer Review 38(2), 207–224

Six questions. Four minutes. No call.

Unlimited surveys and questions on the free plan, no credit card.

Unlimited surveys and questions, free. No credit card.