Voice Search: Is It Worth Optimising for Right Now
Voice search has been promised as the next big shift for roughly ten years running. Every year articles appear saying half of all queries will soon be spoken, and every year companies try to work out what to do about it.
The honest answer turned out to be boring: separate voice optimisation is not needed by almost anyone. But the work proposed under that heading is needed, just for different reasons.
What actually changed
Voice input became normal, but the situations people use it in turned out narrower than predicted. People talk to a device when their hands are busy or typing is awkward: driving, cooking, walking.
The real change happened elsewhere. Conversational queries stopped being voice-only: people now type full sentences into search because they got used to asking models questions. So what changed is not the input method but the phrasing. And that applies to all queries, not only spoken ones.
How a spoken query differs from a typed one
| Attribute | Typed | Spoken |
|---|---|---|
| Length | 2-4 words | 6-10 words |
| Form | Keywords | A full question |
| Language | Telegraphic | Conversational |
| Context | Often absent | Often local: “near me”, “right now” |
| Expectation | A list of options | One answer |
The last row is the only genuinely important difference. A voice assistant reads out one answer rather than showing ten links. Second place here equals zero.
Why separate optimisation is unnecessary
Three reasons why “voice search optimisation” as a standalone project does not pay off:
- It cannot be measured. Analytics platforms do not separate queries by input method. You will not know how much traffic came from voice, so you cannot judge the result.
- The work duplicates other work. Everything recommended for voice, questions as headings, short direct answers, structured data, is exactly what AI answers and rich snippets require.
- The volume is overstated. Predictions about the share of voice queries have failed to materialise for ten years, and budgeting against them is risky.
So the work is worth doing, but calling it “voice optimisation” is not. It is the same work on content structure, and it should be budgeted with the rest.
What is genuinely worth doing
The list is short and completely overlaps with preparing for AI search:
- Questions as subheadings. Phrased the way people actually say them: “how much does it cost”, “how to choose”, “what to do if”.
- A short answer immediately below the question. Two or three sentences, self-contained, with no references back to earlier text.
- FAQ markup. It tells the system directly which part is the question and which the answer.
- Conversational phrasing in the copy. Not “the cost of the service depends on” but “how much this costs depends on”.
- Speed. An assistant will not wait long before taking its answer from another source.
All of this is already covered by the work on AI search and content structure. There is no separate cost line here.
Where voice genuinely matters
Two cases where this is not fashion but a real scenario:
Local search. “Where can I get a phone repaired near me” is a typical spoken query, and it almost always leads to action the same day. What works here is not copy optimisation but a complete business profile, accurate opening hours and reviews. The same work as for maps.
Voice interfaces inside a product. If you have an app used with hands occupied, in logistics, warehousing or fieldwork, voice input delivers real time savings. But that is a product feature, not search optimisation.
How to tell whether this concerns you
Three signs worth paying attention to:
- Your audience searches for the service while moving: transport, repairs, delivery, urgent services.
- Search Console shows queries of six words or more phrased as questions.
- You have a physical address customers visit.
If none of those match, no separate action is needed at all. What you already do for content structure is sufficient.
In short
Voice search is not a separate channel requiring its own budget. It is an input method that makes visible the same requirements AI search already imposes: a direct answer instead of a build-up, structure instead of solid prose, specifics instead of adjectives.
So the right answer to “should we optimise for voice now” is this: do not optimise for voice, write so that a complete answer can easily be extracted from your text. Then a voice assistant will take it, and so will a model in search, and so will a rich snippet.








