Narration Design for Game UI/UX: How to Differentiate Tutorial Voice, Menu Readouts, and Story Delivery

Narration from ¥50,000, delivered in as little as 24 hours.
* If you have a fixed budget, let me know and we can work from there.
In Game UI/UX, Voice Is a Navigation Layer Before It Is an Effect
When people think of game audio, they often focus on character voices, theme songs, or cinematic scenes. In actual production, however, the part that directly affects UI/UX is narration design: what to say, in what order, and with what emotional temperature. Console and mobile/social games differ in session length, control context, and drop-off points, so even the same “good voice” functions differently depending on the platform.
When directing voice for games, I first divide its role into three layers:
1. Tutorial voice, whose goal is comprehension.
2. Menu readouts, whose goal is operational support and accessibility.
3. Story delivery, whose goal is immersion and world-building.
If all three are recorded with the same energy, the same mic distance, and the same information density, UX usually suffers.
Tutorial Voice Should Anticipate Action, Not Merely Explain
The most common tutorial mistake is over-politeness in the script. Players do not want to “understand a sentence”; they want to know what to touch next. In practice, one sentence should ideally stay around 20–35 Japanese characters in density, or about 1.2–2.5 seconds in spoken length. Once a line exceeds 3 seconds, the player’s hands and ears often fall out of sync.
For example, instead of:
“First, tap the Formation button in the lower right and try changing your party.”
a better line is:
“Tap ‘Formation’ on the lower right. Change your party.”
The line ending matters. If tutorial voice sounds too firm, it becomes bossy; if it sounds too soft, it creates hesitation. I often separate delivery into three stages:
- Explanation: flat and neutral
- Guidance: slightly raised ending tone
- Completion: lowered ending tone for closure
On the implementation side, middleware such as Wwise or FMOD should branch by conditions like “no input for 5 seconds,” “failed 3 times,” or “return visit.” Avoid repeating the exact same line. A practical hierarchy is: full line on first visit, shorter version on second, and only SFX plus text from the third onward.
Menu Readouts Must Protect Browsing Speed, Not Add More Information
Menu readouts are not just a kindness feature. They reduce cognitive load in UI. This matters especially on handhelds, smartphones, and console setups viewed from TV distance, where eye movement and text size create real constraints. Voice can support faster navigation.
That said, too much acting hurts usability. Menu readouts prioritize consistency over character flavor. During recording, I usually keep intonation at about 60–70% of normal narration range, make initial consonants clear, and cut line endings short. A single item should typically last around 0.6–1.0 seconds.
Examples:
“Shop” 0.7 sec
“Missions” 0.8 sec
“Presents” 0.9 sec
For accessibility, keep the readout order fixed:
- Selected: item name + state
- Disabled: item name + “unavailable”
- Notification present: item name + “unchecked items”
If this order changes from screen to screen, players cannot build reliable listening habits.
In production, it is safer to manage readout text separately from visual UI labels. The displayed label might be “Gacha,” while the spoken version becomes “Draw gacha.” The screen might say “Enhance,” while the spoken line says “Enhance equipment.” Visual labels and auditory labels do not need to be identical. Designing this split in a spreadsheet from the beginning dramatically reduces revision costs later.
Story Voice Must Be Designed Separately from UI Voice
Story scenes naturally demand the most emotion. But from a UI/UX perspective, smooth transition matters just as much as emotional intensity. If a dramatic battle ends and the player instantly hears a system guide in the exact same tone, immersion breaks.
A useful solution is to separate voice layers by function:
- System layer: carries information
- Guide layer: directs the player while preserving world tone
- Drama layer: maximizes emotion
Even if one actor performs all layers, function can still be separated through EQ and perceived distance. For example, the system layer may emphasize clarity around 2–4 kHz, while the drama layer adds body near 200 Hz and more reverb. At recording stage, keeping the mouth 15–20 cm from the mic for system lines and 10–15 cm for dramatic lines also helps create a stable distinction in editing.
In social games with frequent updates, tone drift across story events is a common issue. The best prevention is to write “functional direction” in the script rather than abstract emotion words.
Bad: “gently”
Better: “make the player feel safe moving to the next choice”
Bad: “cool”
Better: “retain post-victory excitement while landing into the result screen”
If direction remains abstract, long-running live-service titles eventually lose consistency.
A Practical Pre-Implementation Checklist
Here is a checklist I often use before implementation:
- Is each voice line shorter than the player’s expected wait time?
- Does the same screen repeat the same line more than three times?
- Is the voice still intelligible over BGM at roughly -16 to -18 LUFS?
- Does it still work during skip, rapid tapping, or return visits?
- Can the player still take the minimum required action without text?
- Conversely, can the game still progress properly if voice is turned off?
Great game narration is not the voice that stands out most. It is the voice that never lets the player feel lost. Tutorial, menu, and story should not be handled under one generic theory of “performance.” Once you separate them by UX function, voice becomes a powerful design asset. It is not decorative polish added at the end. It is an interface connecting player action, understanding, and immersion.

Masahiro Kobayashi
Professional Narrator
A Japanese male narrator handling over 200 projects a year across corporate videos, commercials and documentaries. Recorded in a broadcast-quality home studio and delivered fast.
Listen to voice samplesRelated Articles
Designing Narration for Audio Guides: Pacing at the Exhibit, Timing for GPS Triggers, and Workflow with Curators
A practical guide to narration design for museum, gallery, and tourism audio guides: pacing at exhibits, pause design for GPS-triggered playback, and efficient workflows with curators.
Multilingual Narration for Expos and International Exhibitions: Coexisting with Simultaneous Interpretation, Unifying Pavilion Tone, and Managing Recording Schedules
A practical guide to multilingual narration for expos and international exhibitions, covering coexistence with simultaneous interpretation, pavilion tone design, and recording schedule management.
Narration Design for Short-Form Social Video: How to Win the First Second in 15, 30, and 60 Seconds
A practical guide to narration for TikTok, Instagram Reels, and YouTube Shorts, covering first-second vocal hooks and vertical-format audio design by 15, 30, and 60-second durations.
CONTACT
Narration Enquiries & Quotes
Corporate VP, commercials, e-learning, product manuals — you do not need everything decided. Send the script length, intended media and target date, and I will come back with a proposal.
- From
- ¥50,000〜
- Turnaround
- 24 hours
- Format
- WAV / MP3
* If you have a fixed budget, let me know and we can work from there.
Or email directly: info@kobatee.jp