Voice Design for Transit Ads and Digital Signage: Optimizing Intelligibility for Stations, Retail Spaces, and Hospitals

Narration from ¥50,000, delivered in as little as 24 hours.
* If you have a fixed budget, let me know and we can work from there.
In Transit Ads and Digital Signage, the Winning Voice Is Not Merely Pleasant but Retentive
Narration for transit advertising and digital signage fails when it is produced with the same mindset as web videos or corporate presentations. The reason is simple: the listening environment cannot be controlled. In stations, train noise and public announcements compete with the message. In retail spaces, background music and reflective surfaces blur articulation. In hospitals, a restrained acoustic environment still contains call chimes, conversations, and movement. So the priority is not whether the voice sounds beautiful as audio content, but whether the message is understood the first time.
That is why intelligibility must be designed around the reverberation characteristics of each installation site. In these projects, I often start not with the script, but by identifying what will mask the voice in the actual space. Is low-frequency HVAC noise dominant? Are there strong metallic reflections around 2–4 kHz? Is background music always present? If you misread those factors, even a well-performed voice will not communicate efficiently.
Reverberation Time and Noise Change the Correct Narration Strategy
In practice, simply understanding RT60 and ambient noise levels greatly improves design accuracy. As rough references, station concourses often fall around 0.8–2.0 seconds, large retail spaces with high ceilings around 1.2–2.5 seconds, and hospital waiting areas around 0.6–1.2 seconds. Of course, direct measurement is ideal, but even early-stage projects benefit from smartphone apps or simple measurement tools. NTi Audio XL2, Rational Acoustics Smaart, or even REW for quick checks can provide useful direction.
Signal-to-noise ratio also matters. To secure speech intelligibility, an effective SNR of around +10 dB is often desirable, and at least +6 dB is frequently needed. However, in signage environments, turning up the volume is not a real solution. In reflective spaces, increasing level alone often makes the overall sound harsher rather than clearer, and consonants do not necessarily become easier to distinguish. That is why success depends less on loudness and more on bandwidth control and speech pacing.
How the Voice Should Change for Stations, Retail Spaces, and Hospitals
In stations, the voice competes with announcements and train noise, so overly dramatic reading with large pitch swings is usually disadvantageous. A better approach is to make sentence openings clear and avoid dropping endings too vaguely. Speech rate should be about 5–12% slower than usual, with shorter sentence units to reduce cognitive load. Compound nouns and loanwords often blur together, so slight segmentation by meaning improves comprehension.
In retail spaces, coexistence with background music is assumed. Here, boosting the top end too aggressively in pursuit of a “cut-through” voice often makes the result fatiguing. From the performer’s side, it is more effective to create a natural presence around 1.5–2.5 kHz rather than relying heavily on 3 kHz or above. Not merely a “bright” voice, but one with core and controlled variation. In mixing, reducing muddiness around 200–300 Hz and taming excessive sibilance above 4 kHz with de-essing usually works better for long playback cycles.
Hospitals require the opposite mindset: excessive assertiveness becomes noise. The audience may include elderly people or those who are unwell, so heavy compression and fast pacing should be avoided. In hospital work, I usually go about another 5% slower than in station projects, use more deliberate pauses, and articulate initial consonants with extra care. Frequency design should also avoid creating an aggressive peak around 2–3 kHz. The goal is not a “quiet” voice, but a voice that feels safe to understand.
Practical Recording and Editing Techniques That Actually Help
During recording, avoid creating too much intimacy. A condenser microphone placed very close to the mouth can produce density, but in signage playback the low-frequency proximity effect may become counterproductive. A distance of roughly 15–25 cm, with a pop filter and slight off-axis positioning, often yields material that is easier to manage. Standard choices such as U87-style microphones can work, but models with clearer contour and less exaggerated gloss—such as the Neumann TLM 103, Sennheiser MKH 416, or certain Earthworks microphones—often suit this application well.
In editing, do not judge solely by LUFS. For example, a file normalized to -16 LUFS may still lack consonant visibility in short-contact advertising playback. Before delivery, I always check not only on full-range monitors, but also on small speakers, soundbars, and mono playback resembling actual facility systems. For EQ, I typically clean below 80 Hz gently, assess muddiness around 250 Hz, tune intelligibility around 2 kHz, and control harshness around 6–8 kHz as needed. Compression should stay moderate—around 2:1 with roughly 2–4 dB of gain reduction.
The Essential Skill Is Reading the Space, Not Just the Script
Audio for transit ads and digital signage does not become complete inside the studio. It is completed when it reflects off hard station walls, when it blends with retail background music, and when it reaches someone in a quiet hospital waiting area without adding stress. That is why a narrator must do more than read text; the narrator must read the acoustic atmosphere of the destination space.
A “good voice” can sometimes remain only a source of satisfaction for the production side. But a voice that arrives with clarity functions in the real world. In transit advertising and signage, what earns trust is not only acting skill, but a voice designed with spatial acoustics in mind. Once you start working from that perspective, the same script can produce dramatically better results.

Masahiro Kobayashi
Professional Narrator
A Japanese male narrator handling over 200 projects a year across corporate videos, commercials and documentaries. Recorded in a broadcast-quality home studio and delivered fast.
Listen to voice samplesRelated Articles
Designing Narration for Audio Guides: Pacing at the Exhibit, Timing for GPS Triggers, and Workflow with Curators
A practical guide to narration design for museum, gallery, and tourism audio guides: pacing at exhibits, pause design for GPS-triggered playback, and efficient workflows with curators.
Multilingual Narration for Expos and International Exhibitions: Coexisting with Simultaneous Interpretation, Unifying Pavilion Tone, and Managing Recording Schedules
A practical guide to multilingual narration for expos and international exhibitions, covering coexistence with simultaneous interpretation, pavilion tone design, and recording schedule management.
Narration Design for Short-Form Social Video: How to Win the First Second in 15, 30, and 60 Seconds
A practical guide to narration for TikTok, Instagram Reels, and YouTube Shorts, covering first-second vocal hooks and vertical-format audio design by 15, 30, and 60-second durations.
CONTACT
Narration Enquiries & Quotes
Corporate VP, commercials, e-learning, product manuals — you do not need everything decided. Send the script length, intended media and target date, and I will come back with a proposal.
- From
- ¥50,000〜
- Turnaround
- 24 hours
- Format
- WAV / MP3
* If you have a fixed budget, let me know and we can work from there.
Or email directly: info@kobatee.jp