How to Commission Narration for Multilingual eLearning Without Failure: Planning for SSML, Subtitles, and LMS Integration

Narration from ¥50,000, delivered in as little as 24 hours.
* If you have a fixed budget, let me know and we can work from there.
Why narration commissioning becomes difficult in multilingual eLearning projects
In corporate training, SaaS onboarding, and educational content for medical and manufacturing fields, there has been a sharp rise in projects that “launch in Japanese first, then expand into English, Chinese, and Southeast Asian languages.” This is where a common bottleneck appears: the video is finished, but the audio gets stuck in the later stages.
The reason is simple. Narration is often commissioned as if it were merely a “reading task.” In reality, multilingual eLearning requires much more than good delivery. You must design for subtitle sync, on-screen dwell time, LMS implementation, pauses before and after quizzes, terminology consistency, and ease of future revisions. This is especially true when operating with Articulate Storyline, Adobe Captivate, Rise, or Moodle-based systems, where even a one-second timing mismatch directly affects both learner experience and revision costs.
In other words, what matters at the ordering stage is not only “what kind of voice you want,” but also “what operational conditions the recording must support.”
Decide the unit of audio first, not just the final voice
The first thing to decide in a multilingual project is the delivery unit: one continuous read, sentence-by-sentence files, or slide-by-slide files. If this is vague, later revisions become a nightmare.
For LMS content, a single long WAV file is usually far less practical than finely split files such as “module03_scene02_015.” There are three reasons. First, if a legal rule changes or the UI is updated and only one sentence needs to be fixed, the re-recording scope can be minimized. Second, when expanding into multiple languages, timing differences can be absorbed scene by scene. Third, resyncing with subtitles and interactions becomes much easier.
At minimum, your brief should specify file naming rules, segmentation unit, silence before and after each file, loudness standards, and delivery format. For example: “48kHz/24-bit WAV, 0.3 seconds of silence at head and tail, target around -19 LUFS, filenames must match CSV.” That single line can significantly reduce cleanup and reassembly costs for the editing team.
Provide the script as an implementation script, not just a reading script
Many projects still hand narrators a Word document containing only plain text. In multilingual eLearning, that is not enough. What you need is not merely a script for reading, but a script for implementation.
Ideally, it should be in table format with at least these columns: ID, screen name, narration text, subtitle text, terminology notes, emphasis points, estimated duration, revision frequency, and reference video URL. The key point is to separate the “narration text” from the “subtitle text,” because what sounds natural to the ear is often different from what reads clearly on screen.
For example, the narration may read, “Next, click the gear icon in the upper-right corner of the admin panel,” while the subtitle can be shortened to “Click the gear icon in the upper right.” With this separation, the narrator can read naturally, while the editor can keep subtitle length under control. It also makes translation into English and other languages much more stable.
The more you use SSML and AI voices, the more important direction for human narrators becomes
Recently, hybrid projects have increased: some languages use AI voices, while Japanese or brand-critical sections are voiced by humans. What is often overlooked is the need to align SSML design for AI with direction for human narrators.
On the AI side, you may use tags such as `
That is why I recommend creating a kind of “pseudo-SSML” direction sheet for human talent whenever AI voices are involved. For example: [PAUSE 0.3], [EMPHASIS term], [SLOW 90%]. This helps reduce the gap between human and AI output and makes it easier to unify the tone across an entire series.
Three common landmines to check before ordering in specialized fields
In fields with heavy terminology—medical, manufacturing, finance, information security—pre-order confirmation directly affects quality. The three most common trouble points are acronym pronunciation, product-name accent, and prohibited expressions.
Acronyms such as VPN, HPLC, and SAML may need to be spelled out letter by letter or pronounced as words depending on the project. Product names may have an internal accent pattern that an outside narrator cannot guess. And in compliance-sensitive projects, certain phrases such as “absolute” or “complete” may be forbidden.
Always attach a pronunciation glossary when placing the order. It does not need to be elaborate. A simple one-page list covering terms, readings, accent notes, prohibited paraphrases, and whether an English notation exists is enough. Even a small glossary can dramatically reduce the correction rate on the first take.
Good ordering reduces future revision risk, not just retake counts
It is risky to judge narration ordering as successful merely because there were few corrections. eLearning content is updated even after release, so what truly matters is whether the recording can withstand future revisions.
That is why you should ask the following when commissioning: Is this training likely to have UI changes in six months? Might more languages be added later? Will subtitles be handled by another vendor? Will voice and BGM be managed separately inside the LMS? Projects ordered with these conditions in mind are far more stable in operation.
Narrators and voice directors are not simply people who “record a nice voice.” They are audio design partners for content that must function over the long term. In multilingual eLearning especially, pre-recording information design determines 80 percent of the final quality. A little more care in the brief changes the recording, the editing, the translation, and ultimately learner comprehension.
Summary: do not order a voice, order a learning experience design
The most important thing in commissioning narration for multilingual eLearning is not just communicating your preferred voice quality, but sharing the operational conditions first. Segmentation unit, implementation script, subtitle policy, SSML-like instructions, terminology glossary, and assumptions about future revisions—if these six elements are prepared, the project’s failure rate drops significantly.
Ordering is not simply the act of saying, “Please read this.” It is the act of sharing how this training will be operated, in what environment, in which languages, and for how long, then designing the right audio together for those conditions. Once you go that far, narration stops being a mere finishing step and becomes a strategic element that supports the quality of the learning experience.

Masahiro Kobayashi
Professional Narrator
A Japanese male narrator handling over 200 projects a year across corporate videos, commercials and documentaries. Recorded in a broadcast-quality home studio and delivered fast.
Listen to voice samplesRelated Articles
Narration Booking Cancellations & Postponements: Fees, Etiquette, and Contract Wording
A practical guide to cancellation and postponement etiquette for narration bookings, covering industry norms, cost allocation, typical fee ranges from 0 to 100%, and contract wording.
The Complete Invoice & Quotation Template for Narration Fees: How to Separate Recording, Studio, Travel, and Retake Costs
A practical guide to standard quotation and invoice formats for narration projects, covering recording fees, studio, travel, retakes, invoice registration numbers, and standard payment terms.
How to Design Narrator Audition Reads Without Regret: Script Length, Direction Detail, and the Free-to-Paid Boundary
A practical guide for video production teams on designing narrator audition reads: ideal script length, content selection, direction detail, and where free test reads should end and paid work should begin.
CONTACT
Narration Enquiries & Quotes
Corporate VP, commercials, e-learning, product manuals — you do not need everything decided. Send the script length, intended media and target date, and I will come back with a proposal.
- From
- ¥50,000〜
- Turnaround
- 24 hours
- Format
- WAV / MP3
* If you have a fixed budget, let me know and we can work from there.
Or email directly: info@kobatee.jp