Add voiceover and subtitles to an Opus 5.5 video project

Align narration and sentence captions in an Opus 5.5 Remotion project. Get the working MP4, WAV, SRT and source, with timing checks and fixes for drift.

Metronome line drawing on pale paper against a sage background, with orange bars and the title Opus 5.5

To add voiceover and subtitles to an Opus 5.5 video, treat audio timing as project data. Create or record the narration, measure it, map each sentence to the timeline, then render the audio and captions together. Asking the model to “make the subtitles sync” without giving it actual timing is not enough.

This tutorial uses the same real-screenshot project as our product demo walkthrough. You can start from its downloadable files without completing that article first. The goal here is a narrated 30-second MP4 with readable sentence captions, an editable timing source and a separate SRT—not a new prompt collection or an unverified promise of automatic dubbing.

Watch and inspect the audio example

Actual Remotion export with an English synthetic reference voice. The voice is deliberately a timing aid, not a demonstration of professional narration quality. Captions follow whole sentences and include reading holds.

Download the complete project, narrated MP4, reference WAV and English SRT. The five-language articles share this English example; we are not presenting it as five localized audio productions.

The actual model call used claude-opus-5-5 in first-party Claude Code to generate Remotion code. It received filenames and written scene instructions, not image pixels or an audio recording. We then supplied the real source captures, generated the reference voice with eSpeak NG, corrected the code and rendered locally. This division of work is central to the tutorial: Opus wrote code; the speech engine made audio; Remotion made the final file.

Choose the right timing workflow

There are two common starting points. If your script and scene durations are still flexible, record or synthesize the approved voice first, then fit the visuals to its pace. If the video must fit an already approved 30-second slot, define the scene windows first and shorten any sentence that does not fit naturally.

Our example follows the second route. It has five existing scenes and five short sentences. Each audio segment must fit inside its assigned caption window; the reference-audio script rejects a segment that exceeds the end time. That check catches a real class of synchronization error before rendering.

Starting materialRecommended approachMain risk to check
Approved continuous narrationMeasure/transcribe it, then place visuals and captions around the actual speechGuessing sentence timing from word count
Fixed-duration product videoWrite short sentences for each scene; measure each recordingRushing or clipping a sentence to keep a fixed cut
Existing video with speechObtain timestamps from the actual soundtrack, then review themTreating an automatic transcript as a verified subtitle file
Multilingual versionRe-record and retime each languageReusing English timing after the spoken duration changes

Remotion’s caption documentation covers importing, displaying and exporting captions. This project intentionally uses a small JSON sentence schedule so the timing is visible and easy to audit. It does not run automatic transcription or forced alignment.

Start with the provided voice, then replace it

Install Node.js and npm before running the project. The optional terminal media checks below also require a separate FFmpeg installation, including ffprobe; npm ci does not install that shell command. You can render through the supplied npm scripts without using the optional inspection commands.

Unzip the archive and run the project from its root:

npm ci
npm start

Choose DemoNarrated. The archive already contains public/narration.wav, so rendering does not require installing a speech engine or buying a voice service. Check the first sentence near the beginning, the documentation sentence around 13 seconds, and the last sentence near 27 seconds.

Real Remotion Studio showing the narrated composition, caption and audio timeline

Real project preview with a WAV audio track. A visible waveform proves that audio is present in the project; it does not establish pronunciation quality or word-level alignment.

To reproduce the reference voice, install eSpeak NG through its documented route for your system, then run:

python3 scripts/make_audio.py
npm run render:narrated

We used eSpeak NG 1.52.0, the en-us voice and a speed setting of 190 words per minute. This formant-based voice is intentionally synthetic. It is useful for testing a timeline before choosing the final narration; the number is a tool setting, not a guarantee that every sentence has an identical speaking rate.

For a finished campaign, use a recording or voice output you are authorized to publish. Listen for the product name, abbreviations, punctuation pauses and the final word of each sentence. A correct transcript cannot establish whether a voice sounds natural. This article’s measured checks concern duration, placement and visible captions; they are not a human pronunciation or performance review.

Use one schedule for the WAV, captions and SRT

Open scripts/make_audio.py. Its LINES array is the reference source: each entry has a start time, a caption end and text. Running the script creates the 30-second WAV, src/captions.json, public/captions.en.srt and audio-timing.json. The composition imports the generated JSON, which avoids keeping a second manually edited caption list in the video code.

The measured reference segments are:

SentenceStartWAV segment durationSegment endsCaption ends
A clear product video starts with a clear brief.0.35 s2.664 s3.014 s3.60 s
Show the real interface. Here, we begin with the model catalog.4.35 s3.681 s8.031 s9.80 s
Then show where a viewer can find the documentation.11.35 s2.917 s14.267 s16.80 s
Connect each scene to an actual page, such as this API reference.18.35 s3.743 s22.093 s23.80 s
Keep the message simple. Plan, build, and verify.25.35 s3.537 s28.887 s29.50 s

These durations include the generated segment’s trailing samples. They are not measurements of individual phonemes. Captions deliberately remain visible after speech; the documentation sentence has about 2.53 seconds of additional reading time after its WAV segment. If you want tighter subtitle timing, shorten that hold after checking the actual narration.

The script creates silence around each segment and places the audio at the listed absolute start. It does not speed up a too-long sentence to force it into a window. If a revised sentence no longer fits, its explicit error tells you to shorten the sentence or move the end time before rendering.

A continuous replacement recording needs a different preparation step. Export a single full-length public/narration.wav, then edit src/captions.json and the SRT to match the actual recording. Do not run make_audio.py afterward unless you intend to overwrite that recording with the reference voice. Keep the original recording and the generated example in separate saved versions.

Convert seconds to frames carefully

The composition runs at 30 fps. A caption at 11.35 seconds falls between video frames. The final code uses Math.floor(start * fps) for the first visible frame, so that sentence appears at frame 340, about 11.333 seconds—slightly before the voice starts. It has no entrance fade that would hide the first spoken word. A short exit fade occurs during its final six frames.

const first = Math.floor(caption.start * fps);
const last = Math.round(caption.end * fps);
const visible = frame >= first && frame < last;

This illustrates the final visibility rule; the complete component also sets text layout and the end fade. Frame rounding is expected. Do not claim sample-accurate audio alignment from a 30-fps visual timeline.

The supplied Remotion 4.0.424 project imports Audio from remotion and resolves the WAV with staticFile. Current Remotion documentation describes that older HTML5 component as Html5Audio and recommends its newer media component for new audio integrations. Keep this archive’s pinned versions while reproducing it; do not combine a current-documentation API change with an unexplained package upgrade.

Keep subtitles readable without hiding words

The landscape composition gives captions a dark band separate from the screenshot. During the documentation and API scenes, the scene title sits on the left and the sentence caption on the right. The portrait version stacks image and captions vertically; use the vertical-video tutorial for its geometry.

Our English captions use 28-pixel text in landscape and 30-pixel text in portrait. These are project values, not universal social-platform rules. The chosen sentence lengths fit the checked layouts. A longer translation might not.

The initial model code limited text to two lines with CSS line-clamping. We removed that rule because it could silently hide the third line after an edit. The correct fix for an overlong subtitle is to split or rephrase it, enlarge the space, or revise the layout—not to make missing words invisible.

When changing text, check the first, longest and last captions at the intended display size. System font fallback can alter line breaks; the archive does not guarantee identical typography on every operating system. If you need a fixed brand font, bundle an appropriately licensed font, load it before rendering, and repeat these checks.

Ask Opus for a bounded timing revision

This adaptation prompt is useful after you have measured your own voiceover. Replace the placeholders with real times; it is not an additional model test.

Update the existing DemoNarrated composition only.
Keep the approved screenshots, 30 fps and current scene order.
The replacement narration is public/narration.wav.
These are measured sentence intervals, in seconds:
[paste start, end, exact spoken text]

Use one caption data source. Preserve every spoken word.
Show each sentence before or at its first spoken word; explain frame rounding.
Do not invent per-word timestamps or automatic transcription results.
Allow a reading hold only where I explicitly provide one.
Do not use line-clamping to hide overflow.
If a sentence crosses its scene cut, identify it instead of silently trimming audio.
Return changed files, the timing assumptions and the frames to inspect.

After a change, compare the text against the narration, then verify the visuals. If a line describes the API page while the catalog is still on screen, both files may be individually valid but the combined story is wrong. Correct the shared scene schedule or the recording; a caption offset alone will not repair mismatched narration.

Export and diagnose the actual file

npm run render:narrated
ffprobe -v error -show_entries stream=codec_name,codec_type,duration,sample_rate \
  -show_entries format=duration -of json out/narrated.mp4

The rendered example has an H.264 video stream and an AAC audio stream. Its video timeline is 30 seconds; compressed audio can report a slightly different stream/container duration because of encoder padding. Our reference WAV is exactly 30 seconds. Check audible content and synchronization rather than interpreting a small AAC tail as a scene-duration error.

SymptomLikely place to inspectWhat to change
Entire MP4 is silentComposition, mute flag, WAV path and stream metadataUse render:narrated, verify the audio asset, and ensure the render was not muted.
Every sentence is late by the same amountLeading silence or a shared start offsetCorrect the single offset and inspect the first and last sentences again.
Drift grows through the videoWrong duration assumptions, audio speed or source timelineMeasure the actual recording; do not patch each caption from guessed word counts.
One sentence crosses a cutThat sentence’s duration and scene boundaryShorten/re-record it or extend the scene and all dependent timing.
Words are missing on screenText wrapping and fixed caption boxSplit the sentence or resize the layout; never conceal the overflow.
Preview has sound but final export does notFinal file and render commandInspect the MP4’s audio stream, not only Studio’s waveform.
SRT and burned-in captions differLast edited source and regeneration orderRegenerate from the same schedule, or update both deliberately for a replacement recording.

Before delivery, play the complete exported file with sound, check sentence starts and ends, and inspect the version uploaded to your publishing platform. Burned-in captions remain visible in the pixels; a separate SRT depends on whether that platform supports and enables it. The example supplies both, but does not claim either was uploaded to a social channel.

Keep the audio work separate from model access

The Opus 5.5 model entry identifies the model relevant to this coding workflow. This example does not establish Ofox TTS availability or billable video-rendering behavior; its recorded model call used first-party Claude Code and its audio was created locally.

Use the main Opus video guide for the overall production path, the screenshot demo tutorial for source selection, and the portrait adaptation before sending the piece to a mobile-first format. A reliable handoff includes the final MP4, the WAV, the matching subtitle file and the exact project version that produced them.

Frequently Asked Questions

Did Opus 5.5 generate the narration?
No. Opus generated the video project code. This example uses an explicitly synthetic eSpeak NG reference voice, and Remotion combines the WAV with the video.
Are these word-synchronized subtitles?
No. The example uses five scheduled sentence captions, with extra reading time after each audio segment. Word-level highlighting needs real word timestamps and separate verification.
Can I replace the voice without regenerating the video design?
Yes. Replace the audio and update the caption timing to match it, keeping the approved visuals. If the new voiceover exceeds the scene or project duration, revise the shared timeline before exporting.