When the same product video enters the launch checklist, it often appears as three files: a finished version with English captions burned into the picture, a clean textless master, and a standalone VTT file. Whether Shopify product video captions should stay in the picture or be displayed by the player cannot be decided by the editor at the last minute before export.

This question is especially relevant to teams building new-product pages, brand story pages, or multilingual standalone sites. After reading this article, you can first determine whether the video is actually hosted natively by Shopify or played through an external player such as YouTube or Vimeo, then decide how many burned-in caption versions, switchable caption files, and language-specific masters to create. This helps prevent the development team from discovering only after export that the player cannot accept the planned caption file.

When viewers watch with the sound off, which information must remain on screen?

Start by isolating the information that becomes impossible to understand without captions. Product models, feature names, operating steps, capacity, or compatibility ranges will be lost during silent viewing if they are explained only in voice-over. That does not mean every line of narration must be burned into the picture. When the visuals already clearly demonstrate a button press, opening, closing, or installation action, captions only need to add the key conditions; there is no need to transcribe the audio word for word.

Production teams usually divide information into three layers: short selling points that must remain visible, complete dialogue that the player can turn on or off, and sound information needed for accessibility. W3C’s explanation of captions includes speech and non-speech sounds needed to understand the content, and emphasizes that they must be synchronized with the audio. The same resource also distinguishes closed captions from open captions that are always displayed. If procurement treats all three layers as simply “adding captions,” it becomes difficult later to determine whether the deliverable is visual packaging, language translation, or a complete caption file. See the W3C guide to captions and caption production.

In day-to-day execution, we first export a silent review version. The team watches the picture without listening to the narration: any selling point that disappears goes back into the information layer, while actions that explain themselves are given room to breathe. The resulting burned-in captions do not cover the entire film, and the standalone caption file is not forced to do work that visual design should handle.

Another easily overlooked category is information where “the sound itself is product feedback.” For example, a smart appliance’s confirmation tone after an operation, a latch closing, an alarm, or a mode-switch sound should be described briefly in square brackets in the captions script if it affects the user’s judgment. If it is only ambient music, it does not need to be described line by line. During editing, it is best to keep dialogue, production sound, and music on separate tracks so caption proofreaders can accurately distinguish sounds that need to be written from those used only to create rhythm.

Is the video hosted by Shopify or by an external player?

The playback path should determine the caption format first. Shopify’s official documentation currently lists two ways to add product videos: upload a video file or embed a YouTube or Vimeo link. Direct uploads support MP4, MOV, and WebM, subject to requirements for duration, file size, and resolution. The theme must also support the relevant media type. In other words, the fact that “the admin can upload a video” does not show how the current theme will load a standalone caption track. Before launch, the development team should confirm the actual player, theme component, and available fields instead of leaving the production team to guess. See the Shopify product media types guide.

If the video is uploaded natively and the current component reads only one video file, the procurement package should include at least a burned-in caption version and a clean textless version. Whether WebVTT or another caption track can be integrated should be determined through development testing. With a YouTube or Vimeo embed, switchable captions are usually managed through the external platform account and player. The production team should deliver proofread caption files, a clean master, and a publishing timeline, while the operations team handles uploading, language settings, and permissions.

The page placement also changes the answer. Product hero media is often short and action-dense, with viewers perhaps staying for only a few seconds, so key selling points are better presented as restrained on-screen information. Brand stories or installation tutorials need pausing, replaying, and language switching, which makes switchable captions more valuable. If the same asset is used in two locations, do not export only one file: duplicate the timeline, then adjust the opening hold, caption density, closing action, and cover candidates for each placement.

Do not revisit this decision as if it were the same as the hosting choice. If the team has not yet decided between native upload and external embedding, first see Four considerations for uploading and embedding Shopify videos; this article then addresses how to split caption assets after the path is determined.

How should delivery differ between one market and multilingual sites?

When there is only one language market, captions will remain fixed, and the page has no switching requirement, burned-in captions require the fewest operational steps. However, they lock the language into the picture. As soon as multiple markets are planned, retain a clean textless master and manage translation, timecodes, and visual packaging separately. Shopify’s localization materials explain that stores can publish in different languages and manage translated content. Whether video changes with the market still depends on the theme, apps, and page implementation; “the store supports multiple languages” must not be interpreted as “the video will automatically switch captions.” See the Shopify localization and translation guide.

In practice, delivery can be divided into three tiers by risk. For a single-market page, deliver one finished version with burned-in captions and one clean textless version. For multilingual projects sharing the same visuals, deliver an SRT or VTT for each language in addition to the textless master, along with a set of burned-in review versions. If selling points, regulatory wording, or shot order differ by market, do not only replace the caption file; output separate market versions.

Before translation begins, freeze a Chinese or English source script with shot numbers, stating the product facts, units, proper names, and approved abbreviations line by line. That way, when translators encounter terms such as “mode,” “level,” or “battery life,” they do not have to infer the product meaning from the picture. Once a native-language reviewer approves the wording, post-production can handle line breaks, in and out points, and layout. If the translation is changed after typesetting, variations in length will repeatedly push captions out of the safe area and increase rework across every language version.

File naming should also be agreed upon in the project. For example, hero_16x9_clean_v05.mp4 indicates the clean master, hero_16x9_en-burned_v05.mp4 indicates the English burned-in review version, and hero_en-US_v05.vtt indicates the caption file for the corresponding language and version. Whenever the timeline changes, the version numbers of all three files must advance together to prevent old timecodes from being applied to a new edit.

Can the caption safe area be shared across landscape, portrait, and mobile?

Do not simply shrink landscape captions and force them into a portrait layout. The product-page player, floating purchase button, mobile browser bar, and cropping behavior of different themes can all take up space at the bottom of the picture. Before production, obtain screenshots of the real desktop and mobile containers. Mark the protection zones for the product, hand movements, and captions first, then decide whether landscape, square, and portrait versions need separate typography layouts.

During the shoot, leave an extra second of stable framing before and after each key action, and keep the product out of the planned caption area whenever possible. Have the hand model enter from the side to prevent the wrist and explanatory text from competing for the same position. In post, lock the single-line length and line-break rhythm of the Chinese or English reference version before creating other languages. Longer texts such as German should not be handled by merely reducing the font size; rebreak the lines. Text in different directions, such as Arabic, also requires joint review by a native-language proofreader and the page team.

Caption timing must follow the shot action. When a button is pressed, the screen changes, or an accessory clicks into place, the text should remain long enough for viewers to connect the action with the result. At a cut, avoid leaving the previous line on screen into the next scene whenever possible. W3C’s explanation of captions for prerecorded media requires captions for prerecorded audio in synchronized media and states that captions should include information needed to identify speakers and important sounds. It can be used to create a review checklist; see the WCAG explanation of captions for prerecorded media.

Do not complete review only on a large monitor. At minimum, choose a commonly used narrow-screen phone, place the video in a page mockup close to its real width, and check whether two caption lines cover product controls, finger actions, or important materials. Then use typical system font sizing, browser zoom, and page overlays to run through playback, pause, scrubbing, and back navigation in full. Any issues found here should be re-typeset for the corresponding aspect ratio; do not uniformly reduce the font size simply to make everything fit.

How can the final acceptance package get operations, development, and production to sign off?

The minimum acceptance package should not consist of only “one final.mp4.” It should include an archived clean master, burned-in caption versions for each page placement, SRT or VTT files for every language, a timecoded script, font and layout specifications, and a file index. If the captions include confirmation tones, mechanical sounds, or speaker identities, mark them in the script as well; do not leave only the dialogue.

Run launch testing by viewing state: when the page is first opened with the sound off, does the core information still make sense? After sound is enabled, are the captions synchronized with the voice-over? Can the player turn captions on and off and switch among the planned languages? On a narrow mobile screen, are captions covered by controls? After switching the market or language from the product page, do the video and page copy still correspond? Automatic captions can serve as a first draft, but product model numbers, proper names, units, and sound cues must be corrected manually.

Responsibilities at the page level should also be documented. The production team is responsible for the master, burned-in versions, caption files, and visual proofing. The operations team handles external accounts and language publishing. The development team handles theme components, caption-track integration, and testing on real devices. To first see how ONCE can place page location, landscape or portrait format, and future use into one delivery checklist, see the Shopify product video solution; you can also browse the ONCE commercial imaging case library to compare actual visual expression.

If you are deciding how to caption a new Shopify product page, send ONCE the page placement, playback path, target languages, and current edit. We will first determine which information must remain on screen, then define the scope for masters, caption files, and real-device testing.