Vocal priority levels must be established during the project initiation phase.

The key to commercial video audio mixing is defining the role of vocals during project initiation. Brands must first determine whether the video's core message is conveyed primarily through voice or through visuals and music. If the selling points are product specs, service commitments, or founder insights, vocals must take top priority. If the focus is atmosphere, emotion, or visual impact, vocals may serve only a supporting role. This assessment dictates budget and schedule allocation for all subsequent audio work.

Commercial video production footage from case materials, observing camera angles, subjects, and lighting relationships.
Case study frame sourced from the research material 'How one VFX artist made these 3 minutes of madness.' This image is for observing cinematography and production techniques only and does not represent an ONCE client project. Source page. Case Study Page

Specifically, clearly define the vocal function in the project brief—whether narration, dialogue, interview, or voiceover. Identify segments where vocals are essential and those where music or sound effects may take precedence. Upon receiving the brief, the production team should submit a vocal priority list detailing vocal types, recording methods, and expected mix ratios for each scene. The brand must approve this list upfront rather than waiting to provide feedback on the rough cut.

The risk is that failing to define priorities during pre-production leads to repeated adjustments in post-production mixing, where music and vocals compete, resulting in a final deliverable that satisfies neither. An exception applies to videos consisting purely of music or sound effects without vocals, which do not require this priority setting. The rule is simple: if vocals appear in the video, their priority must be established during pre-production.

Capture clean audio during filming to support vocal priority in post-production.

Achieving vocal priority relies partly on capturing clean audio during production. Commercial shoots often involve complex environments where air conditioning, traffic, and crew movement can interfere with post-production mixing. Crews should conduct sound checks before rolling, monitor ambient noise with headphones, and identify and eliminate noise sources whenever possible. If elimination is impossible, adjust microphone placement or reschedule the recording time.

Specifically, record a dedicated high-quality audio track synchronized with video for every shot containing dialogue or interviews. Also record at least thirty seconds of room tone to fill gaps in post-production. Keep recording equipment away from power transformers and fluorescent lights, and position the microphone fifteen to thirty centimeters from the talent's mouth to avoid breath noise and plosives. If multiple sound sources exist on set, record them on separate tracks rather than mixing them into one.

The risk is that rushing production by relying solely on camera-mounted or built-in microphones results in excessive reverb and high noise floors, making clean vocal extraction impossible in post-production. The standard is that monitored vocals should sound as clear as face-to-face conversation, with every word intelligible without raising the volume. An exception applies if the project explicitly adopts a documentary style allowing natural ambient sound; in such cases, recording standards may be lowered but must be clearly documented during pre-production.

Production teams must also ensure talent maintains consistent speech pace and volume. Directors should remind talent not to shout into or turn away from the microphone. If lines are changed on set, re-record the entire line rather than patching individual words to prevent tonal inconsistencies during editing.

In post-production editing, lay down vocal tracks before adding music.

During editing, vocal priority must be reflected in the timeline structure. Editors should place vocal tracks first to establish the start and end points of each line or voiceover before adding music and sound effects. This approach ensures that music and effects support the narrative backbone provided by the vocals, rather than the reverse.

Specifically, editors should set up three primary tracks: vocals on the first, music on the second, and sound effects on the third. Set vocal volume as the baseline, keeping music and effects lower; typically, music peaks should not exceed seventy percent of average vocal volume, and effects should not exceed fifty percent. While not absolute, these ratios serve as a starting reference. Editors should use volume automation curves rather than simple global adjustments, as vocal intensity varies across segments.

The risk is that placing music before vocals causes speech to follow the musical rhythm, creating a mismatch between speech pacing and visual timing. The test is whether every word remains intelligible with eyes closed, without music distracting attention. Exceptions include transitions or emotional climaxes where music briefly overpowers vocals, but such moments must last under three seconds and be noted in the script beforehand.

After editing, deliver a preliminary mix with vocal priority for the brand's review. This version does not need fine tuning, but it should let the brand confirm whether vocal clarity and placement meet expectations. If the brand feels a section of vocals is unclear, return to the original footage to check rather than simply raising the volume, because increasing volume will amplify background noise.

Use compression and EQ during the mix stage to make vocals stand out.

Mixing is the core technical stage for setting vocal priority. The mixer uses a compressor to control the vocal dynamic range, reducing the gap between quiet and loud parts so overall levels stay more stable. The compression ratio is usually set between 2:1 and 4:1, with a fast attack and moderate release, so the vocals do not sound unnatural.

Specifically, the mixer first reduces noise on the vocal track, removing continuous background noise and electrical hum. Then the mixer uses EQ to boost the 2,000 to 4,000 Hz range, which is key to vocal clarity. At the same time, cut the muddy 200 to 300 Hz range and sibilance above 8,000 Hz. The mixer also uses a de-esser to keep s and t sounds from becoming too harsh.

The risk is that over-compression makes the sound lose its naturalness, and over-boosting the high end makes it harsh. The standard is to audition on ordinary speakers and phone speakers; vocals should remain clear, not muffled or harsh. The exception is when vocals are already recorded well, the mixer can do less processing and only make minor adjustments. If the vocal source quality is poor, the mixer should inform the brand in advance that post-production can only improve, not fully fix, the issue.

Mixers should also pay attention to sidechain compression for music and sound effects. Sidechain compression means that when the voice appears, the volume of music and sound effects automatically decreases, and returns when the voice stops. This technique ensures the voice always stays at the front of the mix, but the trigger threshold of the sidechain compression should be set reasonably, so that the music doesn't noticeably pump when the voice appears, which would look unprofessional.

Before delivery, check voice loudness according to platform requirements.

Commercial videos are ultimately published on different platforms, and each platform has different requirements for voice loudness. The brand should confirm the target platform's loudness standard before delivery, usually based on LUFS measurements. If the platform has no clear standard, general loudness specifications can be referenced, but the platform's latest official requirements should prevail.

Specifically, after mixing is complete, use a loudness meter to check the loudness of the entire video, ensuring the voice portion is within the standard range. If the loudness is too high, the platform will automatically compress it, causing distortion. If the loudness is too low, viewers will need to turn up the volume, which is a poor experience. Also check whether peaks exceed 0 dB to avoid clipping.

The risk is that many projects only check the loudness of the entire video without separately checking the loudness of the voice track. If the voice is drowned out by music, the overall loudness may be normal but the voice is inaudible, making the mix useless. The criterion is to listen on both phone speakers and headphones; the clarity of the voice should be consistent. The exception is that if the video is being published on multiple platforms simultaneously, and each platform has different standards, multiple masters should be output rather than using one version universally.

Delivery must also include an audio specification document recording the original format of vocal tracks, mixing settings, loudness values, and platform-adapted versions. This document helps the brand quickly locate issues during later revisions without having to listen to the entire film again.

During acceptance, vocal clarity must be confirmed segment by segment.

During acceptance, the brand must not only look at the visuals, but specifically inspect the audio. The acceptance checklist should include whether vocals are clear, volume is stable, there is background noise, music overpowers, and sound effects are jarring. Each item must have clear judgment criteria rather than relying on feeling.

Specifically, the brand should prepare an acceptance form listing each scene or timecode, the corresponding vocal type, and the expected effect. During acceptance, play the video and record the perceived sound of vocals segment by segment, such as a segment being too quiet, having echo, or being covered by music. Then send this feedback to the production team and request revisions.

The risk is that brands often focus only on color grading and editing pace, neglecting audio issues. If vocals are found unclear only after release, revisions require remixing, which is costly and time-consuming. The judgment standard is that during acceptance, the audio should be tested on at least three devices, including professional studio monitors, standard laptop speakers, and mobile earphones. Vocals are only considered acceptable if they are clear on all three devices.

The exception is, if the project is pure music or pure sound effects with no dialogue, then the acceptance checklist does not need a human voice item. But if there is narration or interviews, human voice acceptance must come first. After acceptance, save the final mix project file to facilitate later revisions.

Human voice priority does not apply to all commercial videos.

Human voice priority is not a cure-all; some commercial video types are not suited to forcing voice priority. For example, in pure product showcase videos, the visuals are the star, music and sound effects set the atmosphere, and the human voice may only be a brief product name. In such cases, forcibly raising the narration volume would actually weaken the visual impact.

The specific criterion is the video's communication goal. If the goal is to make viewers remember a line, the human voice must be prioritized. If the goal is to make viewers remember the product's appearance or usage scenario, the visuals take priority and the voice supports. If the goal is to build brand emotion, music and sound effects take priority and the voice can step back. The brand side should reach consensus with the production team during the project initiation stage, rather than waiting until post-production to make changes.

The risk is that some brand clients want every video to prioritize the human voice, causing music and sound effects to be pushed too low and the video to lose its appeal. The exception is fast-paced videos on short-video platforms, where the human voice is often very short or even absent, so voice priority does not need to be set. Also, for AIGC-generated videos, if the voice is synthesized, its clarity may not match a real recording and requires additional processing.

When proposing a project, the production team should actively ask the brand about its communication scenarios and determine whether voice priority is appropriate. If the brand cannot explain clearly, ask for competitor references or placement data to help make the judgment. If the brand insists on voice priority but the video content itself is not suitable, the production team should propose an alternative solution rather than blindly executing.

The next step is to start with a sound test.

If you are preparing a commercial video project, do not jump straight into formal production; instead, run a voice-priority test first. Select a thirty-second sample clip, set voice priority according to the methods described in this article, mix it, and send it to the internal team for review and feedback. This test can help you identify issues with recording equipment, mixing workflow, and acceptance standards ahead of time, avoiding pitfalls in the formal project.

When testing, do not only test ideal conditions; simulate real shooting conditions, such as background music, ambient noise, and multiple speakers. After the test, compile a voice-priority specification as part of the project documentation. This specification can be used for all subsequent video projects, reducing communication costs.

If you need professional support, contact the video production team for audio mixing case studies and process documentation. But remember, no team can guarantee perfectly clear dialogue, because sound quality depends on shooting conditions, equipment, and post-production. What you should do is define acceptance criteria in the contract—such as dialogue clarity, loudness range, and number of revisions—so there is a basis for delivery.

If you are preparing a commercial video production project, start by organizing the Brief, reference visuals, product or company materials, delivery platforms, and copyright scope, then visitthe video production solutions page, moving communication from abstract preferences to executable production boundaries.