Headshots, full-body photos, scene images, camera-movement videos, music and the brand key visual (KV) are piled in front of the monitor, and the generation artist submits them with one instruction: “Use them all as references.” The result is a person wearing the clothes from the scene image, a camera-movement reference bringing its person into the generated shot, and a sound rhythm that throws the action out of order. More assets do not solve the problem; they create more opportunities for the references to compete for control.

Shenzhen TikTok video production should not rely on an unrestricted pool of assets. It should establish a reference-binding protocol: every asset needs a unique ID, a defined role, a priority, permitted inheritance and prohibited inheritance. We recommend that one asset handle one task whenever possible; the clearer the responsibility, the easier it is to trace the cause of rework.

1. First divide references into six roles

How Shenzhen TikTok video production binds multimodal references: original tutorial cover for the AIGC character and shot guide
How Shenzhen TikTok video production binds multimodal references: an AIGC character and shot guide for controlling generation and acceptance with visible production evidence.
  • Identity: The person’s face, product structure and body proportions.
  • Wardrobe: Clothing, hair and makeup, accessories and materials.
  • Scene: Spatial layout, set design, time and lighting.
  • Motion: Camera trajectory or action rhythm.
  • Style: Color, grain, material language and art direction.
  • Audio: Timbre, dialogue, ambient sound and musical beats.

One full-body photo can define body proportions and clothing, but it should not also determine the scene. If it must be reused, the manifest should clearly identify the primary and secondary roles.

2. Replace “image 1” and “image 2” with role labels

First name the files CHAR_A_FACE, CHAR_A_WARDROBE, LOC_LOBBY_NIGHT, CAM_ORBIT_SLOW and SFX_DOOR_CLOSE. Then state in the prompt: “Person A’s face references CHAR_A_FACE only; clothing references CHAR_A_WARDROBE only; CAM_ORBIT_SLOW provides camera movement only and does not inherit its person or scene.” Once the labels are stable, changing shot order will not scramble the references.

Google DeepMind Veo publicly demonstrates controls for scene, character and object references, as well as first and last frames. The production layer should further turn these inputs into asset roles instead of relying on an operator’s temporary memory.

3. Priority resolves conflicts; it does not replace QC

The recommended priority is: product or person identity > functional action > spatial continuity > camera > lighting and materials > stylistic decoration. When identity and style references conflict, preserve identity first; when a camera-movement video conflicts with a scene image, inherit movement from the camera reference only. Put the priority in the shot package so different generation artists do not make separate judgments.

A high priority does not guarantee a correct result. The face may remain stable while the hands still pass through objects; the product structure may be correct while the camera axis changes unexpectedly. Each role still needs independent QC.

4. Separate groups for multiple people and products

Putting four characters, three products and a complex scene into one shot easily causes identity swaps and unwanted duplicates. Generate the character group and product group separately, use an establishing shot to explain the space, and only then create interaction. Dialogue can be completed with separate over-the-shoulder singles, reverse angles, and two-person medium shots; there is no need to force one long take.

Use an independent headshot and full-body photo for each person. Do not use a nine-panel collage of multiple people as an identity reference. Always use stable names for references; do not write “he” or “the other person.”

5. Time binding determines whether sound can truly control the image

Sound references should mark events: reach the product close-up on the first strong beat, match the door-lock sound to completion of the hand action, and drop out the music when the brand line appears. If you provide only one music track without synchronization points, the model or editor can only guess the rhythm. Action references should likewise state the start, peak and end states.

Alibaba Cloud Model Studio (Bailian) provides an entry point to documentation about multimodal and generative AI. Whatever changes occur in a specific model interface, asset roles, timings and output acceptance should remain in the project’s own manifest so the workflow is not tied to one tool.

6. Reference-binding manifest template

Field Example
asset_idCHAR_A_FACE_V03
roleidentity
priority1
inherit Face, hairstyle and age
exclude Clothing, background and lighting
shotsS02,S03,S05
sha256 File checksum

After completing the manifest, use one establishing shot and one action shot for stress testing before expanding to the whole scene. For help collecting identity, scene, camera-movement and sound references, see Short-Video Production Solutions.

Get the multimodal reference-binding template