One Shoot, Many Markets: The Product Video Problem Localisation Never Solved
Anyone who has taken a store beyond its home market knows the sequence. The product pages get translated. The currency switches. The size chart converts. Then someone opens the product video, and the whole effort stalls on a person who is unmistakably from somewhere else.
Text localises cleanly because words are separable from the thing they describe. A presenter is not. The face on screen carries a nationality, an age, a style, and a set of assumptions about who the product is for, and none of that translates with a subtitle track.
The Economics That Kept It Unsolved
The obvious fix is to shoot the video again with a local presenter. Most stores never do, and the reason is simple arithmetic.
A product video is a single fixed cost that the store expects to spread across every unit it sells. Shooting a second version for a second market doubles that cost while the second market, at launch, is usually a fraction of the first. Shoot a third and a fourth and the video budget grows linearly with the number of markets while the revenue behind each one starts small. So stores compromise: one presenter everywhere, subtitles on top, and a quiet acceptance that the conversion rate in the new market will be lower than it needs to be.
That compromise has held for years because there was no third option between reshooting and not reshooting.
What Actually Needs to Change on Screen
It helps to be precise about which parts of a product video carry market-specific weight, because it is fewer than people assume.
The product itself is the same everywhere. The hands demonstrating it are the same. The pacing, the camera moves, the moment the lid comes off and the moment the texture is shown: all of that was decided once and does not need to be decided again. What changes is the presenter’s identity, which lives almost entirely above the collar. Hair, face, the outline of the head against the background. Below that line the video is already localised, because a hand holding a bottle does not have a nationality.
Seen that way, the reshoot is enormously wasteful. It rebuilds an entire production to change the one region of the frame that was ever a problem.
Replacing the Head Rather Than the Shoot
Tools built for head swap ai now do exactly that narrow job: given one photograph of the local presenter and the finished video, they replace the head in every frame and leave the body, the hands, the product and the motion exactly as shot. iMideo’s implementation tracks the head frame by frame rather than pasting a still onto each one, which is the difference between a result that survives the presenter turning to the camera and one that does not.
The reason this has to be a head replacement rather than a face replacement matters for the localisation case specifically. Face swaps keep the original hairline and head shape. For a viewer in a different market, hair is often the strongest cue that the presenter is not local, and a swapped face under the original hair produces something that reads as wrong without the viewer being able to say why. Replacing the whole head removes the cue instead of half-hiding it.
The Workflow This Makes Possible
With the cost of a localised presenter reduced from a shoot to an edit, the production model inverts. Instead of shooting once and living with it, a store shoots once with the intention of producing variants, and that intention changes how the original is made.
The presenter’s body language should be neutral enough to belong to anyone: no gestures that read as regional, no jewellery that anchors the person to a culture, clothing that is plain and unbranded. The head should stay in frame and reasonably lit throughout, because that is the region being replaced. Product close-ups can be as expressive as the store likes, since they are shared across all versions.
Then, for each market, a local presenter provides a photograph and, where the audio matters, a voice track in the local language. The video that ships to each storefront has a presenter who looks like the customer, holding a product demonstrated exactly the way the brand demonstrated it once.
Where the Honesty Line Sits
Two rules keep this from sliding into something a store will regret.
The people whose heads appear must have agreed to it, in writing, for this purpose. A stock photo is not consent, and a former employee’s headshot is not consent either. The presenters in each market should be paid for their likeness the way they would be paid for an afternoon on set, because that is what they are providing.
And the replaced presenter cannot be made to say something the original did not. The demonstration, the claims, the sequence of what is shown: those were decided in the original shoot and they stay. Localisation changes who is presenting. It does not change what is being presented, and a store that uses the same technique to invent a testimonial has left localisation behind entirely.
What Changes for a Store Going Abroad
The product video used to be the last thing to localise and the first thing to be skipped. It sat at the wrong end of the cost curve, expensive to redo and easy to leave alone, so most international storefronts shipped with a presenter from headquarters and hoped nobody minded.
People did mind, quietly, in the conversion rate. The change is not that stores can now afford to reshoot for every market. It is that they no longer have to, because the part of the video that was ever foreign turns out to be small enough to swap.
Leave a Reply