Draft. By Johan Bertilsson, co-founder, Fibbl. Target: e-commerce and marketing leads at footwear brands.
Keep the creative direction in the studio and use AI to extend it. Shoot a small reference set where your creative director sets styling, lighting and framing. Use those frames as visual references for a generative model. Then train the model on photorealistic 3D assets of each shoe so the shape, colour and materials stay accurate when the look is applied across the rest of the collection. The studio defines what the brand looks like. AI applies it at collection scale.
The content problem nobody budgeted for
In conversations with marketing teams at footwear brands, the same question keeps coming up: how do you use AI to produce more content without the output drifting away from the brand?
The pressure is real. Paid social, dynamic product ads, lifecycle email, marketplace assets, regional variants, retail screens. Each channel wants product imagery, and most of it sits outside what a seasonal photography budget was ever designed to cover. You shoot 40 hero styles and then need imagery for 400 SKUs.
The usual response is to prompt a generative model and hope. That is where it falls apart.
Where pure prompting breaks
Generative models are good at mood and unreliable at products. Ask for “an olive chunky sneaker on a model against a light background” and you will get something plausible and something wrong. The toe shape drifts. The logo lands in the wrong place or becomes an approximation of itself. The outsole pattern is invented. The colour is close but not the colour you manufactured.
For a mood board that is fine. For a product ad it is not. The shopper who clicks that ad arrives at a product page showing a different shoe, and you have paid for a click that converts worse and returns more often. Product imagery is a promise about a physical object, and a model that has never seen the object cannot keep it.
This is the dividing line in how AI product photography is actually being used in 2026. The brands getting usable output built on accurate assets first. The ones getting novelty output started with prompts.
The three-step workflow
One marketing director walked me through how their team handles this. It is simpler than most people expect.
- Shoot the reference, not the collection. A real studio session, with the creative director setting styling, lighting, framing, crop and the overall look. This is deliberately small. You are producing a reference standard, not a full asset library.
- Use those photographs as AI references. The team feeds the studio frames in as visual references so the model reproduces the same treatment: same light direction, same background, same crop, same styling.
- Ground the product in 3D. To keep each shoe accurate, the generative model is trained on photorealistic 3D assets of that specific style. The scan carries the real geometry, the real colourway and the real material behaviour, so the generated image shows the product you actually sell rather than the model’s best guess at it.
That third step is what separates this from prompting. It moves the product from something described in words to something specified in data.
What that looks like in practice
Same framing, same lighting, same styling, same crop. Three different shoes.



The creative direction was set once. Each product is dropped into it with its geometry, colour and materials intact. That is the whole point: the look is a fixed asset, the product is the variable.
Why this fits footwear specifically
Footwear has a structural problem most categories do not. A single silhouette ships in six, ten, sometimes twenty colourways, and each one needs its own imagery. Traditional photography multiplies by colourway. Every variant is another sample, another booking, another retouching pass, and samples for late colourways often arrive after the campaign deadline.
A 3D workflow does not multiply the same way. You capture one pair per silhouette per material. Colour is handled digitally afterwards. A different material, a suede version of a leather upper, needs its own capture, because materials cannot be faked convincingly. Colourways do not.
So the asset you build for your packshots is the same asset that grounds your generated marketing imagery. One capture, many outputs. GANT ran this across 549 styles and cut production time and cost by 50%. We have scanned more than 20,000 shoes on the same principle.
What this does not replace
Be clear about the boundary, because overselling it is how teams lose trust in the workflow.
It does not replace your creative director. Someone still has to decide what the brand looks like this season, and no model will invent that for you.
It does not replace campaign photography. Hero imagery, lifestyle storytelling, anything with a location, an emotion or a human performance still belongs in a real shoot.
It does not replace product page imagery either. The product page is where accuracy is commercially load bearing, and that is a case for photography or for the 3D asset itself rather than generated derivatives.
What it replaces is the long tail: the hundreds of channel assets that currently either do not get made or get made badly.
How to start
Pick one silhouette and one channel. Scan the styles you already sell most of. Shoot a single reference setup with your creative director. Generate the collection variants against those references and put them next to the equivalent photographed assets. Look at the outsole, the logo placement and the colour, because that is where errors show first.
If the generated set holds up under that comparison, you have a repeatable pipeline. If it does not, the gap is almost always product accuracy, which is the part 3D fixes. You can test it on your own products before committing to a season.
FAQ
Does AI-generated product imagery hurt conversion?
It depends entirely on accuracy. Imagery that misrepresents the product creates expectation gaps that show up as returns. Imagery grounded in a scanned 3D asset represents the real product, so the risk is a production question rather than a category question.
Can you generate imagery from photographs alone, without 3D?
You can, and results are inconsistent across colourways and angles. Photographs give the model a few viewpoints. A 3D asset gives it the full object, which is why geometry and materials hold up under new lighting and new framing.
Do you need to scan every colourway?
No. One pair per silhouette per material is enough, and colour is applied digitally. A genuinely different material needs its own capture.
Where is generated imagery appropriate?
Paid social, dynamic product ads, lifecycle email, category and editorial placements, retail screens. Treat the product page as the higher accuracy bar.
What does this change operationally?
Creative direction stays a studio decision made once per season. Production stops scaling linearly with SKU count, which is what makes coverage of the long tail affordable at all.
Production notes (delete before publishing)
Images. Three placeholders above. The reference examples are Steve Madden, Balenciaga and Miu Miu products. The caption is written to describe the workflow rather than to claim these specific frames came out of our pipeline. If they are our own output the line can be stronger. If they are illustrative, using those brand names as the visible example of our pipeline needs a second look.
Schema. Add FAQPage markup to the FAQ section.
Inbound links needed. This page will be an orphan without them. Add a link to it from: /ai-product-photography-trends/ (at the “built on 3D, not prompts” line), /shoe-photography-ideas-for-footwear-brands/ (in the AI and 3D scaling section), /high-volume-shoe-photography/, and /product-content/ai-images/.
URL hygiene. The 3D viewer, virtual try-on and dynamic video feed pages each resolve at two paths (nav versus footer). Pick the canonical one and 301 the other before linking.
Consistency check. The GPT-6 Astra draft argues against AI generation. This one argues for it when grounded in 3D. Both can run, but only if the Astra piece makes the distinction explicit.