Beyond Aesthetic Quality: A 12-Output Pilot Study of E-commerce Image Editing
Summary
E-commerce image generation is not mainly an aesthetics problem. A result can look polished and still be unsafe to use because the bottle is taller, the cap has changed shape, the SKU color has shifted, or mandatory copy is wrong. In this pilot, I used one authorized synthetic skincare packshot, three editing tasks, and four models to examine that gap.
Key takeaways
- This pilot contains 12 outputs: four models × three e-commerce editing tasks.
- All four models produced credible creative imagery for background replacement and lifestyle scenes.
- In this sample, two of four models followed the requested packaging text layout exactly; the other two introduced a layout deviation.
- A visually attractive output is not evidence of SKU-level fidelity. Product detail and mandatory text still require review.
This is a small, manually reviewed field study, not a leaderboard or a statistically significant benchmark. It is intended to make the evaluation criteria explicit and invite discussion about what an e-commerce-ready image model should be measured on.
Why visual quality is the wrong first question
When a designer evaluates an image model, the first reaction is usually aesthetic: does it look realistic, well lit, and on-brand? That matters. It is not sufficient for commerce.
For a merchant, a product photo is also a statement of fact. The bottle shape, cap material, label copy, pack size, color, and logo can all be commercially meaningful. A model that improves the mood of a product photo while quietly changing those details may be useful for a campaign concept, but it is not automatically suitable for a product-detail page.
The useful production distinction is not good image versus bad image. It is creative-safe versus SKU-safe. A creative-safe image can be used as a starting point for social, ads, or moodboards. A SKU-safe image preserves product identity well enough for a stricter workflow, usually with human approval.
Pilot protocol
Input
I used one synthetic, unbranded cobalt-blue skincare bottle with a matte ivory cap. Using a synthetic asset avoids third-party product or logo claims while still exposing the kinds of failures that matter in real catalogues.
Figure 1. Source packshot — authorized synthetic, unbranded product image used as the common input for all 12 edits.
Models and tasks
The pilot used four image-editing models accessed through PixPix:
| Model | Background replacement | Lifestyle scene | Packaging text edit |
|---|---|---|---|
| Qwen Image Edit | 1 output | 1 output | 1 output |
| GPT Image 2 | 1 output | 1 output | 1 output |
| Nano Banana Pro | 1 output | 1 output | 1 output |
| Seedream 5.0 Pro | 1 output | 1 output | 1 output |
The three tasks were deliberately common in e-commerce production:
- Background replacement — preserve the product while moving it into a pale travertine bathroom setting.
- Lifestyle scene generation — place the product on wet dark slate with water and botanical shadows.
- Packaging text edit — add a white label with the exact three lines
NOVA,HYDRATING SERUM, and30 ML.
Each model received the same source image and equivalent English task instruction. One 1:1 output per model/task was reviewed manually. I did not select a favorite from multiple generations.
What I reviewed
The review looked at four practical questions:
- Product identity: Did bottle color, geometry, cap, position, and blank surface remain recognizably consistent?
- Instruction following: Did the requested scene and exclusions appear?
- Physical plausibility: Do the light, contact shadow, water, and reflections make visual sense?
- Text fidelity: Are the requested words and line breaks rendered exactly?
Because this is only 12 outputs from one SKU, the observations below should not be interpreted as general model rankings.
Result 1: scene generation was strong, but product preservation was not exact
All four models created plausible bathroom scenes. The product remained recognizable, the blue glass color survived, and the environments read as premium skincare photography. That makes this task promising for advertising concepts and merchandising experiments.
Figure 2. Background replacement — Qwen Image Edit.
Figure 3. Background replacement — GPT Image 2.
Figure 4. Background replacement — Nano Banana Pro.
Figure 5. Background replacement — Seedream 5.0 Pro.
The limitation is subtle but important: every output re-rendered the product to some extent. In several results, the cap-to-body ratio, the bottle scale, or the shoulder geometry shifted. Those differences are easy to miss in a campaign image. They matter when the generated image must represent a physical SKU precisely.
Pilot finding: In all four background-replacement outputs, the requested environment was credible and the product remained recognizable. None should be treated as pixel-identical preservation of the original packshot. For e-commerce, that makes this a creative-safe task, not an automatically SKU-safe one.
Result 2: lifestyle imagery increases creative value and review risk at the same time
The wet-slate task asked for more than a background. It required water, surface interaction, reflections, soft botanical shadows, and a product that still looked physically present in the scene.
Figure 6. Wet lifestyle scene — Qwen Image Edit.
Figure 7. Wet lifestyle scene — GPT Image 2.
Figure 8. Wet lifestyle scene — Nano Banana Pro.
Figure 9. Wet lifestyle scene — Seedream 5.0 Pro.
All four outputs are visually usable as a creative starting point. They contain convincing wet surfaces, reflections, or natural light. Yet the same creative freedom that produces an appealing image also increases the opportunity for object drift. A taller cap, a changed bottle shoulder, or a stylized reflection can make an image unsuitable for a product-detail page even when it looks excellent in a feed.
The practical workflow lesson is to route lifestyle generation toward ad creative, social assets, and early art direction. Keep approved studio packshots, regulated packaging, and precise product imagery in a separate workflow with stricter review gates.
Result 3: text was readable more often than it was layout-perfect
Text is where visual plausibility and commercial correctness diverge most clearly. The prompt required exactly three label lines. In this single-output sample, GPT Image 2 and Nano Banana Pro reproduced the three requested strings and line grouping. Qwen Image Edit introduced / separators, while Seedream 5.0 Pro wrapped HYDRATING SERUM across two lines.
| Model | Observed text result | Pilot interpretation |
|---|---|---|
| Qwen Image Edit | All words readable; added / separators |
Exact-layout failure |
| GPT Image 2 | Three requested lines readable | Pass in this one sample |
| Nano Banana Pro | Three requested lines readable | Pass in this one sample |
| Seedream 5.0 Pro | All words readable; middle line wrapped | Exact-layout failure |
Figure 10. Packaging-text task — Qwen Image Edit. The model added / separators, so the exact three-line layout failed.
Figure 11. Packaging-text task — GPT Image 2. The requested three-line layout is readable in this one output.
Figure 12. Packaging-text task — Nano Banana Pro. The requested three-line layout is readable in this one output.
Figure 13. Packaging-text task — Seedream 5.0 Pro. The middle phrase wraps, so the exact three-line layout failed.
Two of the four outputs matched the requested three-line text layout in this pilot. That is useful evidence for ideation. It is not enough evidence to approve prices, ingredients, claims, legal copy, or any other mandatory text without a deterministic design step and human review.
A production routing rule for e-commerce teams
The pilot suggests a simple routing rule that is more useful than asking for a universal “best” model.
| Task type | Recommended use | Required control |
|---|---|---|
| Background replacement | Ads, landing-page concepts, social creative | Review product shape and scale |
| Lifestyle generation | Campaign concepts and visual exploration | Review object geometry, lighting, and reflections |
| Text and packaging edits | Drafts and design exploration | Rebuild final text in a deterministic design tool |
| Product-detail hero images | Only after an identity-preservation review | Compare against approved source asset |
The goal is not to avoid image generation. It is to send each task to the level of review it deserves. A fast, visually strong model may be the right answer for a paid-social creative and the wrong answer for a regulated package label.
Limitations and the next version
This pilot is intentionally narrow: one synthetic SKU, three tasks, four models, and one output per cell. It does not measure variance across seeds, product categories, resolutions, prompt languages, real brand packaging, or video output. It also uses manual visual review rather than blinded raters or automated similarity metrics.
The next version should expand to at least six categories: cosmetics, apparel, jewelry, footwear, home goods, and consumer electronics. Each task should be repeated across several source assets and multiple outputs. Text evaluation should include exact-string matching, while product identity should use both reviewer scoring and reference-image comparison.
If the community finds this protocol useful, I plan to publish the task definitions, prompts, source-asset rules, and a larger result sheet as a Hugging Face Dataset. The product workflow itself can remain closed while the evaluation method remains inspectable.
Reproducibility appendix
Source asset: an authorized synthetic unbranded skincare bottle, shown above.
Background-replacement prompt: “Edit only the environment. Preserve the bottle's exact cobalt-blue color, matte ivory cap, cylindrical geometry, front-facing position, scale, and blank label-free surface. Replace the plain background with a premium pale travertine bathroom counter, soft morning window light, one out-of-focus green leaf in the far background, believable contact shadow and reflection. Do not add text, logos, extra products, hands, or change the bottle.”
Lifestyle-scene prompt: “Create a premium e-commerce lifestyle scene. Preserve the exact cobalt-blue bottle, matte ivory cap, proportions, blank label-free surface, and upright front-facing orientation. Place it on wet dark slate beside a small pool of water with realistic droplets and soft botanical shadows. The bottle must remain fully visible and physically plausible. Do not add text, logos, extra products, hands, or alter the bottle.”
Text-edit prompt: “Keep the exact same cobalt-blue bottle and matte ivory cap, including geometry, color, size, and centered front-facing composition. Add one clean white rectangular front label with exactly these three lines of readable black uppercase text: NOVA / HYDRATING SERUM / 30 ML. Use a minimal seamless light-gray studio background and one soft shadow. Do not add any other text, logo, or objects.”
Disclosure
I work on PixPix. The tests in this article were run through PixPix, a paid multi-model workflow for e-commerce image and video production. The purpose of the article is to publish the method, outputs, and limitations transparently; it is not a claim that any model is universally best.












