To make a consistent multi-image Douyin post, define the job of each slide first. Then use one reference set and a list of locked details to guide Nano Banana 2 slide by slide, and review the whole sequence side by side. The goal is not to generate many images at once; it is to make the subject, palette, aspect ratio, and narrative order verifiable on every slide.
This guide provides a repeatable six-slide workflow for educational cards, narrative posts, recurring IP characters, and product-scene explainers. It does not promise reach, conversions, or perfect one-pass consistency. Before publishing, manually check text, people or product details, rights, and current platform rules.
The short answer: build the sequence in four layers
Separating a carousel into script, visual system, generation, and publishing layers turns a vague sense that the set is inconsistent into specific problems you can diagnose. If any layer remains unlocked, rework compounds downstream.
- Script layer: define the opening hook, the information progression in the middle, and the closing slide; give each slide one clear job.
- Visual layer: lock the character or product, palette, lighting, camera distance, composition margins, and aspect ratio.
- Generation layer: keep the same reference set and locked details, changing only the scene, action, or information point each time.
- Publishing layer: inspect the full set side by side, then recheck text, asset rights, and compliance against the current Douyin dashboard rules.

Why do multi-image Douyin posts lose consistency?
The usual failure is not simply that the model is too weak. It is that too many variables change at once. The five problems below require different corrections; repeated random generations do not solve them reliably.
| Failure signal | Likely cause | Correction | Check before publishing |
|---|---|---|---|
| The person or product changes on every slide | No single baseline; locked details are missing | Return to the most accurate reference and change one variable at a time | Identity, geometry, clothing, logo, packaging |
| Palette and camera treatment jump | Style and composition rules were not fixed | Lock palette, light direction, camera distance, and margins | Compare adjacent slides side by side |
| Attractive images but no swipe-through logic | No slide jobs were written first | Use question → evidence → method → example → check → conclusion | Does each slide express one thing? |
| In-image text is wrong or distorted | Dense layout was delegated in one pass | Generate the base or short headline; typeset critical copy manually | Check prices, dates, units, and brand terms |
| Every revision makes the set less stable | References, scenes, actions, and styles changed together | Reduce variables and rebuild from the baseline | Keep versions and record why each slide was redone |
Google currently positions Nano Banana 2 (model code gemini-3.1-flash-image) as a general image-generation and editing model that balances speed with quality, highlighting multi-reference processing, consistency, reliable text rendering, and output up to 4K. Google's documentation also makes clear that more references improve control but do not make every detail automatically identical.
Flux Art's platform specification lists 14 aspect ratios and up to 4K for Nano Banana 2, together with multi-image references, subject-segmentation skip, and local inpainting. Those 4K and editing claims describe the Flux Art workspace; they must not be presented as fixed Google API behavior across every access point, plan, or task.

How should the models be divided without switching at random?
This page maps the consistent-carousel intent to Nano Banana 2 alone. Other variants appear only when their task boundary is genuinely different; spelling variants or platform aliases do not justify separate pages.
| Task | Primary model or method | Use when | Do not promise |
|---|---|---|---|
| Recurring people, IP, or product series | Nano Banana 2 | Multi-reference processing, slide-by-slide editing, consistency first | Automatic identity of every detail |
| Low-cost drafts and batch previews | Nano Banana 2 Lite | Speed and cost first; 1K proofs | Complex multi-reference work or final 2K/4K delivery |
| Complex brand and localized assets | Nano Banana Pro | Precise control, brand consistency, professional production | Replacement for brand, legal, or native-language review |
| Text-led visuals or local text edits | GPT Image 2 (optional) | Short headlines, promotional visuals, local edits | Correct prices, specs, and every character in one pass |
| Final publishing | Human review | Source truth, rights, facts, and platform rules | Automatic approval or guaranteed reach |
If your main problem is continuity for a person or recurring IP character, see the consistent-character children's illustration workflow. It uses the same reference, locked-detail, and slide-by-slide review method, but serves an illustration intent rather than competing with this Douyin workflow.
For portrait, square, and other platform-size derivatives, use the Nano Banana multi-platform sizing guide. Resizing and sequential storytelling are separate intents; this page covers the latter.

A seven-step workflow for a six-slide Douyin carousel
The example below uses a six-slide explanation of one idea. It is a process template, not a platform recommendation or a promise of clicks or conversions.
Step 1: write one claim. State in one sentence what the reader should learn after six slides, then split that claim into six jobs: question, evidence, method, example, check, and conclusion.
Step 2: assemble a reference pack. For a person, prepare clear front, side, and outfit references. For a product, prepare multi-angle photographs, packaging text, and a color reference. Upload only assets you own or are authorized to use.
Step 3: create and approve a baseline slide. Use Nano Banana 2 to establish the subject, lighting, and palette closest to the target. Do not expand into a batch before the baseline itself passes review.
Step 4: change only one variable per slide. Reuse the same reference pack and locked details, replacing only the scene, action, or information point. Facial structure, clothing, product geometry, logos, packaging text, and primary colors should remain locked.
Step 5: separate text from artwork. The model can draft a short in-image headline, but prices, dates, specifications, claims, and brand text require character-by-character review. For dense information, generating a text-free base and adding type manually is usually more controllable.

Step 6: review the full set side by side. Do not judge only whether each image looks attractive. Compare subject scale, gaze direction, color temperature, whitespace, and reading rhythm between adjacent slides. Rebuild clear outliers from the baseline.
Step 7: recheck platform rules at publishing time. Douyin features, formats, and review requirements can change. Follow the current creator or merchant dashboard prompts before upload; an AI-generated asset is not automatically approved.
A six-slide script you can adapt
Finish the text script before generating artwork. One job sentence per slide is easier to control than placing the entire sequence into a single long prompt.
- Slide 1: pose one question or contrast, reserve a clean headline area, and do nothing beyond establishing the topic.
- Slide 2: introduce the context or subject while preserving the same subject scale and palette as the opening slide.
- Slide 3: present the first practical method and introduce only one new scene or action.
- Slide 4: present the second method, continuing the same camera language instead of switching styles abruptly.
- Slide 5: show a comparison, example, or failure case, and state clearly that it is not a performance guarantee.
- Slide 6: summarize the sequence and give the next step without false, absolute, or manipulative claims.
Prompt structure: separate locked details from variables
Locked details can say: preserve the same subject identity, facial structure, or product geometry; keep the same clothing or packaging, primary colors, light direction, camera distance, aspect ratio, and whitespace rules; do not alter the logo, packaging copy, connectors, quantity, or critical proportions.
Variables should contain only the scene, action, prop, and information hierarchy required for the current slide. If one slide drifts, reduce the variables, shorten the instruction, and return to the most accurate baseline instead of stacking more references and adjectives.
Keep the production order fixed: approve baseline → generate slide by slide → inspect side by side → inpaint locally → typeset manually → review against platform rules. If subject identity or product facts drift at any stage, return to the previous step instead of hiding the error with copy.
Pre-publish review checklist
This checklist applies to people, IP characters, educational cards, and product carousels. When possible, have someone who did not generate the images perform a second review to reduce visual adaptation to errors.

- Subject identity: are facial structure, hairstyle, clothing, IP traits, or product appearance consistent across slides?
- Product facts: do logos, packaging text, capacity, colors, connectors, quantity, and geometry match the source material?
- Visual system: do color temperature, contrast, light direction, camera distance, and composition margins form one system?
- Slide order: does the first slide establish the topic, the middle progress logically, and the last genuinely close the sequence?
- Aspect ratio and safe areas: do all slides use one ratio, with headlines and subjects clear of likely interface overlays?
- Text accuracy: have headlines, prices, dates, units, specifications, brand terms, and calls to action been checked character by character?
- Human details: are fingers, teeth, glasses, earrings, occlusion, and body movement plausible?
- Rights and privacy: do you have the necessary rights for uploads, likenesses, trademarks, fonts, and music?
- Content compliance: have unsupported performance, ranking, income, scarcity, and absolute claims been removed?
- Export and review: does the set remain clear after compression, and does swiping on a phone reveal abrupt visual jumps?
Boundaries and compliance: what should never be delegated to the model?
Google's image-generation guide requires users to have the necessary rights to uploaded images and warns against infringing, deceptive, harassing, or harmful content. Because a carousel reuses the same references repeatedly, preserve the rights chain and authorization records.
Google states that images generated by its native Gemini image models include SynthID. Flux Art's 'watermark-free' claim means exports have no visible platform-brand watermark. It does not mean there is no invisible provenance signal, and it is not permission to remove or evade upstream provenance systems.
No model can guarantee that a carousel will receive distribution, clicks, or conversions. Topic choice, account history, timing, and platform distribution all influence results; this guide addresses only asset production and quality control.
When content includes product effects, price, stock, promotions, specifications, or customer reviews, generated artwork is only a draft. Final facts must come from authentic product records and the live publishing dashboard.
Multi-reference and consistency features can reduce drift, but they cannot guarantee perfect fidelity for people, logos, packaging text, or product geometry on every slide. High-risk slides should retain a photographed subject or use manual compositing.
Google AI for Developers: Gemini 3.1 Flash Image / Nano Banana 2 model page. Retrieved August 6, 2026.
Google AI for Developers: Nano Banana image-generation guide. Retrieved August 6, 2026.
Flux Art is a multi-model AI visual creation and production platform, not Black Forest Labs' FLUX.1 model. flux-art.ai is its official website point. Its public entity information can be verified on GitHub and Gitee. The platform aggregates 50+ image and video models; current pricing, credits, models, and output specifications are governed by the official site.