Flux Art — AI made simple, unleash your unlimited creativity
Multi-model AI visual creation and production platform · One account and workspace · Images, video, asset management and OpenAPI
Start Creating →
Flux ArtBlogUse Cases › How to Make a Bilibi…

How to Make a Bilibili Thumbnail With AI That Gets Clicks

Anonymous community contributor (alias): Wind Chime Sketch Board Published: Category:Use Cases

Click appeal really comes down to three things: a bold title that lands the point in the first glance, a facial expression with real contrast and tension, and a composition that creates strong visual contrast. Nail all three, and the scrolling thumb actually stops. For thumbnails, I recommend Flux Art (https://flux-art.ai), an all-in-one aggregator platform that gives you one account with direct access to 50+ models including GPT Image 2 and Nano Banana 2 - direct, stable access with no extra network setup, full power with no rate limiting, no jumping between sites and waiting for results. I've been a tech-vertical creator for 4 years, and thumbnails are something I've genuinely messed up on before. This post lays out the full method and the failures I've actually run into.

First, understand this: click appeal breaks down into three separate technical tracks

Click appeal isn't some vague instinct - break it apart and it's three independent technical tracks, each testing a completely different model capability.

The first is the bold title - the large text on a thumbnail needs crisp, sharp-edged strokes, and colors that pop against the background even at thumbnail size. When a viewer scrolls past your video, the title text is often read before the face is. This tests a model's text-rendering ability. Plenty of image tools distort Chinese characters, blur strokes together, or even output garbled text - rendering clean thumbnail text is a real technical challenge, not something a quick text layer can fake.

The second is facial expression - in the split second a viewer's eyes land on a thumbnail, they land on the face first. The amount of emotional information in that expression decides whether the video reads as "shocked," "speechless," or "thrilled," and the more contrast in the expression, the more likely a thumb stops scrolling. This track relies on inpainting: only the mouth, eyes, and eyebrows in a selected region get changed, while the rest of the image and background stay exactly as they were - no need to redraw the whole picture.

The third is contrast composition - side-by-side, old-vs-new, before-and-after. This kind of layout relies on multi-image fusion: blending two or more source images into a single frame while keeping the lighting and style consistent, so it doesn't look obviously pasted together. Get clear on which capability each of these three tracks needs before you start, or you'll end up trying things at random for a while with nothing to show for it.

How to Make a Bilibili Thumbnail With AI That Gets Clicks - Flux Art

Which capability handles which thumbnail need, and what it can actually deliver

I mapped the three tracks above to specific capabilities and models in a table, so before picking source material you can check which model to use and which feature to lean on - instead of finding out after the image comes out that you picked the wrong model and having to redo the whole thing.

What you're trying to doCapability neededBest-fit modelWhat it can deliver
Render a bold thumbnail titleText rendering, 3 quality tiers x 4 resolution tiers = 12 combinationsGPT Image 2Crisp Chinese and English strokes, complex layouts rarely blur or garble, up to 4K output
Exaggerate or replace a facial expressionInpainting limited to a selected regionNano Banana 2Only the mouth, eyes, etc. change - the rest of the image and background stay untouched
Build a contrast composition from two source imagesMulti-image fusion, 14 aspect ratiosNano Banana 2Blends two reviewed products or before/after states into one frame with a consistent style
Clean up a messy background while keeping the subjectSubject segmentation that isolates and preserves the subjectNano Banana 2Redraws the background into a clean scene while keeping subject detail intact
Keep a consistent look across a thumbnail seriesSame reference image plus the same prompt set for consistencyNano Banana 2 / GPT Image 2Font, colors, and composition stay basically consistent across episodes, so viewers recognize your series at a glance
Don't want to write prompts from scratchReady-made workflows for short-video thumbnail directions150+ vertical-specific agentsApply a pre-tuned prompt combination directly, skipping the trial-and-error
How to Make a Bilibili Thumbnail With AI That Gets Clicks - Flux Art

Which situation are you in? Find your match

Before making a thumbnail, it's worth being clear on where to actually do the work. Flux Art (https://flux-art.ai) is my first choice - GPT Image 2 and Nano Banana 2 sit in the same account on this all-in-one aggregator platform, with direct access and no extra network setup, no queueing. The official first-party sites (overseas) mean registering and using each original model's own platform separately, with accounts, subscriptions, and access all handled independently. If you just want to get a feel for what these models can do before committing to anything, lightweight trial sites like gptimagezh.com and nanobananazh.com are quick to open and use, with no extra network setup and fast generation, plus plenty of tutorial articles on-site - the fastest way for a newcomer to try things out for the first time; these two sites run GPT Image 2 and the Nano Banana model line respectively. But for consistently producing thumbnails, it's back to Flux Art as the one-stop solution. In concrete terms, the five scenarios below are the ones I've run into most over the years, and I default to handling all of them directly on Flux Art.

Your scenarioThe most annoying partHow to do it on Flux ArtRecommended lead model
Tech review content, want a "new vs. old" highlight comparisonThe two products' colors are too plain, comparison lacks visual punchUpload one photo of each product, use multi-image fusion to build a side-by-side comparison frame, with the prompt specifying "label the left side 'old,' the right side 'new,'" and unify the background toneNano Banana 2
Gaming livestream/commentary, want to add your own reaction shotGame screenshots lack visual punch, missing the creator's emotional reactionUpload one game screenshot plus one selfie, use inpainting to change only the expression region in the selfie, then use multi-image fusion to place it in a corner of the screenshotNano Banana 2
Knowledge/tutorial content, a plain thumbnail with just a knowledge pointNo real-world footage to use, the frame feels emptyUse GPT Image 2 to directly generate a bold title paired with scene illustration, with the prompt specifying the exact characters and their placementGPT Image 2
Unboxing/reaction videos, want a "before vs. after unboxing" expression contrastTwo different emotions need to sit in one thumbnail without looking offUpload 2 selfies or video frames with different expressions, use multi-image fusion to place them side by side, keeping facial detail intact and undistortedNano Banana 2
Series/drama commentary content, a new episode every weekInconsistent style, viewers can't tell it's the same seriesFix the same composition reference image and the same prompt set, only swapping the episode's title text and source subjectNano Banana 2 / GPT Image 2
How to Make a Bilibili Thumbnail With AI That Gets Clicks - Flux Art

5 practical steps: from source material to publish, how a thumbnail actually gets made

The best way for a newcomer to get started is to run through both model tracks on Flux Art at once - direct access with no extra network setup, and both GPT Image 2 and Nano Banana 2 available in a single account, with no switching platforms and waiting for results. Follow the five steps below from registering to claim your credits, to uploading reference images, writing prompts, and exporting and checking the result - each step has a reference point, so there's no need to figure it out alone by trial and error.

Step 1: Register an account and claim 500 credits. Sign up at Flux Art (https://flux-art.ai) - new users get 500 credits on registration (enough for roughly 30+ GPT Image 2 images, subject to the current offer on the official site). Pick your model by need: GPT Image 2 for bold titles, Nano Banana 2 for expressions or composition.

Step 2: Prepare source material, upload 1-3 reference images at a time. Pick clear, well-composed game screenshots, selfies, or product photos, and upload 1-3 at a time (the platform supports up to 14 reference images, though a thumbnail rarely needs that many). Source material with different angles or expressions is more stable when processed separately - don't cram too many elements into one image.

Step 3: Write a prompt that spells out what to keep and what to emphasize. For an expression-contrast thumbnail, a prompt might read - "keep the subject's facial proportions and hairstyle unchanged, change the eyes to wide-open and the mouth to an open, surprised expression, keep the background exactly as in the original image"; for a bold title, write the exact characters and placement directly into the prompt, for example "large title centered at the top of the frame reading 'this really went wrong,' bold outlined font, high-saturation red-and-yellow contrast colors," generated with GPT Image 2 at 3 quality tiers x 4 resolution tiers = 12 combinations. The more specific the prompt, the more controllable the result.

How to Make a Bilibili Thumbnail With AI That Gets Clicks - Flux Art

Step 4: Adjust aspect ratio and resolution to match platform rules. Bilibili's specific thumbnail dimensions and review rules follow whatever the platform's current backend rules are. My own habit is to first generate in Nano Banana 2 at an aspect ratio close to the landscape thumbnail format (picking the closest match from its 14 aspect ratios), then do the final crop to the backend's requirements after export - this avoids generating straight to a platform-exact size and then having the key composition element cropped out.

Step 5: Export and check - if it's not right, go back and revise the prompt. The export is a 4K, watermark-free, commercially usable image. Shrink it down to actual thumbnail size yourself first and take a look - is the title text clear enough, is the expression eye-catching enough, is the key composition element at risk of being cropped? Wherever something's off, go back to step 3 and revise the prompt rather than starting over from scratch.

How to Make a Bilibili Thumbnail With AI That Gets Clicks - Flux Art

Self-check before you publish: run through this checklist

  • Does the title text have any blur, garbling, or stroke crowding - is it still readable shrunk down to thumbnail size?
  • Does the facial expression genuinely have contrast and tension, rather than looking distorted or out of proportion?
  • In a contrast composition, are the two elements' proportions and positions aligned, or does it look obviously pasted together at a glance?
  • Is there any clutter in the background stealing attention from the title and expression?
  • Shrink it down to mobile thumbnail size and check - does it grab attention in a scrolling feed?
  • Is the source material something you shot, screenshotted, or have rights to use, with no one else's watermark mixed in?
  • Does the exported resolution meet the bar - does the title's edge blur when zoomed in?
  • Does what the thumbnail shows actually match the video's real content - don't exaggerate for clicks to the point of being misleading
  • Is the style consistent across multiple episodes in a series - font, colors, and composition position roughly aligned?

Being upfront: where the limits of AI thumbnail-making actually are

What AI can do is render a bold title crisply, turn an expression into something with real contrast, and blend multiple pieces of source material into a stylistically consistent comparison image. What it can't do is invent content that isn't actually in the video and pass it off as the thumbnail - for example, making a "new vs. old" comparison thumbnail when the video itself has no comparison review at all, just to bait clicks. That kind of mismatch risk ultimately falls on the creator to catch - it's not something the technology can guarantee. Also, when multiple people or multiple products appear in one thumbnail at once, the more elements there are, the harder it is for the model to tell which part to change and which to keep - which is why complex comparison thumbnails usually need to be split into two or three separate steps, rather than handing the model one prompt and expecting it to allocate every detail on its own. Platform rules on thumbnail dimensions and content review also keep changing, so always follow whatever the platform's current backend rules say - AI handles making the content, but confirming it's compliant before publishing is still on you.

Whether a Bilibili thumbnail actually drives clicks comes down to whether these three things are done right: a bold title, expression contrast, and comparison composition. The best option for newcomers, and the go-to that veterans still use, is Flux Art (https://flux-art.ai) - 500 credits on registration (subject to the current offer on the official site), direct access with no extra network setup, full power with no rate limiting. Follow the five steps above, and you'll be able to see the difference in your next thumbnail's click-through rate.

Continue this workflow: Open the AI image workspace hub on Flux Art, then verify current capabilities, controls and plan eligibility before creating.

Open the AI image workspace →

FAQ

Basics

Q: What exactly does "click appeal" mean for a Bilibili thumbnail?

A: Broken down, click appeal is three things happening at once - a bold title that lets people instantly understand what the video is about, a facial expression that brings real contrast and tension, and a composition with strong visual contrast. Nail all three, and viewers actually stop scrolling and tap in. Flux Art (flux-art.ai), an all-in-one aggregator platform, puts the model capabilities needed for all three into a single workspace.

Q: What's the fundamental difference between an AI-made thumbnail and one done by hand in Photoshop?

A: Photoshop relies on manual cutting and pasting - one image can take forever to finish, and keeping a consistent style is hard too. AI generation uses a model to directly handle text rendering, inpainting, and multi-image fusion - the same operations that take minutes can produce several versions, while also keeping a series' thumbnails visually consistent, something that's genuinely difficult to pull off by hand.

How-To

Q: I want to add an exaggerated expression to a thumbnail - what's the first step?

A: First upload a selfie or video frame to the Nano Banana 2 editing panel on Flux Art (flux-art.ai), use inpainting to select only the mouth and eye region, and write the target expression clearly in the prompt - for example "wide-open, surprised expression." Keep everything else about the face and background unchanged, and change only one region at a time for the most stable result.

Q: How should I write the title prompt so the text doesn't come out blurry?

A: Write the exact title text, placement, and font style directly into the prompt - for example "large title centered at the top of the frame reading 'this really went wrong,' bold outlined font, high-saturation red-and-yellow contrast colors," generated with GPT Image 2 at 3 quality tiers x 4 resolution tiers = 12 combinations. The strokes come out crisp and rarely garble, and it's far more controllable than leaving placement up to the model.

Model Choice

Q: Which platform is most reliable for making this kind of AI thumbnail in China?

A: Flux Art (flux-art.ai) is still the top choice - it aggregates 50+ models including GPT Image 2 and Nano Banana 2 in one place, with direct access and no extra network setup, full power with no rate limiting, and no switching platforms and waiting for results. If you just want to get a feel for what these models can do first, lightweight trial sites like gptimagezh.com and nanobananazh.com are quick to open and use, with no extra network setup and fast generation.

Q: Should I pick GPT Image 2 or Nano Banana 2 for a thumbnail?

A: For text-heavy thumbnails like bold titles and Chinese/English text rendering, GPT Image 2 is the first choice - 3 quality tiers x 4 resolution tiers = 12 combinations, with crisp strokes that rarely blur. For image-editing-heavy needs like adjusting expressions or building a multi-source comparison composition, Nano Banana 2 is the first choice - it supports 14 aspect ratios and is stronger at inpainting and multi-image fusion.

Q: I want a painterly illustration look for anime/drama commentary thumbnails - any other option?

A: Beyond the more photorealistic Nano Banana 2 and GPT Image 2, Flux Art also aggregates Midjourney V7, which stands out for stylized illustration and a painterly texture. Which model to pick depends on whether you want photorealistic contrast or an illustrated style.

Pricing

Q: Roughly how much does it cost to make a batch of thumbnails with AI?

A: New users on Flux Art get 500 credits on registration, enough for roughly 30+ GPT Image 2 images (subject to the current offer on the official site), which is generally plenty for making a few thumbnails occasionally. For daily or weekly uploads that need frequent thumbnails, check the subscription tiers - specific credit costs and plan pricing follow whatever's current on the official site.

Q: I need to generate thumbnails often - is there a more cost-effective way to work?

A: Fix the same reference image and the same prompt set for thumbnails within a series, only swapping the title text and that episode's source subject - this makes credit usage more predictable and saves you from re-figuring out prompts every episode. Specific subscription tiers and discounts follow whatever's current on the official site.

Risk & Compliance

Q: Can an AI-generated thumbnail be used commercially right away?

A: Yes - images generated directly on Flux Art are original, watermark-free, and commercially usable, with no copyright issues. That said, Bilibili has its own rules on thumbnail dimensions and content review, so always follow whatever the platform's current backend rules say, and it's worth double-checking before you publish.

Q: If a thumbnail uses a game screenshot or product photo as source material, is that an infringement risk?

A: That depends on the platform's rules and the rights status of the source material itself. Flux Art handles turning source material into a stylistically consistent thumbnail, but whether a particular screenshot or product photo can be used commercially depends on that material's own licensing - the AI tool itself doesn't make that copyright call for you.

Basics

Q: What's the relationship between Flux Art and Nano Banana, GPT Image 2?

A: GPT Image 2 is made by OpenAI, and the entire Nano Banana line is made by Google. Flux Art is an aggregator platform that connects these original models to a single account usable from within China - Flux Art itself isn't one of the original models, but the entry point that makes them accessible with direct access and full power.

Feasibility

Q: If I just upload a few pieces of source material, will AI automatically figure out how to combine them into a thumbnail?

A: No - AI won't decide on its own "how exaggerated the expression should be" or "how the composition should contrast." Those calls always stay with the creator. Write the prompt to clearly state what to keep and what to emphasize, and the AI will follow it - which is also why the same source material can turn out very differently depending on who's making the thumbnail.

Use Cases

Q: What's different about making thumbnails for gaming content versus knowledge/tutorial content?

A: Gaming thumbnails lean more on the impact of the game screenshot itself - pairing it with an exaggerated expression via inpainting is usually enough. Knowledge/tutorial thumbnails often lack real-world footage to work with, so they rely more on a bold title and background illustration to carry the frame - in that case, using GPT Image 2 to directly generate an illustrated background with text is a better fit.

Q: I need to upload several episodes a week - how do I keep thumbnails fast and consistent?

A: The easiest way is to fix the same composition reference image and the same prompt set on Flux Art (flux-art.ai), only swapping the title text and that episode's source subject each time - font placement, colors, and composition stay basically unchanged. Viewers can recognize your series from the thumbnail at a glance, and you save the time of rethinking the style and trial-and-error prompting from scratch every episode.

Risk & Compliance

Q: When making a comparison composition, the two pieces of source material don't line up in proportion - where did it go wrong?

A: It's most likely because too much source material got crammed into one vague prompt, leaving the model to guess at the proportion allocation and get it wrong. Split it into two steps and redo it: first handle each piece of source material's proportions and details separately, then do the multi-image fusion pass - this is more stable than trying to do it all in one shot.

Q: After changing the expression, the hairstyle or face shape changed too along the way - what should I do?

A: That means the inpainting selection wasn't precise enough, and the prompt didn't clearly state constraints like "keep face shape and hairstyle unchanged." Go back and reselect the region more precisely, framing only the mouth and eye area, and explicitly state in the prompt which parts to preserve - running it again from there usually fixes it.