Write speaker, action and distance together
For two-person dialogue, define who speaks, how the listener reacts and how room ambience sits around them.
Update your browser, then reload this page. Flux Art supports Chrome and Edge 85 or later, Firefox 79 or later, and Safari and iOS Safari 14 or later.
Reload Flux ArtCheck your network connection and reload this page. If the problem continues, clear your browser cache and try again.
Reload Flux ArtAI E-commerce product suites are now live
Upload a product image, choose a platform and the images you need, then create listing-ready main images, white-background images, selling-point images, and lifestyle scenes. Product suites currently support Taobao/Tmall, JD.com, Pinduoduo, and Douyin E-commerce, with every result saved together for easy review, refinement, and export.
Generate image, Chinese or English speech, effects and ambience together in clips up to 10 seconds.
Kling V2.6 is not about attaching a track to silent footage; it aligns dialogue, action sound and ambience with visible events in the same generation.
For two-person dialogue, define who speaks, how the listener reacts and how room ambience sits around them.
Place falling beans, machine startup and selling-point narration on distinct time cues instead of stacking everything.
Rain, splash and bus brakes each connect to a visible event so the soundscape feels grounded.
The input route defines existing visual facts. The prompt should supply motion and change instead of competing with the source.
Generate image, speech, effects and ambience together from text.
The frame owns subject and composition; the prompt owns motion and sound events.
Pro mode uses two boundary images to constrain visual result and sound timing.
Use it for short ads, character dialogue, product narration and environment-led shots where sound carries narrative information.
Name the speaker, write the line and visible action, then specify spatial distance and emotion separately.
Bind selling-point narration to product action while controlling mechanical, material and room sound.
Tie rain, traffic, wind or crowd sound to visible sources instead of unexplained background noise.
In Pro mode, use two boundary images to constrain the result while writing transition action and sound cues.
Write in three columns—event, camera and sound: what is visible, how the camera moves and what is heard at that moment.
The first frame owns teahouse, reporter and tea master. Eight seconds, slow medium push. As he pours tea, the master says in Chinese, 'When the rain is heavy, the tea fragrance stays.' Keep natural speech, porcelain contact, rain outside and a distant river boat; no music.
Six-second product shot. Beans fall into a cream grinder as camera moves slowly from side to front. Female narration: 'Freshness begins with every bean.' Add falling beans, burr startup and subtle walnut resonance; preserve product proportions and mark placement.
The first frame owns the red shelter, parent and child. The child in yellow boots jumps a puddle as the parent steps back under a clear umbrella; low lateral camera. Rain hits glass, boots splash and bus brakes approach from the right; no dialogue or music.
Kuaishou released Kling Video 2.6 on December 3, 2025, officially describing text- and image-led audiovisual generation, Chinese and English voice, dialogue, narration, singing, effects and ambience up to 10 seconds. Flux Art exposes text, first-frame and boundary-frame routes; boundary frames use Pro mode, with negative prompts available.
Create with Kling V2.6Model capabilities are summarized from public Kling AI and Kuaishou materials. Available modes, parameters and pricing follow the current Flux Art workspace.
Published
Kling Video 2.6 is the model Kling released on December 3, 2025 with simultaneous audio-visual generation from text or images, including speech, effects and ambience.
Flux Art currently provides text-to-video, first-frame video and start-and-end-frame video. The boundary-frame route is available in Pro mode.
Kling's official release describes output up to 10 seconds. Available workspace controls remain the source of truth if the provider changes its options.
The official release states support for Chinese and English voice generation.
Kling officially lists speech, character dialogue, narration, singing, rap, ambient effects and mixed sound effects.
Yes in the current Flux Art Pro route. Use two boundary images and describe the transition action and sound cues between them.
Name the speaker, place the line at a clear moment, describe visible mouth or body action, and specify the distance and emotion of the voice.
Tie each sound to a visible cause and sequence events instead of asking all dialogue, effects and ambience to begin at once.
Commercial use depends on current Flux Art and model-provider terms and your rights to prompts and reference media. Check the latest terms before client or advertising use.