Back to Home

Kling V2.6

Generate image, Chinese or English speech, effects and ambience together in clips up to 10 seconds.

Explore simultaneous audiovisual creation
Native soundtrack

Write why a sound happens before the sound itself

Kling V2.6 is not about attaching a track to silent footage; it aligns dialogue, action sound and ambience with visible events in the same generation.

Up to 10 secondsChinese and English voiceDialogue and narrationEffects and ambiencePro boundary frames
Character dialogue

Write speaker, action and distance together

For two-person dialogue, define who speaks, how the listener reacts and how room ambience sits around them.

Product narration

Give narration and product action one rhythm

Place falling beans, machine startup and selling-point narration on distinct time cues instead of stacking everything.

Environmental sound

Use sounds with visible causes

Rain, splash and bus brakes each connect to a visible event so the soundscape feels grounded.

Sound inputs

Decide which input owns the information

The input route defines existing visual facts. The prompt should supply motion and change instead of competing with the source.

T2AV

Text to audiovisual

Generate image, speech, effects and ambience together from text.

FIRST

First-frame audiovisual

The frame owns subject and composition; the prompt owns motion and sound events.

PRO A→B

Pro start and end

Pro mode uses two boundary images to constrain visual result and sound timing.

Sound-picture control

Where Kling V2.6 fits

Use it for short ads, character dialogue, product narration and environment-led shots where sound carries narrative information.

01

Chinese and English dialogue

Name the speaker, write the line and visible action, then specify spatial distance and emotion separately.

02

Product narration and effects

Bind selling-point narration to product action while controlling mechanical, material and room sound.

03

Ambience-led scenes

Tie rain, traffic, wind or crowd sound to visible sources instead of unexplained background noise.

04

Pro boundary transition

In Pro mode, use two boundary images to constrain the result while writing transition action and sound cues.

Sound scripts

Kling V2.6 audiovisual briefs

Write in three columns—event, camera and sound: what is visible, how the camera moves and what is heard at that moment.

Image video with Chinese dialogue

Rainy teahouse interview

The first frame owns teahouse, reporter and tea master. Eight seconds, slow medium push. As he pours tea, the master says in Chinese, 'When the rain is heavy, the tea fragrance stays.' Keep natural speech, porcelain contact, rain outside and a distant river boat; no music.

Product audiovisual ad

Coffee grinder narration

Six-second product shot. Beans fall into a cream grinder as camera moves slowly from side to front. Female narration: 'Freshness begins with every bean.' Add falling beans, burr startup and subtle walnut resonance; preserve product proportions and mark placement.

Ambience-led image video

Rainy bus-stop ambience

The first frame owns the red shelter, parent and child. The child in yellow boots jumps a puddle as the parent steps back under a clear umbrella; low lateral camera. Rain hits glass, boots splash and bus brakes approach from the right; no dialogue or music.

Mix check

How to use Kling V2.6 in Flux Art

Kuaishou released Kling Video 2.6 on December 3, 2025, officially describing text- and image-led audiovisual generation, Chinese and English voice, dialogue, narration, singing, effects and ambience up to 10 seconds. Flux Art exposes text, first-frame and boundary-frame routes; boundary frames use Pro mode, with negative prompts available.

Create with Kling V2.6
  1. Choose text or image sourceStart from text, a first frame or Pro boundary frames, defining subject, space and final state first.
  2. Write sound by eventFor dialogue, narration, effects and ambience, specify source, timing, distance and intensity.
  3. Review image and sound separatelyCheck lip rhythm, action effects and environmental space, then revise only the mismatched layer.
Source notes

Official sources and page information

Model capabilities are summarized from public Kling AI and Kuaishou materials. Available modes, parameters and pricing follow the current Flux Art workspace.

Published

Kling V2.6 FAQ

What is Kling V2.6?

Kling Video 2.6 is the model Kling released on December 3, 2025 with simultaneous audio-visual generation from text or images, including speech, effects and ambience.

Which Kling V2.6 workflows are available in Flux Art?

Flux Art currently provides text-to-video, first-frame video and start-and-end-frame video. The boundary-frame route is available in Pro mode.

How long can Kling V2.6 videos be?

Kling's official release describes output up to 10 seconds. Available workspace controls remain the source of truth if the provider changes its options.

Which spoken languages does Kling V2.6 support?

The official release states support for Chinese and English voice generation.

What audio can Kling V2.6 generate?

Kling officially lists speech, character dialogue, narration, singing, rap, ambient effects and mixed sound effects.

Can Kling V2.6 use start and end frames?

Yes in the current Flux Art Pro route. Use two boundary images and describe the transition action and sound cues between them.

How should I write dialogue prompts?

Name the speaker, place the line at a clear moment, describe visible mouth or body action, and specify the distance and emotion of the voice.

How can I improve audio-video synchronization?

Tie each sound to a visible cause and sequence events instead of asking all dialogue, effects and ambience to begin at once.

Can Kling V2.6 results be used commercially?

Commercial use depends on current Flux Art and model-provider terms and your rights to prompts and reference media. Check the latest terms before client or advertising use.