Command Palette
Search for a command to run...
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Abstract
We present Kandinsky 6.0 Video, a family of foundation difusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and imageto-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920×1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audiovideo data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and difusers integration under the MIT license.
One-sentence Summary
Kandinsky Lab introduces Kandinsky 6.0 Video, a family of foundation diffusion models comprising 3B-parameter Lite and 29B-parameter Pro variants for synchronized text-to-audio-video and image-to-audio-video generation, which uses a dual-stream CrossDiT architecture with bidirectional cross-attention, continuous pretraining, and reinforcement-learning-based post-training to deliver Full-HD output under an MIT license.
Key Contributions
- The paper presents Kandinsky 6.0 Video, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters), for synchronized text-to-audio-video and image-to-audio-video generation with 5-second clips, 44 kHz audio, lip-sync, and latent-space super-resolution to Full-HD 1920×1080.
- The models use a dual-stream CrossDiT architecture that connects a pretrained video stream with a from-scratch audio stream through bidirectional cross-attention, trained with continuous audio pretraining, joint audio-video pretraining, supervised fine-tuning merged into a model soup, reinforcement learning adapted from OmniNFT, and two-stage distillation to 10 function evaluations.
- Kandinsky 6.0 Video Pro outperforms Kandinsky 5.0 Video Pro and is preferred over LTX 2.5 in visual and speech quality, achieves best scores among evaluated models on VABench for speech quality, audio aesthetics, alignment, lip-sync, desynchronization, and visual realism, and reduces generated speech word error rate by 47%; code, checkpoints, and diffusers integration are released under the MIT license.
Introduction
Diffusion models and flow matching have become the dominant approach for image and video generation, but extending them to synchronized text-to-audio-video generation is substantially harder because it requires semantic, temporal, and emotional consistency across modalities. Closed commercial systems such as Veo 3.1, Sora 2, and Wan 2.6 achieve strong unified audio-visual synthesis, yet their source code is unavailable, while open-source models still lag behind in clip duration, resolution, and synchronization accuracy. The authors introduce Kandinsky 6.0 Video, a family of open foundation models with Lite and Pro variants that generate 5-second clips with synchronized 44 kHz audio, using a dual-stream CrossDiT architecture, continuous pretraining, and a multi-stage post-training pipeline for text-to-audio-video and image-to-audio-video generation.
Dataset
The authors build multiple datasets for Kandinsky 6.0 Video, mostly derived from the earlier Kandinsky 5.0 collection. The data covers pretraining, supervised fine-tuning, reinforcement learning, super-resolution, and culturally aware generation.
Pretraining datasets
-
Text-to-Image, T2I: 50 million images selected from the Kandinsky 5.0 collection and re-annotated with a new captioner.
- Filtering inherits the Kandinsky 5.0 pipeline: watermark and text-heavy scenes are removed, perceptual deduplication is applied, YOLOv8 and CLIP-based classifiers balance object and scene categories, TOPIQ and Q-Align scores enforce technical and aesthetic quality, and SAM 2 complexity filtering retains visually rich compositions.
- Captions are generated by Qwen3-235B-A22B in Russian and English, then rewritten by Qwen3-30B-A3B into four detail levels: long, medium, short, and shortest.
-
Text-to-Video, T2V: 20 million video scenes selected from the Kandinsky 5.0 T2V collection and re-captioned.
- Additional filtering removes clips with non-standard frame rates or scene transitions detected by TransNetV2.
- Properties include stable single-scene clips, standardized frame rates, no internal scene breaks, no watermarks or text-heavy scenes, perceptual deduplication, DOVER and Q-Align quality checks, YOLOv8 and CLIP-based category balancing, VideoMAE-based camera and object motion diversity, and MS-SSIM-based structural dynamics filtering.
- Captions use the same Qwen3-235B-A22B and Qwen3-30B-A3B pipeline as T2I, with cleaning of introductory phrases and non-English or non-Russian characters.
-
Text-to-Audio, T2A: 40 million audio tracks of 5 to 20 seconds extracted from Kandinsky 5.0 videos.
- Silent, corrupted, or low-quality tracks are discarded.
- Each track is annotated by Qwen3-30B-A3B with descriptions of sounds, music, speech, and environmental acoustics.
- Used to bootstrap the audio modality before joint audio-video training.
-
Text-to-Audio-Video, T2AV: 7 million audio-video segments derived from the Kandinsky 5.0 video collection.
- Filtering targets audio quality, meaningful audio-visual correspondence, and lip-sync suitability.
- Audio quality filters remove silent, corrupted, or low-quality audio and keep clips with sufficient dynamic range, low clipping, and adequate amplitude.
- Lip-sync and synchronization filters use Lip-Sync alignment score, LSA, and Desync from VABench. Clips with people require LSA of at least 3 out of 5. Clips with audio-visual events require Desync below 0.1 out of 1. Clips without visible faces do not require lip-sync and make up about 25% of the T2AV dataset.
- Visual quality inherits T2V filters: DOVER, Q-Align, CRAFT, YOLOv8, and CLIP-based classifiers.
- Captioning combines Gemma-4-E4B-it speech transcripts inside
<S>...<E>, Qwen3-30B-A3B audio descriptions inside<AUDCAP>...<ENDAUDCAP>, and Qwen3.5-35B-A3B for a unified visual and audio caption.
-
Image-to-Audio-Video, I2AV: 116,000 selected samples from the Kandinsky 5.0 video collection.
- Each sample pairs a starting frame, the first frame of a video scene, with a combined caption describing that frame and the expected audio-visual continuation.
- Captioning is a two-stage local pipeline: audio descriptions are generated by Gemma-4-E4B-it and Qwen3-Omni-30B-A3B, then combined with camera motion labels and passed to Qwen3.5-35B-A3B for a unified video and audio caption. Captions are rewritten into four detail levels.
- Starting frames are filtered with Qwen3-VL-Embedding-2B quality embeddings, PyIQA no-reference image quality metrics, blur detection via variance of Laplacian, and Qwen3-VL-8B-Instruct semantic checks for a clear non-mid-action starting frame.
Supervised fine-tuning datasets
- The authors construct two SFT datasets: one for video-only tasks and one for audio-visual tasks.
- Both use the same 11-domain taxonomy from Table 1 and follow a common pipeline:
- Technical filtering with DOVER, Q-Align, CRAFT, and motion dynamics models; audio is filtered with Meta Audiobox Aesthetics and silent or corrupted tracks are removed.
- Domain classification into 11 domains such as people action, animals, cartoons, people speech, music, food, nature, art, architecture, interiors, and tech, using Qwen2.5-VL-32B-Instruct visual embeddings and Qwen3-Omni-30B-A3B audio embeddings.
- Two-stage expert evaluation by technical screeners and cinematography experts.
- Captioning via Gemma-4-E4B-it, Qwen3-Omni-30B-A3B, and Qwen3.5-35B-A3B, with motion classification labels and four detail levels.
- Stage 1 contains 4,663 scenes and is used mainly for visual quality enhancement. Stage 2 contains 7,686 scenes and adds speech and music content to improve audio-video synchronization and lip-sync.
Reinforcement learning dataset
- The RL dataset contains 2,984 curated clips sourced from the SFT collection.
- Each sample has T2AV-specific captions in English and Russian.
- All samples include an audio caption with
<AUDCAP>...<ENDAUDCAP>markers. - Precomputed text embeddings are provided for efficient online training.
- Used for online RL with per-prompt rollouts and multiobjective rewards.
Super-resolution dataset
- The super-resolution dataset contains 150,000 unique clips from three sources:
- 100,000 diverse high-resolution 4K video clips.
- Around 35,000 high-resolution clips from the Kandinsky 5.0 T2V collection, selected through artifact filtering and Q-Align quality filtering.
- 15,000 synthetic camera-motion clips created from Aesthetic-4K high-resolution images by moving a spatial crop along a sampled camera trajectory.
- Quality filtering checks for compression artifacts, visible banding in smooth gradients, Q-Align scores, and a separate ripple detector for the 4K collection.
- Crop preparation extracts spatial crops of 512 by 768, 512 by 512, or 768 by 512 pixels. Crop locations use center cropping, random cropping, optical-flow-based selection, or selection of high-detail regions. Dark borders are excluded.
- Accepted crops become HQ targets, and degraded counterparts become LQ inputs. Both are encoded with K-VAE and stored as paired video latents.
- Two degradation pipelines are used: a Real-ESRGAN-style pipeline with blur, resizing, noise, JPEG compression, and video compression, and a CreativeVR-style temporal degradation pipeline with spatial grid warping, directional motion blur, adjacent-frame blending, frame dropping with interpolation, and spatio-temporal downsampling.
Russian Cultural Code dataset
- The Russian Cultural Code dataset contains 1.2 million images and 530,000 video scenes.
- It is curated to represent Russian cultural elements such as faith, language, historical memory, nature, architecture, and other culturally significant aspects.
- Samples are annotated with accurate object names and comprehensive scene descriptions.
- The dataset is used during T2V pretraining and is also injected through images during joint T2AV training, with English and Russian captions.
Usage summary
- Pretraining uses T2I, T2V, T2A, T2AV, and I2AV datasets to bootstrap image, video, audio, joint audio-video, and image-conditioned audio-video generation.
- Supervised fine-tuning uses the two SFT datasets in a two-stage procedure: first for visual quality, then for audio-video synchronization and lip-sync.
- The RL dataset is used for online reinforcement learning after SFT.
- The super-resolution dataset is used for constructing HQ/LQ training pairs for super-resolution.
- No explicit mixture ratios across these datasets are reported in the provided excerpt.
Method
The authors design Kandinsky 6.0 Video around a dual-stream CrossDiT architecture that processes video and audio modalities simultaneously. The model comprises asymmetric DiT backbones for video and audio, connected by blockwise bidirectional cross-attention.
Each stream inherits its architecture from the previous generation but differs in hidden and feed-forward dimensions. The video stream utilizes 3D Rotary Position Embedding (RoPE) indexed by frame number, while the audio stream uses 1D RoPE based on temporal position.
To achieve Full-HD output, the authors employ a separate super-resolution pipeline operating in the latent space of a video autoencoder. This pipeline consists of a Latent Upscaler (LU) and a Super-Resolution Diffusion Transformer (SR-DiT). The SR-DiT is a text-free model that takes a noisy latent, an HQ anchor tensor, and a binary mask to predict the flow-matching velocity.
The Latent Upscaler brings the low-quality latent to the working resolution of the SR-DiT directly in latent space, avoiding pixel-space decoding and re-encoding. It features separate branches for ×2 and ×4 upscaling, derived from the K-VAE decoder geometry.
The training process follows a sequential multi-stage approach.
Pre-training begins with separate training of the video and audio streams. The video stream is initialized from a previous checkpoint, while the audio stream is trained from scratch. They are then fused and trained jointly in text-to-audio-video and image-to-audio-video modes. Following pre-training, the model undergoes supervised fine-tuning across 11 domains to enhance visual quality and audio-video synchronization, culminating in a model soup via uniform weight averaging. Reinforcement learning post-training is then applied using an adapted OmniNFT method to optimize for modality-wise advantages and lip-sync accuracy. Finally, the model is distilled using π-Flow for trajectory compression and Sim-LADD for adversarial refinement to recover high-frequency details.
Experiment
The experiments validate Kandinsky 6.0 video models through regularization and reinforcement learning post-training studies, super-resolution and latent upscaler validation, consumer GPU optimization tests, benchmark comparisons, and large-scale human side-by-side evaluations. The regularization and RL stages improve speech intelligibility and audio-visual alignment over supervised fine-tuning, while distillation preserves model quality. Super-resolution is assessed with complementary no-reference quality and temporal diagnostics plus human review, and offloading-based deployment scales to consumer GPUs with only rounding-level differences, apart from NF4 text-encoder quantization. Human and benchmark results show mixed competitiveness: the model is strong on speech and synchronization, but several competitors lead on visual fidelity and general audio quality.
The SFT dataset is unevenly distributed across domains, with people action dominating both stages and nature starting smallest. Stage 2 synchronization data is larger than Stage 1 for every listed domain except music, with especially strong relative growth in nature, people speech, cartoons, and food. Music remains unchanged between stages. People action is the largest domain in both stages, while nature has the smallest Stage 1 count. Nature grows sharply in Stage 2, overtaking music and approaching food-level data volume. Music is the only listed domain whose sample count stays the same across Stage 1 and Stage 2. People speech, cartoons, food, and nature receive substantially more Stage 2 samples than Stage 1.
Kandinsky 6.0 Video uses a dual-stream CrossDiT with separate visual and audio backbones. The Pro configuration is consistently larger than Lite across hidden widths, feed-forward dimensions, visual and audio depth, and text depth. Modality-specific pretraining stages omit the unused stream, while joint and image-to-audio-video stages configure both visual and audio streams with the same per-modality dimensions and block counts. Pro uses wider video and audio model dimensions, wider feed-forward layers, and deeper visual and audio stacks than Lite. The text encoder block count remains fixed across all training stages for a given model size. Video-only training includes only visual stream parameters and audio-only training includes only audio stream parameters; joint stages include both streams. Once a visual or audio stream is present, its model dimension and number of blocks stay constant across video-only, audio-only, joint, and image-to-audio-video stages.
The continuous pre-training recipe for Kandinsky 6.0 Video Lite and Pro moves from separate text-to-video and text-to-audio stages to a longer joint text-to-audio-video stage, followed by a short mixed stage that includes image-to-audio-video samples. Hyperparameters vary by stage: audio-caption training uses a higher learning rate and a different optimizer beta schedule, while the final mixed stage uses a lower learning rate. Lite and Pro differ mainly in step counts across stages, with shared regularization settings such as weight decay and EMA decay. Text-to-audio pre-training on audio captions uses a higher learning rate than the other stages and a different optimizer beta schedule compared with text-to-video and joint stages. The joint text-to-audio-video stage is the longest stage, while the final image-to-audio-video mixed stage is much shorter, uses a lower learning rate, and nearly eliminates warmup. Lite uses more joint training steps but fewer audio-video-caption and final mixed steps than Pro, while weight decay and EMA decay stay constant across stages.
Validation metrics show stage-dependent trade-offs between spatial quality and temporal consistency for the SR-DiT. At ×4, the later stages improve no-reference image and video quality scores, and distillation reduces Laplacian artifacts, but warp error increases across stages. At ×2, the stage changes are smaller, with distillation again lowering artifacts while yielding the highest warp error. At ×4, NABLA fine-tuning improves MUSIQ, DOVER, and CLIP-IQA, but increases Laplacian artifacts and warp error; subsequent π-Flow distillation reduces Laplacian artifacts below the Stage 1 level while warp error continues to rise. At ×2, spatial quality metrics remain close across stages, while the distilled stage achieves the lowest Laplacian artifacts and the highest DOVER but also the highest warp error.
The latent upscaler is compared against bilinear pixel-space upsampling using PSNR averaged per clip. In aggregate, the upscaler improves PSNR slightly at ×2 and more clearly at ×4, driven mainly by synthetic 4K and camera-motion clips. For weak-source clips, ×2 results can favor bilinear, but ×4 results are neutral or positive. Aggregated across all validation clips, the upscaler improves over bilinear at both scales, with a much larger benefit at ×4 than at ×2. The upscaler helps most on synthetic 4K and camera-motion content, while weak-source clips at ×2 can favor bilinear.
The experiments validate the Kandinsky 6.0 Video pipeline across data, architecture, training, and super-resolution components. Stage 2 synchronization data expands most domains while music remains unchanged and people action stays dominant, and the staged training recipe moves from separate video and audio pretraining to a longer joint audio-video stage followed by a short low-learning-rate mixed stage. Model configuration comparisons confirm Pro is consistently wider and deeper than Lite with stable per-modality dimensions, while super-resolution evaluations show a quality tradeoff where later stages improve spatial and no-reference scores but increase warp error, and the latent upscaler benefits most at higher scale factors and on synthetic 4K and camera-motion content.