HyperAIHyperAI

Command Palette

Search for a command to run...

SenseNova-U1.5: Auf dem Weg zu nativer, vereinheitlichter visueller Intelligenz

Zusammenfassung

Wir stellen SenseNova-U1.5 vor, ein natives, vereinheitlichtes multimodales 8B-MoT-Modell, das visuelle Inhalte innerhalb einer Encoderund VAE-freien Architektur versteht, darüber nachdenkt und generiert. Wir stärken seine visuelle Schnittstelle durch räumlich kohärente Patch-Rekonstruktion und skalieren das Training mit sorgfältig kuratierten Generierungsund Bearbeitungsdaten, verbesserter Aufgabenformulierung, struktureller Prompt-Anreicherung und nativen Auflösungen von bis zu 4K. Für das Post-Training optimieren wir spezialisierte Experten für visuelle Ästhetik, zweisprachige Textdarstellung, Infografik-Generierung und Bildbearbeitung und konsolidieren deren Fähigkeiten durch Multi-Experten-On-Policy-Destillation. In umfangreichen Evaluierungen verbessert SenseNova-U1.5 die Bildtreue, Textdarstellung, komplexe Komposition, Bearbeitung mit mehreren Referenzen und verschränkte Generierung erheblich, während es gleichzeitig die Befolgung von Anweisungen verbessert und Subjektidentität, Geometrie und unveränderte Bereiche bewahrt. Trotz begrenzter Exposition gegenüber strukturierten Formaten in seinen Generierungsdaten generalisiert SenseNova-U1.5 effektiv auf lange, komplexe und strukturierte visuelle Anweisungen, was weiter belegt, dass multimodales Verstehen auf visuelle Planung und Kreation übertragbar ist. Zusammengenommen positionieren diese Erkenntnisse die native, vereinheitlichte Modellierung als vielversprechenden Weg zu Systemen, die in einem vollständig Ende-zu-Ende-Framework wahrnehmen, schlussfolgern und kreieren. Wir werden den Trainingscode, einschließlich überwachtem Feintuning, bestärkendem Lernen und On-Policy-Destillation, als Open Source veröffentlichen.

One-sentence Summary

SenseTime et al. introduce SenseNova-U1.5, an 8B-MoT native unified multimodal model that unifies visual understanding, reasoning, and generation within an encoder-free, VAE-free architecture, enhanced by spatially coherent patch reconstruction and multi-expert on-policy distillation, achieving significant advances in image fidelity, text rendering, and complex editing while generalizing to long structured visual instructions.

Key Contributions

  • The paper introduces SenseNova-U1.5, an 8B-MoT native unified multimodal model that performs visual understanding, reasoning, and generation within a single encoder-free and VAE-free architecture, and strengthens its visual interface through spatially coherent patch reconstruction.
  • It develops a training pipeline that scales with curated generation and editing data, improved task formulation, structural prompt enhancement, and native resolutions up to 4K, followed by post-training that optimizes specialized experts for aesthetics, bilingual text rendering, infographic generation, and image editing, and consolidates them via multi-expert on-policy distillation.
  • Extensive evaluations demonstrate that SenseNova-U1.5 substantially improves image fidelity, text rendering, complex composition, multi-reference editing, and interleaved generation, while also generalizing to long and structured visual instructions, providing evidence that multimodal understanding can transfer to visual planning and creation.

Introduction

The authors address the growing scope of visual creation, which now spans multilingual typography, high-resolution rendering, multi-reference editing, and interleaved generation. Most systems handle understanding via pretrained vision encoders and generation via VAEs, creating separate representational spaces that hinder seamless coordination. Prior native unified models like SenseNova-U1 removed these separate pipelines but reconstructed each visual token as an independent RGB patch, leading to visible seams and texture discontinuities at high resolutions. The authors present SenseNova-U1.5, an 8B-MoT model that introduces spatially joint reconstruction: tokens are projected onto a 2D feature field and progressively upsampled through convolutions and Pixel Shuffle, enabling information exchange across neighboring regions for coherent 4K generation. A second advance replaces joint reward optimization with a specialize-then-unify strategy, where dedicated RL experts are optimized for distinct visual capabilities and then consolidated into a single policy via multi-expert on-policy distillation, preserving complementary strengths without entangling competing objectives.

Dataset

The authors construct a large-scale multimodal training corpus for SenseNova-U1.5, expanding the earlier SenseNova-U1 with a focus on data diversity, visual quality, high-resolution generation, complex instruction following, and fine-grained image editing. The data is curated through quality-aware filtering, distribution rebalancing, stronger image-text alignment, and targeted synthesis for underrepresented capabilities. The corpus is organized into four main subsets, each with distinct composition, processing, and usage.

  • Image Generation Data

    • Composition: approximately 59 million additional text–image pairs collected from 78 sources, covering general scenes, human-centric content, objects, text-rich imagery, infographics, and specialized domains. High-resolution samples dominate: 88.2% of effective training volume exceeds 1024² resolution and 64.4% exceeds 2048².
    • Processing: captions are provided at multiple granularities (long descriptions, concise captions, semantic tags) with strengthened bilingual Chinese–English coverage, especially for text-rich images. Synthetic data is generated for rare concepts, complex layouts, and text-intensive designs. All samples pass a shared filtering pipeline that evaluates perceptual quality, prompt fidelity, text correctness, layout quality, and visual diversity.
    • Usage: this subset exposes the model to rich multi-scale visual statistics, supervising both global scene composition and fine-grained local details.
  • Image Editing Data

    • Composition: approximately 38 million examples covering four settings—general editing (~43%), spatially controlled infographic editing (~42%), reference-conditioned editing (~15%), and spatially controlled editing with bounding boxes and visual markers. Reference-conditioned examples include up to ten reference images.
    • Processing: editing instructions are diversified in formulation, granularity, and complexity. For multi-target or multi-constraint edits, structured prompt-enhancement and chain-of-thought examples explicitly describe editing intent, target regions, desired attributes, and content to preserve. Curation jointly evaluates instruction fulfillment, target quality, reference/identity fidelity, and preservation of unedited regions.
    • Usage: the data trains the model on a broad range of editing scenarios, from localized attribute changes to scene-level transformations and reference-guided customization.
  • Interleaved Data

    • Composition: multimodal trajectories drawn from lifestyle-oriented sequences (~44%), infographic data (~29%), video-derived sequences (~19%), and reasoning-intensive examples (~8%).
    • Processing: all sources are converted into a unified trajectory format where text and visual states are progressively interleaved. A shared pipeline handles source preprocessing, domain-specific transformation or synthesis, and trajectory-level validation. Filtering ensures linguistic validity, visual quality, cross-modal consistency, and overall coherence.
    • Usage: the interleaved trajectories train the model on long-range semantic coherence, visual consistency, temporal evolution, and reasoning-conditioned generation across multiple steps.
  • RL Training Data

    • Aesthetic, OCR, and Infographic Samples:
      • Aesthetic: ~280K prompts (from HPSv3++, Pick-a-Pic, bilingual expansions, and internal sources); multiple images per prompt are scored with HPSv3++, and groups with weak reward variation are removed.
      • OCR: ~60K bilingual text-rendering prompts (20K short-text prompts from Flow-GRPO expanded to 40K longer, denser prompts); specified text is extracted as reference transcription for online OCR reward.
      • Infographic: the first GRPO stage reuses the OCR corpus; later stages jointly iterate OCR and aesthetic prompts. Additionally, ~120K filtered preference pairs from the Linear-DPO corpus are used for DPO, balanced over portrait/non-portrait content and Chinese/English prompts.
    • Image Editing RL Samples:
      • ~120K examples, with local editing (~80%) and global editing (~20%). Instructions cover Chinese, English, and mixed languages, with balanced lengths and resolutions from 512×512 to 2048×2048.
      • Curation: corrupted, duplicated, low-resolution, or degraded samples are removed; entity, text, and attribute grounding is verified; both source and edited images are evaluated for aesthetics and technical quality; retained samples are rebalanced across task categories and resolutions.
    • Usage: these RL-specific subsets are used in GRPO and DPO stages to align the model with human preferences and task-specific rewards, improving visual quality, text rendering, and editing precision.

Overall, the dataset is used to train SenseNova-U1.5 across image generation, image editing, and interleaved multimodal tasks, with the RL subsets providing targeted preference optimization.

Method

The authors design SenseNova-U1.5 with a native unified multimodal paradigm, integrating understanding and generation within a single framework. The architecture features a near-lossless visual interface that directly transforms raw images or noise-corrupted visual inputs into compact token sequences without relying on an external visual encoder or VAE. Two convolutional projections with GELU activations downsample the input, yielding one visual token for each 32×\times×32 image region. Two-dimensional sinusoidal positional embeddings preserve spatial coordinates, and special tokens delimit individual visual blocks. Text and visual representations are projected into a shared hidden space and jointly processed by the unified backbone.

As shown in the figure below:

For generation, the model explicitly incorporates a resolution-dependent noise-scale embedding to condition the denoising process. The noise scale σR(H,W)\sigma_R(H, W)σR(H,W) is normalized and encoded using a dedicated sinusoidal MLP, then combined with the diffusion timestep representation to provide explicit awareness of both diffusion time and resolution-dependent noise statistics. To address patch-wise factorization artifacts at high resolutions, the authors replace the original MLP head with a lightweight spatially coupled decoder. Given backbone hidden states, the model restores their two-dimensional topology and progressively reconstructs the full-resolution RGB image through Pixel Shuffle stages with upsampling factors of 2, 2, and 8. Between successive upsampling stages, 3×\times×3 convolutions enable information exchange across neighboring token regions, restoring local spatial interactions before pixel synthesis.

The native Mixture-of-Transformers design integrates understanding and generation within a single Transformer backbone. Clean image-text context and noise-conditioned visual states are interleaved within a unified sequence, allowing semantic representations and generative dynamics to interact directly through shared self-attention. The attention pattern is structured to reconcile causal language modeling with bidirectional visual interaction. Text tokens attend only to preceding context, while tokens within each clean image block attend bidirectionally. Crucially, architectural unification does not imply parameter sharing everywhere; understanding and generation retain separate attention projections, normalization layers, and feedforward modules, dynamically routed by token type at each layer.

The model is trained with a unified objective that jointly couples autoregressive language modeling, native pixel-space flow learning, and perceptual supervision. For multimodal understanding, the authors optimize the conditional likelihood of the target text sequence:

LAR=1Nn=1Nlogpθ(xnx<n,c)\mathcal{L}_{\mathrm{AR}} = - \frac{1}{N} \sum_{n=1}^{N} \log p_{\theta}(x_n \mid x_{<n}, \mathbf{c})LAR=N1n=1Nlogpθ(xnx<n,c)

For visual generation, the model learns pixel-space flow matching directly in RGB space with resolution-adaptive noise:

zt=tx+(1t)σR(H,W)ϵ\mathbf{z}_t = t \mathbf{x} + (1 - t) \sigma_R(H, W) \boldsymbol{\epsilon}zt=tx+(1t)σR(H,W)ϵ

The model predicts the clean endpoint to induce the velocity estimate, supervised by the pixel-space flow-matching loss:

LFlow=E[vθv22]\mathcal{L}_{\mathrm{Flow}} = \mathbb{E} \left[ \| \mathbf{v}_{\theta} - \mathbf{v}^{\star} \|_2^2 \right]LFlow=E[vθv22]

A perceptual objective using LPIPS is further introduced to improve structural consistency:

LPerc=LPIPS(x^θ,x)\mathcal{L}_{\mathrm{Perc}} = \mathrm{LPIPS}(\hat{\mathbf{x}}_{\theta}, \mathbf{x})LPerc=LPIPS(x^θ,x)

The entire training process is jointly optimized under the unified objective:

L=λARLAR+λFlowLFlow+λPercLPerc\mathcal{L} = \lambda_{\mathrm{AR}} \mathcal{L}_{\mathrm{AR}} + \lambda_{\mathrm{Flow}} \mathcal{L}_{\mathrm{Flow}} + \lambda_{\mathrm{Perc}} \mathcal{L}_{\mathrm{Perc}}L=λARLAR+λFlowLFlow+λPercLPerc

The authors progressively build native multimodal capabilities through a multi-stage training pipeline, followed by capability-specific learning and distillation.

Refer to the framework diagram:

The initial stages involve Generation Pre-Training, Unified Mid-Training, and Unified Supervised Fine-Tuning. In Stage 1, the generation branch is randomly initialized and trained via pixel-space flow matching, conditioned on representations from the frozen understanding branch. This phase introduces a dedicated native 4K training phase, extending to higher-resolution text-to-image data and incorporating image-editing and interleaved-generation tasks. In Stage 2, the authors jointly optimize both branches using a mixed corpus of text-only, multimodal-understanding, text-to-image, and image-editing data, balancing general multimodal competence with diverse generative capabilities. Stage 3 further fine-tunes the model on high-quality instruction-following data to strengthen instruction adherence.

Following these unified stages, the authors implement Multi-Expert Reinforcement Learning in Stage 4. They train four specialized experts for aesthetics, text rendering, infographic generation, and image editing, each with task-specific data, rewards, and optimization strategies. The aesthetic expert interleaves aesthetic-preference and typography data, utilizing HPSv3++ and a bilingual OCR reward. The OCR expert specializes in accurate bilingual text rendering using a normalized multiset intersection-over-union score. The editing expert balances instruction adherence, content preservation, and visual quality using a multi-dimensional VLM-based reward framework and a progressive sliding-window strategy. The infographic expert undergoes a complete training process from infographic-oriented mid-training to task-specific RL, refining text rendering and overall visual quality through direct preference optimization and group relative policy optimization.

In Stage 5, the authors apply Multi-Expert On-Policy Distillation to consolidate the four specialized experts into a single unified model. Each training sample is hard-routed to the frozen expert corresponding to its capability. The student generates its own trajectory, and both the student and the routed expert are evaluated at the same student-generated state, timestep, and condition using on-policy velocity-field distillation:

LOPD=E[vθ(sg(x^θ,t),t,c)vm(sg(x^θ,t),t,c)22]\mathcal{L}_{\mathrm{OPD}} = \mathbb{E} \Big[ \| \mathbf{v}_{\theta}(\mathrm{sg}(\hat{\mathbf{x}}_{\theta, t}), t, \mathbf{c}) - \mathbf{v}_{m}(\mathrm{sg}(\hat{\mathbf{x}}_{\theta, t}), t, \mathbf{c}) \|_2^2 \Big]LOPD=E[vθ(sg(x^θ,t),t,c)vm(sg(x^θ,t),t,c)22]

To learn both high-noise states that establish global structure and low-noise states that determine local details, the query distribution progressively shifts over training from emphasizing early high-noise portions to late low-noise states. The student is trained using fixed 30-step deterministic ODE denoising steps, directly optimizing the velocity under classifier-free guidance to align training and inference.

Experiment

SenseNova-U1.5 is evaluated across multimodal understanding, image generation, editing, and interleaved generation benchmarks. The model retains strong visual and language understanding while adding robust generation and editing capabilities, with notable improvements in compositional consistency, bilingual text rendering, and reasoning-driven editing. Interleaved generation experiments further show that generation can serve as an intermediate representation to support multimodal reasoning, validating the effectiveness of native multimodal unification.

SenseNova-U1.5 demonstrates leading open-source and often state-of-the-art performance across multimodal editing, reasoning, and interleaved generation tasks. It particularly excels when chain-of-thought reasoning is applied, and its generation capabilities can serve as an intermediate representation to enhance multimodal understanding. On OmniRef-Bench, the model achieves leading open-source results with pronounced gains in style consistency and background preservation, while reliably maintaining subject and pose consistency. Chain-of-thought reasoning yields substantial improvements on RISEBench for causal, logical, and temporal edits, but spatial edits benefit less consistently. With chain-of-thought, SenseNova-U1.5 attains top overall performance on OpenING, surpassing proprietary pipelines in image-text coherency, human alignment, and multi-step consistency. On VBVR-Pro-Bench, it sets a new state of the art across cognitive faculties and generalizes strongly to out-of-domain tasks, outperforming leading proprietary models. The model uses generation as an intermediate representation, achieving leading performance on RealUnify-GEU while remaining competitive on Uni-MMMU-GaU.

The training recipe progresses through generation pre-training, unified mid-training, and supervised fine-tuning, with learning rates decreasing from 2e-4 to 2e-5 and sequence lengths growing from 8K to 32K tokens. Early stages keep the understanding branch frozen while training the generation branch, then both branches are jointly optimized with a mixed corpus that emphasizes text-to-image data. In the joint stages, generation loss is weighted ten times higher than understanding loss to strengthen synthesis while retaining comprehension. Stage 1 generation pre-training uses three phases that lower the learning rate and raise the sequence length, moving from a constant 2e-4 at 8K tokens to a cosine decay ending at 2e-5 at 20K tokens. Unified mid-training mixes 40% text-to-image, 30% text and understanding, 20% image-editing, and 10% interleaved data, and applies a constant learning rate of 2e-5 with a 32K-token context. In both unified mid-training and SFT, the generation loss coefficient is 1.0 while the understanding loss coefficient is 0.1, prioritizing generative capability without discarding understanding.

SenseNova-U1.5 maintains or improves upon SenseNova-U1 across multimodal benchmarks while adding generation and editing capabilities, and it outperforms the encoder-free Gemma4-12B on key STEM and VQA tasks. Unified training does not degrade language understanding, with U1.5 achieving strong results on MMLU-Pro, C-Eval, and instruction-following evaluations. SenseNova-U1.5 matches or exceeds SenseNova-U1 on MMMU, MathVista, MMBench, and other multimodal benchmarks, while substantially extending its generative abilities. Language understanding remains robust after unified training: U1.5 scores 86.67 on MMLU-Pro, 90.41 on C-Eval, and 93.35 on IFEval, improving instruction following over U1.

SenseNova-U1.5 with prompt enhancement achieves an overall score of 60.22 on the English subset of Qwen-Image-Bench, the highest among open-source models. It surpasses several closed-source systems like Nano-Banana-Pro and Seedream 5.0, and approaches GPT-Image-1.5, demonstrating strong bilingual generation quality. With prompt enhancement, SenseNova-U1.5 reaches the best overall performance among open-source models, outperforming closed-source systems such as Nano-Banana-Pro and Seedream 5.0. Even without prompt enhancement, SenseNova-U1.5 consistently improves over its predecessor SenseNova-U1 and remains competitive with larger open-source baselines.

On the Chinese subset of Qwen-Image-Bench, SenseNova-U1.5 with prompt enhancement achieves an overall score of 60.13, surpassing several closed-source models such as Nano-Banana-2.0 and GPT-Image-1.5, and trailing only GPT-Image-2. Without prompt enhancement, the model still improves over its predecessor and remains competitive with larger open-source baselines, narrowing the gap with leading closed-source systems. SenseNova-U1.5 with prompt enhancement reaches 60.13 overall on the Chinese subset, outperforming closed-source models like Nano-Banana-2.0 (59.82) and GPT-Image-1.5 (59.65). Even without prompt enhancement, SenseNova-U1.5 consistently improves over SenseNova-U1 and narrows the performance gap with top closed-source systems.

SenseNova-U1.5 is trained through a progressive recipe of generation pre-training, unified mid-training, and supervised fine-tuning, with generation loss heavily weighted to preserve synthesis while maintaining understanding. Evaluations across multimodal editing, reasoning, and generation benchmarks show that the model achieves leading open-source and often state-of-the-art performance, particularly when chain-of-thought reasoning is applied, and its generative capabilities can serve as an intermediate representation to enhance multimodal understanding. The model also matches or surpasses its predecessor on understanding benchmarks without degradation, and with prompt enhancement it attains the best open-source results on bilingual image generation quality tests, outperforming several closed-source systems.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp