HyperAIHyperAI

Command Palette

Search for a command to run...

FuseReg: 層融合の正則化が表現オートエンコーダにおける再構成–生成ギャップを緩和する

概要

表現オートエンコーダ(RAE)は、事前学習済み視覚エンコーダから得た特徴を再構成および拡散の潜在表現として再利用し、強力な視覚表現を画像生成に統合する。しかしRAEでは、生成器とピクセルデコーダが共有する潜在空間をどのエンコーダ層から構成するかを決める必要がある。この選択にはトレードオフがある。浅い層は微細なピクセル詳細を保持しやすい一方、深い層は生成指標が優れる傾向がある。したがって、固定的なヒューリスティック層融合は、異なる情報を必要とする2つの段階を結合してしまう。本研究では、ヒューリスティックな特徴選択をエンコーダ層のランダムな部分集合で学習するFuseRegを導入する。そのメカニズムを理論的に解析し、部分集合サンプリングが層間不一致への感度に明示的なペナルティを課すことを示す。ImageNet-256とDINOv3-Lを用いた実験では、単一のFuseRegデコーダが再学習なしで全層・スパース・単層融合から再構成し、固定融合に特化したデコーダよりも高いPSNRを達成した。この柔軟性は生成にも寄与し、RAEv2 DiT-XL生成器を変更せずにデコーダを置き換えるだけで非ガイドgFIDを27%削減した。同じ正則化原理は拡散学習にも拡張でき、両段階の同時正則化によりDiT-Baseで非ガイドgFIDが29%削減された。これらの結果は、層融合への頑健性のために下流モデルを学習させることが、事前学習済みエンコーダを変更せずに再構成–生成ギャップを狭めることを示している。

One-sentence Summary

Researchers at USC PSI Lab and collaborators propose FuseReg, which regularizes representation autoencoders by training on random subsets of encoder layers to penalize cross-layer disagreement, enabling a single decoder to reconstruct from full, sparse, and single-layer fusions without retraining while improving PSNR and reducing unguided gFID by up to 29%.

Key Contributions

  • FuseReg is a layer-fusion regularizer that trains representation autoencoder decoders and generators on normalized fusions of randomly sampled encoder layers, replacing fixed heuristic layer selection.
  • A theoretical analysis shows that randomized subset sampling preserves the full-layer mean while making cross-layer disagreement an explicit regularization penalty and establishes a second-order separation from deterministic global fusion methods.
  • On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining and achieves higher PSNR than decoders specialized to fixed fusions; decoder replacement alone reduces unguided gFID from 3.01 to 2.21 at k=23, while joint regularization reduces DiT-Base gFID from 13.96 to 9.93.

Introduction

The authors work with representation autoencoders (RAEs), where a frozen visual encoder defines the latent space of a diffusion model but its hierarchical features must be collapsed into a single latent through layer fusion. This interface matters because the decoder reconstructs pixels from generated latents while the generator models the latent distribution, and prior fixed fusion choices such as RAEv2 create a reconstruction-generation gap: reconstruction favors shallow pixel-aligned layers, whereas diffusion favors deeper semantic layers. For example, in the DINOv3-L setting, expanding the fused subset from k=7 to k=23 improves reconstruction PSNR from 22.58 to 27.04 dB but worsens unguided gFID from 1.65 to 3.01. The authors introduce FuseReg, a layer-fusion regularization approach that trains the decoder and generator on randomly sampled, normalized subsets of encoder layers. This preserves the full-layer representation in expectation while reducing reliance on any single layer composition, yielding a single decoder that supports full, sparse, and single-layer fusions without retraining, and improving generation quality without architectural changes or additional computational cost.

Method

The authors address the limitations of fixed layer fusion in Representation Autoencoders by introducing FuseReg, a method that leverages randomized layer fusion as a regularization technique. In traditional setups, a fixed fusion rule or a learned global gate exposes downstream models to only a single layer composition during training. However, learned gates often fail to utilize deep semantic features effectively, instead collapsing to rely almost exclusively on shallow, pixel-aligned layers.

As shown in the figure below:

To overcome this dependency on fixed shortcuts, the authors replace deterministic fusion weights with a mean-preserving distribution over layer subsets. Given an image and a frozen encoder, the layer-normalized tokens from KKK chosen layers are mapped to a single latent representation. While a standard normalized fusion uses fixed weights aaa such that za=∑k=1Kakhkz_a = \sum_{k=1}^K a_k h_kza​=∑k=1K​ak​hk​ with 1⊤a=1\mathbf{1}^\top a = 11⊤a=1, FuseReg samples a separate layer-dropout mask for each training example. For a drop rate p∈[0,1)p \in [0, 1)p∈[0,1), the method draws an independent and identically distributed Bernoulli mask mmm and conditions it on being nonzero. The resulting fused latent is computed as:

zm=∑kmkhk∑kmkz_m = \frac{\sum_k m_k h_k}{\sum_k m_k}zm​=∑k​mk​∑k​mk​hk​​

This normalization is critical because it ensures that the expected value of the sampled latent equals the full-layer mean. Consequently, the variation introduced by the sampling strictly targets directions where the encoder layers disagree without shifting the average representation.

The framework applies this randomized fusion to both the pixel decoder and the diffusion transformer, but employs separate drop rates for each stage since they solve fundamentally different prediction problems. For the decoder, subset fusions are used to reconstruct pixels. The decoder is trained to map the sampled latent zmz_mzm​ back to the image using a combination of pixel, perceptual, and adversarial objectives. This forces the decoder to reconstruct images from multiple layer compositions, reducing its reliance on shallow-layer shortcuts.

For the diffusion transformer, the generator receives noisy subset fusions but is trained to predict the full-layer aggregate target. Using a flow-matching interpolant built from the subset aggregate, the model regresses the full-layer target with a time-weighted prediction loss. Setting the drop rate to zero recovers the standard full-fusion objective, while any positive rate trains the generator to map multiple subset fusions to the same full-layer target.

Theoretical analysis confirms that this normalized subset sampling preserves the all-layer representation on average while adding variance precisely along the components where layers disagree. Furthermore, the second moment of these random fusion weights spans the entire layer-contrast subspace. Because deterministic input-independent fusion rules possess a rank-one second moment, no fixed global fusion can match both the mean and the variance properties of the randomized approach, making FuseReg a uniquely effective regularizer for bridging the reconstruction and generation gap.

Experiment

The evaluation uses a frozen DINOv3-L encoder and a ViT decoder trained on ImageNet-256 with matched DiT-Base and DiT-XL budgets, and it isolates decoder, generator, and joint effects by testing one decoder across different layer fusions, swapping decoders while fixing the generator and latents, and jointly regularizing both stages. The experiments show that a randomized-fusion decoder reconstructs robustly across full, subset, and single-layer fusions, whereas fixed-fusion decoders specialize and degrade on shifted fusions; improving decoder robustness also improves generation with the generator fixed, especially for shifted fusions. Joint regularization yields complementary gains on DiT-Base but more nuanced, metric-dependent behavior on DiT-XL, motivating separate selection of decoder and generator regularization rates depending on scale, evaluation metric, and guidance.

The experiment compares reconstruction quality across full, subset, and single-layer fusions. Fixed-fusion RAEv2 decoders perform well only on the fusion they were trained on and degrade sharply when transferred to other fusions. A single FuseReg decoder, trained across nontrivial fusions, retains consistently strong PSNR, SSIM, and rFID across all evaluated fusion types. Fixed-fusion RAEv2 decoders specialize heavily to their training fusion, with large drops in PSNR and rFID on off-training fusions. The FuseReg decoder matches or exceeds the deterministic decoders across all tested fusions, keeping rFID below 0.6 and SSIM above 0.67 in every case.

In this unguided ImageNet-256 Inception Score sweep, decoder regularization has a non-monotonic effect: small decoder rates improve over the fixed-fusion baseline, but larger decoder rates reduce scores. Generator regularization generally lowers Inception Score, so the strongest results come from no or little generator regularization paired with a modest decoder rate. The best Inception Score appears at zero generator rate with a small decoder rate, improving substantially over the reproduced fixed-fusion baseline. For each generator rate, scores tend to peak at decoder rates of 0.05 or 0.1 and then decline as decoder regularization increases. High generator regularization combined with high decoder regularization produces the lowest scores in the sweep.

The first experiment evaluates reconstruction quality across full, subset, and single-layer fusions, showing that fixed-fusion RAEv2 decoders overfit to their training fusion and degrade sharply when transferred, while a single FuseReg decoder trained across nontrivial fusions maintains consistently strong PSNR, SSIM, and rFID. The second experiment sweeps decoder and generator regularization on unguided ImageNet-256 using Inception Score, finding that mild decoder regularization improves over the fixed-fusion baseline but larger amounts hurt, and generator regularization generally reduces scores. The best setup uses no or little generator regularization with a modest decoder rate, whereas high regularization on both sides yields the lowest scores.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています