HyperAIHyperAI

Command Palette

Search for a command to run...

Tokenize Once, Recommend Anywhere: Unified Item Tokenization for Multi-domain LLM-based Recommendation

Yu Hou Won-Yong Shin

Abstract

Large language model (LLM)-based recommender systems have achieved high-quality performance by bridging the discrepancy between the item space and the language space through item tokenization. However, existing item tokenization methods typically require training separate models for each item domain, limiting generalization. Moreover, the diverse distributions and semantics across item domains make it difficult to construct a unified tokenization that preserves domain-specific information. To address these challenges, we propose UniTok, a Unified item Tokenization framework that integrates our own mixture-of-experts (MoE) architecture with a series of codebooks to convert items into discrete tokens, enabling scalable tokenization while preserving semantic information across multiple item domains. Specifically, items from different domains are first projected into a unified latent space through a shared encoder. They are then routed to domain-specific experts to capture the unique semantics, while a shared expert, which is always active, encodes common knowledge transferable across domains. Additionally, to mitigate semantic imbalance across domains, we present a mutual information calibration mechanism, which guides the model towards retaining similar levels of semantic information for each domain. Comprehensive experiments on wide-ranging real-world datasets demonstrate that the proposed UniTok framework is (a) highly effective: achieving up to 51.89% improvements over strong benchmarks, (b) theoretically sound: showing the analytical validity of our architectural design and optimization; and (c) highly generalizable: demonstrating robust performance across diverse domains without requiring per-domain retraining, a capability not supported by existing baselines.

One-sentence Summary

Researchers at Yonsei University propose UniTok, a unified item tokenization framework that employs a mixture-of-experts architecture with a series of codebooks to convert multi-domain items into discrete tokens via a shared encoder and domain-specific experts, supplemented by a mutual information calibration mechanism to balance semantic information across domains, achieving up to 51.89%51.89\%51.89% improvements over strong benchmarks without requiring per-domain retraining.

Key Contributions

  • Introduces UniTok, a unified item tokenization framework that combines a customized mixture-of-experts architecture with codebook-based token extraction to map items from multiple domains into discrete semantic tokens. The design routes inputs through domain-specific experts and an always-active shared expert, while a mutual information calibration mechanism reduces semantic imbalance across domains.

  • Experiments on diverse real-world datasets demonstrate that UniTok achieves up to 51.89% gains over strong baselines in NDCG@10, reduces trainable parameters by 9.63×, and maintains robust zero-shot performance without per-domain retraining.

  • Theoretical analysis shows that UniTok induces a higher entropy token space, achieves lower quantization error, and reduces mutual information variance across domains, ensuring semantic consistency and stable performance.

Introduction

Large language models (LLMs) are increasingly used for generative recommendation, but they require item tokenization to convert items into discrete identifiers that bridge the item space and language space. Existing tokenization methods are tailored to single-domain settings, requiring separate tokenizers for each item domain. As recommendation tasks span multiple domains, this siloed approach leads to inefficiencies in training, deployment, and maintenance, hindering scalability. Two key challenges arise: (C1) training overhead from repeatedly learning domain-specific tokenizers, and (C2) semantic alignment, where a shared token space across domains risks semantic mixing and biased token assignments.

The authors introduce UniTok, the first unified item tokenization framework for multi-domain LLM-based recommendation. To address C1 and partially C2, they propose Token-MoE, a mixture-of-experts architecture that separates domain-specific experts from a shared expert, enabling the model to retain domain specialization while sharing common knowledge across domains. To fully solve C2, they add a mutual information calibration mechanism that encourages latent embeddings from each domain to remain semantically informative and consistent, reducing inter-domain performance variability. In experiments, UniTok achieves up to 51.89% improvement in NDCG@10, reduces trainable parameters by 9.63x compared to codebook-based baselines, and shows strong zero-shot generalization without retraining. The authors also provide theoretical guarantees for higher entropy token space, lower quantization error, and semantic consistency via reduced MI variance.

Method

The authors propose UniTok, a unified framework designed to address the challenges of multi-domain item tokenization for large language model-based recommender systems. As illustrated in the framework diagram, the architecture comprises four core components: a shared autoencoder, a Mixture of Experts (MoE) module named TokenMoE, codebook-based identifiers, and a Mutual Information (MI) calibration mechanism. This design enables the model to project heterogeneous item domains into a consistent latent space while preserving domain-specific nuances.

The pipeline begins with a shared autoencoder that maps continuous semantic embeddings from diverse item domains into a unified latent space. Given an input item embedding xik\mathbf{x}_i^kxik​ from domain kkk, the encoder fθf_\thetafθ​ produces a latent representation zik\mathbf{z}_i^kzik​. The decoder gϕg_\phigϕ​ subsequently reconstructs the original embedding from the processed latent representation. This component establishes a common structural foundation across domains, optimized via a reconstruction loss that minimizes the L2L_2L2​ distance between the original and reconstructed embeddings.

To capture domain-specific semantics without conflating them, the framework employs TokenMoE. This module routes each latent embedding zik\mathbf{z}_i^kzik​ through a learnable router that computes a softmax distribution over KKK domain-specific experts. The router selects the top-NNN experts based on the highest activation scores. Concurrently, a shared expert is deterministically activated to facilitate cross-domain knowledge transfer. The final processed latent representation z^ik\hat{\mathbf{z}}_i^kz^ik​ is computed as a weighted combination of the outputs from the selected domain-specific experts and the shared expert. Each expert module is initialized with the mean feature of its corresponding domain to provide an inductive bias for specialization.

Within each expert, the model utilizes codebook-based identifiers to discretize the continuous latent embeddings into compact token sequences. This is achieved through Residual Quantization (RQ), which recursively applies a series of codebooks to quantize the input residual vectors. At each quantization stage, the nearest code vector is selected from the corresponding codebook, and the residual is updated. This hierarchical quantization process generates a discrete codeword sequence that effectively represents the item. The quantization module is trained using an RQ loss that aligns the code vectors with the target residuals while enforcing commitment from the encoder and router.

A critical challenge in multi-domain settings is semantic imbalance, where the quality of latent embeddings varies significantly across domains. To mitigate this, the authors introduce an MI calibration mechanism. This component measures the statistical dependence between the original semantic embeddings and the learned latent embeddings using the Hilbert-Schmidt Independence Criterion (HSIC) as a proxy for mutual information. By mapping both the input and latent spaces to a reproducing kernel Hilbert space (RKHS), the model computes the cross-covariance between the two spaces. The MI calibration loss is designed to penalize the variance of mutual information estimates across different domains, thereby enforcing semantic balance. Additionally, it encourages each domain to retain sufficient domain-specific information.

The overall training process optimizes the model end-to-end by minimizing a combined objective function. This total loss aggregates the reconstruction loss, the RQ loss, and the MI calibration loss, weighted by respective hyperparameters. Once trained, UniTok tokenizes each item into discrete semantic tokens, which are then utilized by downstream LLM-based recommenders to predict user interactions based on item token sequences.

Experiment

The experiments validate UniTok's unified tokenization framework across ten recommendation domains, showing consistent accuracy gains over both collaborative filtering and tokenization-aided methods, with notable improvements in semantic capture and generalization. Efficiency experiments reveal a substantial reduction in trainable parameters compared to traditional per-domain codebook tokenizers, while maintaining superior performance even under a shared training budget. Zero-shot tests on unseen domains confirm robust transferability without retraining. Ablation studies demonstrate that both the TokenMoE module and mutual information calibration are essential, as removing either degrades recommendation quality.

UniTok consistently outperforms all competing methods across ten diverse benchmark datasets in NDCG@10, with a maximum improvement of 51.89% on the Tools dataset. The results confirm that unified tokenization with shared semantics is more effective than dataset-specific tokenizers, while LLM-based recommender systems generally surpass traditional collaborative filtering baselines. UniTok achieves the highest NDCG@10 on every dataset, with the largest gain of 51.89% on Tools. LLM-based tokenization methods (TIGER, LC-Rec, LETTER, UniTok) outperform traditional collaborative filtering models like MF and LightGCN. Unified tokenization maintains strong accuracy across domains, whereas competitors trained jointly show substantial degradation.

UniTok reduces trainable parameters by about an order of magnitude compared to codebook-based methods by using a single shared model across datasets, while competitors accumulate parameters per dataset. The savings come almost entirely from a smaller autoencoder, with codebook and router parameters being minimal. UniTok's total trainable parameters are roughly one tenth of codebook-based competitors (9.11M vs 87.78M), achieved through parameter sharing across all ten datasets. The autoencoder accounts for nearly all parameter reduction (8.75M vs 87.45M), while codebook size is similar (0.36M vs 0.33M) and the router adds only 0.01M.

Under a unified training setup with comparable parameter budgets, UniTok consistently outperforms traditional codebook-based tokenizers across all evaluated domains, while competitors suffer substantial performance degradation. The gains are particularly pronounced on Cellphones, where UniTok improves Recall@10 by up to 84.51%. UniTok achieves large relative gains over all competitors, with Recall@10 improvements ranging from 65.60% on Beauty to 84.51% on Cellphones. Competing methods degrade notably in the unified multi-domain setting, likely due to difficulty in distinguishing items from different domains with shared tokenization. UniTok's modular TokenMoE architecture enables domain-specific experts to learn independently while sharing a unified token space, preserving performance.

In a zero-shot generalization test across three unseen domains (Clothing, Health, Sports), UniTok consistently achieves the best recall and NDCG at 10 compared to existing item tokenization-based methods, with relative gains ranging from about 8% to nearly 18%. The improvement is most pronounced on the Health domain, where NDCG@10 increases by 17.87% over the strongest baseline. UniTok attains the highest Recall@10 and NDCG@10 on all three unseen datasets. The largest relative gain is 17.87% in NDCG@10 on the Health domain. Baselines like TIGER, LC-Rec, and LETTER perform lower across all metrics, with UniTok showing up to 15.88% improvement in Recall@10 on Sports.

An ablation study progressively removing the TokenMoE module, shared expert, and MI calibration shows that each component contributes to recommendation accuracy. The largest performance drop occurs when TokenMoE is removed, while MI calibration provides additional consistent gains across all datasets. The full model achieves the best results on every dataset. Removing TokenMoE causes the most significant performance decline across all datasets. Adding the shared expert improves accuracy over using only TokenMoE without it. MI calibration further boosts performance on top of TokenMoE and the shared expert. The complete UniTok model consistently outperforms all ablated variants.

The evaluation shows UniTok consistently outperforms all baselines across ten benchmark datasets, with large gains in NDCG@10 and Recall@10, and superior zero-shot generalization on unseen domains. It achieves these results while reducing trainable parameters by nearly an order of magnitude through shared tokenization and a modular architecture. Ablation studies confirm that the TokenMoE module, shared expert, and MI calibration each contribute meaningfully to accuracy, with the full model yielding the best performance across all settings.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp