HyperAIHyperAI

Command Palette

Search for a command to run...

TokenRouter : un système de service efficace pour le routage de LLM au niveau des jetons

Tianyu Fu Tengxuan Liu Ruoxi Wang Yixin Dong Yi Ge Yichen You Yu Wang

Résumé

Le routage de grands modèles de langue (LLM) répartit la charge d'inférence entre différents modèles, repoussant la frontière de Pareto coût-qualité dans le service des LLM. Bien que le routage à gros grain au niveau de la session ou de la requête soit largement adopté dans les systèmes de production, des travaux algorithmiques récents montrent qu'un routage à grain fin au niveau des jetons peut apporter des gains substantiels en efficacité et en qualité. Cependant, l'exécution efficace d'inférences routées au niveau des jetons pose des difficultés importantes aux systèmes existants. Construits sur l'hypothèse d'un LLM unique, les systèmes actuels souffrent d'une forte désynchronisation des étapes et de fréquents retards d'admission des lots avec le routage au niveau des jetons, tout en imposant une grande complexité d'implémentation aux développeurs. Pour répondre à ces défis, nous concevons TokenRouter, un système de service d'inférence efficace et convivial pour les développeurs, destiné à l'inférence de LLM routée au niveau des jetons. TokenRouter suit le principe d'une programmation orientée requête et d'une exécution orientée modèle : les développeurs décrivent la logique de routage du point de vue d'une requête unique, tandis que le moteur d'exécution lance un sous-serveur pour chaque LLM et distribue les requêtes de manière asynchrone. Chaque sous-serveur utilise un ordonnanceur de mise en lots différée, dont les hyperparamètres optimaux sont déduits d'un modèle mathématique de débit du système. Sur divers algorithmes de routage, charges de travail et paires de modèles, TokenRouter atteint un débit de décodage de 2,01 à 64,15 fois supérieur à celui des systèmes existants, améliorant ainsi considérablement l'efficacité de service du routage de LLM au niveau des jetons. Notre code est disponible à l'adresse https://github.com/thu-nics/TokenRouter.

One-sentence Summary

Researchers from Tsinghua University and Carnegie Mellon University propose TokenRouter, an efficient serving system for token-level LLM routing that employs request-centric programming and model-centric execution to dispatch requests asynchronously to per-model subservers, each using a delayed-batching scheduler whose optimal hyperparameters derive from a mathematical throughput model, achieving 2.012.012.01 to 64.1564.1564.15 times higher decoding throughput than existing systems.

Key Contributions

  • TokenRouter is a serving system for token-level LLM routing that offers a request-centric route-send-receive programming interface and a model-centric runtime with one subserver per model for asynchronous request dispatch.
  • It contributes a decoupled tri-loop execution design with a handoff-resume mechanism to address step desynchronization across cooperating LLMs, along with a delayed-batching scheduler whose optimal hyperparameters are derived from a throughput-based mathematical model.
  • Experiments across diverse routing algorithms, workloads, and model pairs show that TokenRouter achieves 2.01 to 64.15 times higher decoding throughput than existing implementations.

Introduction

Large language models vary widely in size, expertise, latency, and cost, and serving systems increasingly route requests across models to improve the cost-quality tradeoff. Prior routing is mostly coarse-grained at the session or query level, so responses are generated by a single model; token-level routing can exploit difficulty variation within a response and enable complementary models to collaborate, but current serving systems assume synchronous single-model decoding. This makes token-level routing inefficient due to step desynchronization, batch admission delays, and implementation complexity. The authors propose TokenRouter, a serving system that separates a request-centric route-send-receive programming interface from model-centric asynchronous execution, with delayed batching to reduce admission latency. It provides a drop-in server interface and achieves 2.01 to 64.15 times higher decoding throughput than existing implementations.

Method

The authors introduce TokenRouter, a system designed to support diverse token-level routing algorithms while keeping their implementation simple and intuitive. The core of the programming interface relies on a sequential request-centric programming principle. Instead of requiring developers to reason about low-level batching, scheduling, or asynchronous execution, the interface allows them to describe a routing algorithm by tracing how a single request moves among cooperating large language models. The lifecycle of a request follows a sequence: receive, decode, route, send, peer receive, peer decode, and so on.

To implement this, developers only need to fill in three components: route, send, and receive. The route function is called after each decoding step and determines a destination index for each request based on the forward-pass result. A destination of 0 means the request continues locally, while a nonzero destination points to a peer model. The send function is triggered when a request is delegated, returning a customized message that carries the request ID, the suffix of tokens the peer model has not yet seen, and a status field. The receive function converts the incoming peer request to the local request format so the receiver model can continue decoding.

The system design of TokenRouter is organized as a model-centric runtime. As shown in the figure below, the runtime consists of multiple subservers behind a single external interface. Each subserver hosts one candidate LLM, owning a scheduler, an LLM runner with a private KV-cache pool, and the three user-defined functions. Subservers progress independently and communicate through peer requests. Externally, TokenRouter exposes a single server interface, making it a drop-in replacement for existing single-LLM servers.

To address the step desynchronization challenge inherent in token-level routing, the authors implement a decoupled tri-loop execution and a handoff-resume mechanism. Standard servers typically have a client-server loop and a decoding loop. TokenRouter adds a third decoupled loop, the inter-model loop, to send and receive peer requests between subservers. This allows a routed request to leave the local batch and resume when the peer request returns, while remaining local requests continue executing. For the handoff and resume mechanism, the system introduces a pending state. A pending request is skipped by the scheduler but its serving state is preserved. Upon resumption, the state is toggled to running or finished, and new tokens are appended, reducing inter-model transitions to little more than a token append.

While asynchronous execution removes synchronization across subservers, each subserver still requires a scheduling policy to handle sparse and irregular routed-token arrivals. The authors define batch admission delay as the waiting time between a request arrival and the start of its execution. Under eager asynchronous scheduling, if a subserver starts a new decoding step as soon as any peer request arrives, requests arriving during an ongoing batch cannot be admitted until that step finishes, leading to fragmented batching.

To reduce this admission delay, the authors propose a delayed-batching scheduler. The key idea is to wait for a short, controlled interval so that more requests targeting the same model can be executed together. The scheduler buffers received requests and launches a batch only when the buffer size reaches a threshold BBB. As illustrated in the figure below, this approach allows later-arriving requests to be admitted with much less waiting, dropping the average batch admission delay.

To determine the throughput-optimal delayed-batching threshold B∗B^*B∗, the authors model the routing process as a Discrete-Time Markov Chain. Given a token-level routing algorithm, concurrency NNN, routing probability PPP, and per-step decoding latency LiL_iLi​ of each LLM, the model derives the throughput as a function of BBB and searches for the threshold that maximizes throughput. This balances the trade-off where too small a BBB causes large batch admission delay, while too large a BBB traps requests in the queue.

Experiment

The evaluation compares TokenRouter with five token-level routing algorithms across three workloads and two baselines, using Qwen3 small and large model pairs. TokenRouter consistently improves serving efficiency, maintaining throughput with longer outputs and outperforming official implementations under their original settings. Ablations show that engineering optimizations, asynchronous execution, and delayed batching account for most gains, and the benefits persist across model pairs and parallelization choices. Finally, TokenRouter shifts token-level routing to a better throughput-accuracy Pareto frontier, making it competitive with query-level routing.

The evaluated token-level routing algorithms use diverse routing signals but are all expressible through a common route function. CITER, R2R, and Co-LLM follow an asymmetric pattern where one model invokes another on uncertain tokens and control returns after one generated token. R-Stitch uses symmetric entropy-based switching between two models, while ME supports more than two models and reselects the next model after every token based on ensemble weights. CITER, R2R, and Co-LLM use asymmetric handoffs triggered by low-confidence, predicted-divergent, or high-deferral tokens, with control returning after one token. R-Stitch switches symmetrically based on token entropy, while ME generalizes routing to multiple models and selects after every token via ensemble weights.

At concurrency 4, TokenRouter outperforms the official R2R and CITER implementations, improving throughput and reducing end-to-end latency while maintaining similar time-to-first-token. Reported gains across the evaluated original token-level routing settings reach 2.73 to 21.97 times for throughput and 2.78 to 21.61 times for latency reduction. For R2R, TokenRouter more than doubles throughput and cuts end-to-end latency by roughly two-thirds over the official implementation, with TTFT unchanged. For CITER, TokenRouter keeps TTFT identical while lowering end-to-end latency from seconds to under half a second and improving throughput nearly ninefold over the official code.

Across the evaluated small language model and large language model pairs, TokenRouter consistently provides higher throughput than the official R2R implementation. The improvement ranges from about 2x to 3.2x, with the largest gain on the smallest model pairing and substantial gains continuing as model scales increase. These results suggest the optimizations are broadly effective across different model sizes. TokenRouter achieves 1.99x to 3.21x higher throughput than R2R across all tested SLM-LLM pairs. The largest throughput gain occurs with the 0.6B SLM paired with the 8B LLM, while even the smallest gain nearly doubles throughput.

The experiments compare token-level routing strategies that can be expressed through a common route function, including asymmetric handoff methods, symmetric entropy-based switching, and multi-model ensemble routing. Evaluations against official R2R and CITER implementations show that TokenRouter substantially improves throughput and reduces end-to-end latency while keeping time-to-first-token similar. Across multiple small and large language model pairs, TokenRouter consistently maintains throughput gains over R2R, with the largest benefit on the smallest tested pairing. Overall, the results validate that the optimizations are broadly effective across routing patterns and model scales.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp