HyperAIHyperAI

Command Palette

Search for a command to run...

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

Haoyi Zhou; Shanghang Zhang; Jieqi Peng; Shuai Zhang; Jianxin Li; Hui Xiong; Wancai Zhang

Abstract

Many real-world applications require the prediction of long sequence time-series, such as electricity consumption planning. Long sequence time-series forecasting (LSTF) demands a high prediction capacity of the model, which is the ability to capture precise long-range dependency coupling between output and input efficiently. Recent studies have shown the potential of Transformer to increase the prediction capacity. However, there are several severe issues with Transformer that prevent it from being directly applicable to LSTF, including quadratic time complexity, high memory usage, and inherent limitation of the encoder-decoder architecture. To address these issues, we design an efficient transformer-based model for LSTF, named Informer, with three distinctive characteristics: (i) a ProbSparseProbSparseProbSparse self-attention mechanism, which achieves O(Llog⁡L)O(L \log L)O(LlogL) in time complexity and memory usage, and has comparable performance on sequences' dependency alignment. (ii) the self-attention distilling highlights dominating attention by halving cascading layer input, and efficiently handles extreme long input sequences. (iii) the generative style decoder, while conceptually simple, predicts the long time-series sequences at one forward operation rather than a step-by-step way, which drastically improves the inference speed of long-sequence predictions. Extensive experiments on four large-scale datasets demonstrate that Informer significantly outperforms existing methods and provides a new solution to the LSTF problem.

One-sentence Summary

Researchers from Beihang University, UC Berkeley, and Rutgers University propose Informer, an efficient Transformer for long-sequence time-series forecasting that reduces quadratic self-attention complexity to O(Llog⁡L)O(L \log L)O(LlogL) via ProbSparse attention, distills dominating features, and uses a generative decoder for one-shot long-sequence prediction, outperforming prior methods on large-scale datasets.

Key Contributions

  • The ProbSparse self-attention mechanism achieves O(L log L) time and memory complexity for dependency alignments while maintaining comparable alignment performance to canonical self-attention.
  • The self-attention distilling operation halves the input of cascading layers and privileges dominating attention, reducing overall space complexity to O((2 − ε)L log L) for handling extremely long input sequences.
  • A generative style decoder generates long sequence outputs in a single forward step, drastically boosting inference speed and avoiding cumulative error; experiments on four large-scale datasets demonstrate that these components collectively outperform existing methods on long sequence time-series forecasting.

Introduction

Long sequence time-series forecasting (LSTF) underpins critical applications in sensor networks, energy management, finance, and disease modeling, where models must predict far into the future from lengthy historical inputs. Prior approaches, including recurrent networks and vanilla Transformers, falter as prediction horizons grow: Transformers capture long-range dependencies well but incur quadratic O(L²) time and memory from canonical self-attention, O(J·L²) memory when stacking layers, and slow autoregressive decoding that hampers long-output inference. Recent efficiency-focused Transformers (Sparse Transformer, LogSparse, Longformer, Reformer, Linformer) primarily reduce attention complexity, often to O(L log L), yet still suffer from the memory and decoding bottlenecks that limit real-world LSTF deployment. The authors address all three challenges by introducing Informer, which combines a ProbSparse self-attention mechanism with O(L log L) complexity, an attention distilling operation that drastically lowers per-layer memory in deep stacks, and a generative decoder that outputs the entire long forecast in one forward pass, eliminating error accumulation and enabling practical, high-capacity LSTF.

Method

The authors propose Informer, an encoder-decoder architecture tailored for long sequence time-series forecasting (LSTF). At its core, Informer addresses the quadratic complexity of canonical self-attention through a probabilistic sparsity measurement and an efficient query selection mechanism. Combined with a self-attention distilling operation in the encoder and a generative-style decoder, the model achieves significant reductions in memory usage and inference time without sacrificing predictive performance.

Efficient Self-attention Mechanism

The canonical scaled dot-product self-attention forms a probability distribution over keys for each query, expressed as a kernel smoother: A(qi,K,V)=∑jk(qi,kj)∑lk(qi,kl)vj=Ep(kj∣qi)[vj],\mathcal{A}(\mathbf{q}_i, \mathbf{K}, \mathbf{V}) = \sum_j \frac{k(\mathbf{q}_i, \mathbf{k}_j)}{\sum_l k(\mathbf{q}_i, \mathbf{k}_l)} \mathbf{v}_j = \mathbb{E}_{p(\mathbf{k}_j|\mathbf{q}_i)} [\mathbf{v}_j],A(qi​,K,V)=∑j​∑l​k(qi​,kl​)k(qi​,kj​)​vj​=Ep(kj​∣qi​)​[vj​], where k(qi,kj)=exp⁡(qikj⊤/d)k(\mathbf{q}_i, \mathbf{k}_j) = \exp(\mathbf{q}_i\mathbf{k}_j^\top / \sqrt{d})k(qi​,kj​)=exp(qi​kj⊤​/d​). The attention output is dominated by a small subset of dot-product pairs; most queries contribute trivial, near-uniform attention. To identify these “important” queries, the authors measure how far a query’s attention distribution ppp deviates from the uniform distribution qqq using the Kullback-Leibler divergence. After dropping constants, the sparsity measurement for the iii-th query becomes: M(qi,K)=ln⁡∑j=1LKeqikj⊤d−1LK∑j=1LKqikj⊤d.M(\mathbf{q}_i, \mathbf{K}) = \ln \sum_{j=1}^{L_K} e^{\frac{\mathbf{q}_i \mathbf{k}_j^\top}{\sqrt{d}}} - \frac{1}{L_K} \sum_{j=1}^{L_K} \frac{\mathbf{q}_i \mathbf{k}_j^\top}{\sqrt{d}}.M(qi​,K)=ln∑j=1LK​​ed​qi​kj⊤​​−LK​1​∑j=1LK​​d​qi​kj⊤​​. A larger M(qi,K)M(\mathbf{q}_i, \mathbf{K})M(qi​,K) indicates a more “diverse” attention probability and a higher likelihood of containing the dominant dot-product pairs. Based on this, ProbSparse self-attention restricts each key to attend only to the top-uuu queries under this measurement: A(Q,K,V)=Softmax⁡(QˉK⊤d)V,\mathcal{A}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \operatorname{Softmax}\left(\frac{\bar{\mathbf{Q}} \mathbf{K}^\top}{\sqrt{d}}\right) \mathbf{V},A(Q,K,V)=Softmax(d​Qˉ​K⊤​)V, where Qˉ\bar{\mathbf{Q}}Qˉ​ contains only the u=c⋅ln⁡LQu = c \cdot \ln L_Qu=c⋅lnLQ​ most important queries. This reduces the dot-product computation for each query-key lookup to O(ln⁡LQ)\mathcal{O}(\ln L_Q)O(lnLQ​) and the layer memory to O(LKln⁡LQ)\mathcal{O}(L_K \ln L_Q)O(LK​lnLQ​). Different sparse query-key pairs are learned per head, preventing severe information loss.

Computing the exact M(qi,K)M(\mathbf{q}_i, \mathbf{K})M(qi​,K) still requires O(LQLK)\mathcal{O}(L_Q L_K)O(LQ​LK​) dot-products. To make the measurement practical, the authors derive an empirical approximation. Lemma 1 provides an upper bound that motivates the max-mean measurement: Mˉ(qi,K)=max⁡j{qikj⊤d}−1LK∑j=1LKqikj⊤d.\bar{M}(\mathbf{q}_i, \mathbf{K}) = \max_j \left\{ \frac{\mathbf{q}_i \mathbf{k}_j^\top}{\sqrt{d}} \right\} - \frac{1}{L_K} \sum_{j=1}^{L_K} \frac{\mathbf{q}_i \mathbf{k}_j^\top}{\sqrt{d}}.Mˉ(qi​,K)=maxj​{d​qi​kj⊤​​}−LK​1​∑j=1LK​​d​qi​kj⊤​​. Under the long-tail distribution of attention scores, it suffices to compute Mˉ\bar{M}Mˉ using only a random sample of U=LKln⁡LQU = L_K \ln L_QU=LK​lnLQ​ dot-product pairs, padding the rest with zeros. The top-uuu queries are then selected from this approximate set. This makes the total time and space complexity of ProbSparse self-attention O(Lln⁡L)\mathcal{O}(L \ln L)O(LlnL) when LQ=LK=LL_Q = L_K = LLQ​=LK​=L.

Encoder: Self-attention Distilling for Longer Sequences

The encoder processes the full input sequence Xent∈RLx×dmodel\mathbf{X}_{\text{en}}^t \in \mathbb{R}^{L_x \times d_{\text{model}}}Xent​∈RLx​×dmodel​. To further reduce memory and focus on dominant features, the encoder employs a distilling operation after each ProbSparse self-attention block. The output of the jjj-th layer is passed through a 1D convolution (kernel width 3) with an ELU activation, followed by a max-pooling layer with stride 2: Xj+1t=MaxPool⁡(ELU⁡(Conv1d⁡([Xjt]AB))).\mathbf{X}_{j+1}^t = \operatorname{MaxPool}\left( \operatorname{ELU}\left( \operatorname{Conv1d}\left( [\mathbf{X}_j^t]_{\text{AB}} \right) \right) \right).Xj+1t​=MaxPool(ELU(Conv1d([Xjt​]AB​))). This halves the temporal dimension after each layer, yielding an overall memory complexity of O((2−ϵ)Llog⁡L)\mathcal{O}((2-\epsilon) L \log L)O((2−ϵ)LlogL). To enhance robustness, the encoder replicates the main stack with progressively halved input lengths and drops one distilling layer per replica, forming a pyramid structure. The final encoder output is the concatenation of all stack outputs, providing a compact, multi-scale representation of the long input sequence.

Decoder: One Forward Generative Inference

The decoder follows a standard transformer design with two multi-head attention layers. To generate long output sequences without slow auto-regressive decoding, Informer adopts a generative inference scheme. The decoder input is constructed as: Xdet=Concat⁡(Xtokent,X0t)∈R(Ltoken+Ly)×dmodel,\mathbf{X}_{\text{de}}^t = \operatorname{Concat}(\mathbf{X}_{\text{token}}^t, \mathbf{X}_{\mathbf{0}}^t) \in \mathbb{R}^{(L_{\text{token}} + L_y) \times d_{\text{model}}},Xdet​=Concat(Xtokent​,X0t​)∈R(Ltoken​+Ly​)×dmodel​, where Xtokent\mathbf{X}_{\text{token}}^tXtokent​ is a known slice from the input sequence preceding the prediction window (acting as a “start token”), and X0t\mathbf{X}_{\mathbf{0}}^tX0t​ is a placeholder filled with zeros that provides the target time stamps. Masked ProbSparse self-attention sets the masked dot-products to −∞-\infty−∞, preventing each position from attending to future positions without auto-regression. The decoder then predicts the entire output sequence of length LyL_yLy​ in one forward pass. The model is trained to minimize the mean squared error (MSE) between the predicted and ground-truth target sequences, with gradients back-propagated through the entire encoder-decoder pipeline.

Experiment

The Informer model is evaluated on four real-world and public datasets for long sequence time-series forecasting, comparing against statistical, RNN-based, and other Transformer baselines. Experiments demonstrate that Informer consistently outperforms competitors, with prediction error growing slowly as the horizon lengthens, validating the benefits of ProbSparse self-attention, self-attention distilling, and generative-style decoding. Ablation studies confirm the individual contributions of these components, while runtime analysis shows improved training and inference efficiency over related self-attention methods.

Informer and its variant achieve the lowest MSE and MAE across all prediction lengths on both ETTh1 and ETTh2, while Reformer and Prophet suffer large errors as the horizon grows. LogTrans performs competitively, and classical methods like ARIMA show strong results on ETTh1 but fail on ETTh2, highlighting dataset sensitivity. Informer and Informer† maintain low errors across all horizons; for ETTh1 at horizon 720, Informer† MSE is 0.257 versus Reformer's 2.112 and Prophet's 2.735. Reformer degrades sharply with longer horizons, reaching MSE 1.860 at horizon 336 and 2.112 at 720 on ETTh1. ARIMA yields MSE 0.108 at horizon 24 on ETTh1 but 3.554 on ETTh2, showing strong dataset dependence. Prophet performs reasonably at short horizons but deteriorates dramatically at horizon 720 on ETTh1 (MSE 2.735). LogTrans stays close to Informer, with MSE 0.273 at ETTh1 horizon 720, outperforming Reformer and classical methods.

On the ETTh1 and ETTh2 multivariate forecasting tasks, Informer consistently yields the lowest errors across all prediction horizons compared to LogTrans, Reformer, LSTMa, and LSTnet. Its advantage grows with longer horizons, where competing models show steep error increases, demonstrating more robust long-range prediction. A variant of Informer maintains comparable performance, confirming the overall effectiveness. Informer achieves the lowest MSE and MAE on both ETTh1 and ETTh2 at every horizon, with errors often less than half those of Reformer and LSTnet at long forecast lengths. As the prediction horizon extends from 24 to 720 on ETTh1, Informer's errors rise gradually while LogTrans, Reformer, LSTMa, and LSTnet degrade sharply, widening the performance gap.

Informer with ProbSparse self-attention processes encoder inputs up to length 2880 without memory failure, while canonical self-attention (Informer†) and LogTrans run out of memory at encoder input 1440. As the encoder length increases, Informer’s prediction error decreases consistently for both prediction horizons 336 and 720, demonstrating that longer input sequences improve accuracy. Informer achieves lower MSE and MAE than Informer† and LogTrans at every tested encoder length, and its error continues to drop when encoder input grows from 336 to 1440 or 2880. Informer† and LogTrans encounter out-of-memory failures at encoder input length 1440, while Informer scales to 2880 without any failure, enabling the use of much longer historical context.

The ProbSparse self-attention in Informer lowers time complexity to O(L log L) and avoids the O(L^2) memory bottlenecks that caused out-of-memory failures in LogTrans. Combined with a generative decoder that produces the whole forecast in one step, Informer achieves efficient training and inference on long sequences. Informer's ProbSparse self-attention reaches smaller memory usage than full-attention variants, preventing out-of-memory errors observed with the public LogTrans implementation. The Informer decoder generates predictions in a single step (1), whereas Transformer, Reformer and LSTM rely on L autoregressive steps. LogTrans suffers from O(L^2) test-time complexity and O(L^2) memory in its standard implementation, while Informer reduces both to O(L log L) with constant-step decoding.

Informer‡ maintains stable MSE and MAE across all prediction offsets for both 336 and 480 horizon lengths. In contrast, Informer§ only returns valid results when the prediction offset is zero and fails completely for all positive offsets, indicating its dynamic decoding cannot handle shifted forecasting. Informer‡ achieves consistent performance with MSE around 0.20–0.21 at horizon 336 and 0.20–0.21 at horizon 480, with MAE rising only slightly as offset increases. Informer§ obtains an MSE of 0.201 and MAE of 0.393 at offset +0 for horizon 336, but yields unacceptable metrics (‘-’) for all offsets from +12 onward. The failure of Informer§ on non-zero offsets highlights the importance of the generative style decoder for capturing long-range dependencies without error accumulation.

The experiments evaluate multivariate time series forecasting on ETTh1 and ETTh2 datasets across varying prediction horizons, along with scalability tests on encoder length and decoding style. Informer consistently outperforms baselines like Reformer, LogTrans, and ARIMA, with its advantage growing at longer horizons, while its ProbSparse self-attention reduces memory and time complexity to O(L log L), enabling processing of much longer input sequences without out-of-memory failures. The generative decoder further improves robustness by producing forecasts in a single step, whereas dynamic decoding fails for shifted predictions, confirming that these design choices together yield accurate, efficient, and scalable long-horizon forecasting.


Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing

HyperAI Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp