Command Palette
Search for a command to run...
Convolutional Neural Networks for Sentence Classification
Convolutional Neural Networks for Sentence Classification
Yoon Kim
Abstract
We report on a series of experiments with convolutional neural networks (CNN) trained on top of pre-trained word vectors for sentence-level classification tasks. We show that a simple CNN with little hyperparameter tuning and static vectors achieves excellent results on multiple benchmarks. Learning task-specific vectors through fine-tuning offers further gains in performance. We additionally propose a simple modification to the architecture to allow for the use of both task-specific and static vectors. The CNN models discussed herein improve upon the state of the art on 4 out of 7 tasks, which include sentiment analysis and question classification.
One-sentence Summary
Researchers at New York University introduce a simple convolutional neural network for sentence classification that, with minimal hyperparameter tuning, combines static pre-trained word vectors and fine-tuned task-specific vectors through a straightforward architecture modification, outperforming the state of the art on four of seven benchmarks including sentiment analysis and question classification.
Key Contributions
- A simple CNN with one convolutional layer on static pre-trained word vectors, and minimal hyperparameter tuning, achieves strong results across multiple sentence classification benchmarks, suggesting the vectors act as effective generic feature extractors.
- Fine-tuning the pre-trained vectors for each task yields additional performance improvements over using static vectors alone.
- A multi-channel CNN architecture is introduced that combines static and task-specific vectors. Models with this architecture advance the state of the art on 4 out of 7 tasks, including sentiment analysis and question classification.
Introduction
The rise of deep learning in computer vision and speech recognition spurred interest in applying similar techniques to natural language processing, where dense word vectors learned by unsupervised neural language models had already begun to replace sparse one-hot encodings. Convolutional neural networks, while originally developed for image data, had shown promise on NLP tasks such as semantic parsing and sentence modeling, but it was uncertain how well a minimalist CNN architecture could perform when relying solely on generic, pretrained word embeddings without task-specific tuning. The authors demonstrate that a single convolutional layer placed directly on static word2vec vectors, without further architecture engineering, yields excellent results across multiple sentence classification benchmarks. They further introduce a simple multi-channel extension that combines static and fine-tuned embeddings, reinforcing the view that unsupervised pretrained word representations serve as universal feature extractors for text classification.
Dataset
The authors evaluate their model on a collection of standard sentence-level classification benchmarks, summarized in the paper’s Table 1. The following datasets are used:
- MR: Movie reviews, one sentence per review, binary positive/negative classification.
- SST-1: Stanford Sentiment Treebank extension of MR, with five fine-grained labels (very positive, positive, neutral, negative, very negative) and provided train/dev/test splits.
- SST-2: Same as SST-1 but with neutral reviews removed and binary labels.
- Subj: Subjectivity dataset where the goal is to classify a sentence as subjective or objective.
- TREC: TREC question dataset classifying questions into six types (person, location, numeric information, etc.).
- CR: Customer reviews of various products, binary positive/negative sentiment.
- MPQA: Opinion polarity detection subtask.
All tasks involve single sentences (or short questions). The authors follow existing splits when available (e.g., SST-1), and rely on standard practices for the others.
For data preprocessing, the model uses the publicly available word2vec vectors trained on 100 billion words from Google News. These 300-dimensional vectors, trained with the continuous bag-of-words architecture, initialize all in-vocabulary words; any word absent from the pre-trained set is initialized randomly. No other cropping, filtering, or dataset-specific augmentation is applied beyond the original dataset composition. The model is then fine-tuned on each task’s training set, evaluating on the corresponding test splits.
Method
The authors leverage a convolutional neural network architecture for sentence classification. Let xi∈Rk be the k-dimensional word vector corresponding to the i-th word in a sentence. A sentence of length n (padded where necessary) is represented as the concatenation of its word vectors: x1:n=x1⊕x2⊕⋯⊕xn where ⊕ is the concatenation operator.
Refer to the framework diagram for a visual representation of the model architecture.
The model processes these word vectors through a convolutional layer. A convolution operation involves a filter w∈Rhk, which is applied to a window of h words to produce a new feature. Specifically, a feature ci is generated from a window of words xi:i+h−1 by: ci=f(w⋅xi:i+h−1+b) Here b∈R is a bias term and f is a non-linear function such as the hyperbolic tangent. This filter is applied to each possible window of words in the sentence {x1:h,x2:h+1,…,xn−h+1:n} to produce a feature map c=[c1,c2,…,cn−h+1]. The model then applies a max-over-time pooling operation over the feature map and takes the maximum value c^=max{c} as the feature corresponding to this particular filter. This pooling scheme naturally handles variable sentence lengths by capturing the most important feature for each feature map. The model uses multiple filters with varying window sizes to obtain multiple features, which form the penultimate layer. These features are passed to a fully connected softmax layer to output the probability distribution over labels.
In a multichannel variant, the authors experiment with two channels of word vectors: one kept static throughout training and one fine-tuned via backpropagation. In this architecture, each filter is applied to both channels, and the results are added to calculate the feature ci.
For regularization, the authors employ dropout on the penultimate layer along with a constraint on the l2-norms of the weight vectors. Dropout prevents co-adaptation of hidden units by randomly setting a proportion p of the hidden units to zero during forward and backward propagation. Given the penultimate layer z=[c^1,…,c^m] (where m is the number of filters), instead of using y=w⋅z+b for the output unit y, dropout uses: y=w⋅(z∘r)+b where ∘ is the element-wise multiplication operator and r∈Rm is a masking vector of Bernoulli random variables with probability p of being 1. Gradients are backpropagated only through the unmasked units. At test time, the learned weight vectors are scaled by p such that w^=pw, and w^ is used to score unseen sentences. Additionally, the l2-norms of the weight vectors are constrained by rescaling w to have ∣∣w∣∣2=s whenever ∣∣w∣∣2>s after a gradient descent step.
Experiment
The study uses a convolutional neural network with fixed hyperparameters (filter widths 3–5, 100 feature maps each, dropout 0.5) selected via grid search on the SST-2 dev set and applied uniformly across all benchmarks. Models with pretrained word2vec vectors, even when kept static, significantly outperform a randomly initialized baseline, demonstrating that these vectors serve as strong universal feature extractors, and fine-tuning further improves performance by adapting word semantics to the task. The multichannel architecture yields mixed results, while dropout effectively regularizes larger networks and additional capacity proves crucial for achieving competitive results.
The benchmark datasets span sentiment analysis, question classification, and subjectivity detection, with substantial variation in size, sentence length, and vocabulary. Coverage of pre-trained word vectors is generally high but dips notably for the subjectivity corpus, while MPQA stands out for its extremely short sentences and near-complete vocabulary overlap. TREC and MPQA exhibit the shortest average sentence lengths (10 and 3 words), contrasting with Subj and MR (23 and 20 words). CR is the smallest dataset with 3,775 instances, while SST-1 is the largest among those with a fixed test set (2,210 test, 11,855 total). Subj has the highest vocabulary size (21,323) despite a moderate dataset size (10,000), and only 84% of its words appear in the pre-trained vectors—the lowest coverage across all datasets. MPQA achieves the highest pre-trained word coverage (97.4%) and a compact vocabulary (6,246) for its 10,606 examples, reflecting its very short, lexicon-restricted sentences.
Pre-trained word vectors provide large and consistent improvements for CNN sentence classifiers, with a static model matching or outperforming complex recursive and pooling-based methods across diverse sentence classification tasks. Fine-tuning these embeddings further lifts performance on most datasets, while a multichannel architecture yields mixed results. Dropout regularization adds 2–4% relative gain and enables training larger networks without overfitting. CNN-static, using fixed pre-trained vectors, outperforms recursive neural tensor networks on SST-2 (86.8% vs 85.4%) and rivals the best tree-based models without requiring syntactic parsing. Fine-tuning (CNN-non-static) sets a new high of 48.0% on SST-1, far above the 37.4% of a comparable CNN with random weights, demonstrating the value of both pre-training and task-specific adaptation.
Fine-tuning on sentiment analysis shifts word similarities from syntactic to sentiment-based relationships. In the static channel, good is nearest to bad due to syntactic interchangeability, while the fine-tuned non-static channel places good closest to nice and bad to terrible. Randomly initialized tokens like punctuation also acquire task-relevant meanings, such as exclamation marks associating with effusive expressions. Static word2vec vectors capture syntactic similarity: good is most similar to bad, but after fine-tuning good is closest to nice, aligning with sentiment. Fine-tuning teaches randomly initialized punctuation useful associations: exclamation marks become linked to effusive words and commas to conjunctive expressions.
The evaluation spans diverse sentence classification tasks (sentiment analysis, question classification, subjectivity detection) with varying dataset characteristics, and each experiment validates the impact of pre-trained word vectors and fine-tuning. Pre-trained embeddings deliver large and consistent gains, enabling a static CNN to match or surpass complex recursive models without syntactic parsing; fine-tuning further lifts performance while transforming word similarities from syntactic to sentiment-based relations (good near bad becomes good near nice). Dropout adds a reliable relative gain, a multichannel architecture proves inconsistent, and qualitative analysis shows that fine-tuning even imparts task-relevant meaning to randomly initialized punctuation, underscoring how task-specific adaptation turns pre-trained vectors into semantically focused representations.