Command Palette
Search for a command to run...
지각된 음성의 해석 가능한 MEG 디코딩: 피질 신호원과 검색을 유도하는 자극 특징
지각된 음성의 해석 가능한 MEG 디코딩: 피질 신호원과 검색을 유도하는 자극 특징
Ilia Semenkov Daria Kleeva Zarina Maksudova Ivan Dakhtin Alex Ossadtchi
초록
지각된 짧은 음성 구간은 wav2vec 2.0 오디오 임베딩에 대해 CLIP 방식의 목적 함수로 훈련된 심층 신경망을 통해 비침습적 뇌자도(MEG) 기록으로부터 검색될 수 있다[1–3]. 이러한 연구들은 가중치를 고전적인 전기생리학에서 사용되는 개념으로 변환하지 않는다. [1]의 전단부에 있는 공간 필터는 원칙적으로 신호원 지형도로 매핑될 수 있지만, 그러한 신호원의 동역학적 특성은 여전히 파악하기 어렵다. 음성 스트림의 어떤 속성이 디코딩에 기여하는지도 마찬가지로 불분명하다. 본 연구에서는 측정 물리학과 신호원의 생리학에 의해 전단부가 제약된 디코더를 면밀히 조사한다. Petrosyan 등의 프레임워크[4, 5]를 기반으로, [1]의 2D 푸리에 공간 어텐션을 구면 조화 함수[6]로 매개변수화된 계층으로 대체하고, 피험자 특이적 표현을 270개에서 K = 25개의 분기로 줄이며, 훈련 가능한 시간 필터 계층을 추가하여 네트워크의 각 분기가 공간뿐만 아니라 시간적으로도 신경 신호원에 대응되도록 한다. 안구 및 심장 성분은 훈련 전에 제거되는데, 이는 둘 중 하나가 자극에 고정된 정보를 제공하여 그렇지 않으면 피질 신호로 오인될 수 있기 때문이다. 정제된 MEG-MASC 데이터셋[7]에서 이 모델은 6개의 훈련된 솔루션에 걸쳐 1005개의 후보 중 39.75 ± 0.34%의 Top-1 정확도에 도달하며, 디코더의 훈련 가능한 매개변수 수는 약 20배 더 적다. 이 모델의 가중치는 신호원 공간[4, 8]으로 매핑되어 표준적인 음성 지각 네트워크와 일치하는 생성원을 복원하며, 좌측에 국한된 분기는 우측에서는 뚜렷하지 않은 더 높은 주파수의 리듬 성분을 전달한다. 특징이 표시된 음성 구간을 특징 존재 구간 및 특징 부재 구간의 일치하는 기증 구간으로 대체하는 쌍체 MEG 폐색 실험은 19개의 자극 특징 중 15개가 기여하며, 가장 큰 효과는 묵음, 소리 강도, 모음 및 음향적 개시에서 나타남을 보여준다. 흥미롭게도, 무작위로 배열된 단어 목록은 반대 방향으로 작용한다. 즉, 여기에 내러티브 MEG를 대입하면 검색 성능이 향상되므로, 내러티브 구조가 제거된 단어들에 의해 유발된 활동은 일관된 음성에 의해 유발된 활동보다 복원 가능한 정보를 더 적게 전달한다. wav2vec 타겟은 검색 정확도의 손실 없이 약 12개의 학습된 특징 차원으로 축소될 수 있는 반면, 강한 시간적 압축은 명확한 성능 저하를 초래한다. 따라서 물리적 및 생리학적으로 제약된 디코더는 지식 발견 도구로 기능할 수 있다. 즉, 학습된 가중치는 피질 신호원과 시간적 동역학으로 매핑될 수 있으며, 입력 개입 실험은 결정 규칙이 무엇에 의존하는지 밝혀준다.
One-sentence Summary
HSE University and ITMO University researchers propose a physiologically constrained MEG speech decoder that replaces 2D Fourier attention with a spherical harmonics layer, adds trainable temporal filters, and reduces subject-specific branches to K=25, mapping learned weights to canonical speech-perception cortical sources while achieving 39.75±0.34% Top-1 accuracy with approximately 20× fewer parameters, and occlusion reveals that silence, sound intensity, vowels, and acoustic onsets drive retrieval most while narrative structure strongly impacts recoverable information.
Key Contributions
- A decoder with a spherical-harmonics front end and trainable temporal filters reaches 39.75% top-1 accuracy on MEG-MASC, using roughly 20× fewer parameters than previous work.
- The architecture’s interpretability allows learned weights to be mapped to cortical sources, recovering the canonical speech-perception network and showing that left-hemisphere branches carry higher-frequency rhythmic dynamics.
- Paired MEG occlusion and narrative-substitution experiments reveal that the model relies on acoustic features such as silence, intensity, vowels, and onsets, that coherent speech context improves retrievable information, and that feature-use patterns remain stable across six training seeds.
Introduction
Decoding perceived speech from non-invasive magnetoencephalography (MEG) has reached high retrieval accuracy, but the resulting models remain opaque: their learned weights do not correspond to recognizable neural sources, rhythms, or time courses, so the high scores cannot be linked to specific cortical computations. While compact factorized architectures that separate spatial from temporal filtering have been used in brain decoding, earlier attempts to interpret them overlooked the mutual dependence of jointly trained filters and were not applied to whole-head MEG during complex natural speech tasks. The authors address this gap by extending a physiologically grounded front-end that factorizes spatial and temporal processing, equipping it with spherical-harmonic attention, trainable depthwise temporal filters, a small bottleneck of 25 branches, and explicit removal of ocular and cardiac artifacts. This architecture achieves competitive retrieval accuracy while allowing the trained weights to be directly mapped onto cortical source locations and their second-order dynamics. Through paired MEG substitution experiments, the authors further reveal which stimulus properties — spanning acoustics, phonetics, and contextual surprisal — the decoder actually uses, transforming the network from a black-box benchmark into an instrument for neurophysiological discovery.
Dataset
The authors use the MEG-MASC dataset, a collection of simultaneous audio and magnetoencephalography (MEG) recordings from 27 English-speaking participants listening to narrated stories from the MASC corpus. The dataset comprises 49 session recordings (22 participants contributed two sessions, five contributed one), each roughly one hour long.
Key dataset characteristics and processing steps:
- Audio preprocessing: The speech is resampled to 16 kHz and cut into 3-second windows with a 1-second stride. Windows whose peak absolute amplitude falls below 10−4 are discarded. A window is kept only if at least 50 % of the duration of one or more annotated words falls inside it.
- MEG preprocessing: Ocular and cardiac ICA components are removed. The signals are downsampled from 1000 Hz to 100 Hz. For each participant–session–story recording, the per-channel mean over the 0.5 s before the first stimulus onset is subtracted, channels are robust-scaled with the median and interquartile range, standardized to zero mean and unit variance, and clipped to ±20 standard deviations.
- Target representation: For every retained audio window, the target is obtained by passing the audio through the wav2vec 2.0 Base model and averaging the outputs of the last four hidden layers at each model time step.
- Train/validation/test split:
- The development set (2698 segments) consists of the full stories LW1, Cable Spool Fort, and Easy Money, plus the first five pieces of Black Willow. Within this set, the fifth piece of Black Willow is held out as the validation set.
- The test set (1005 segments) comprises the last seven pieces of Black Willow.
- For Black Willow, scaling parameters are fitted only on samples before the first test piece to avoid leakage.
- Audio–MEG pairing: Each 3-second audio segment is paired with the 3-second MEG segment starting 150 ms later to account for auditory response latency.
- Test segment alignment: Unlike common practice, test segments are not aligned to word onsets, making the retrieval setting more challenging.
Method
The authors address the retrieval task by training a network to construct embeddings for MEG data that align with audio embeddings produced by wav2vec 2.0, utilizing a CLIP-style objective. The proposed architecture replaces standard spatial-attention layers with a physically motivated 3D spatial attention layer, augments the model with a temporal-filtering layer, and modifies the convolutional decoder.
As shown in the figure below:
The interpretable front-end processes the input MEG data through a factorized spatial-temporal structure. The spatial filtering stage begins with a 3D spatial attention layer. Because MEG sensors occupy a three-dimensional, approximately spherical arrangement, the authors parameterize each of the J=270 virtual channels using real spherical harmonics. The unnormalized coefficient for virtual channel j and sensor m is computed as:
cjm=ℓ=0∑L−1q=−ℓ∑ℓγjq,ℓYℓq(θm,φm)where (θm,φm) are the polar and azimuthal angles of sensor m, Yℓq is a real spherical-harmonic basis function, and γjq,ℓ is a learned parameter. The coefficients are normalized across the M sensors using a softmax function and applied to the input signal.
Following the spatial attention, a shared 1×1 unmixing convolution applies a learned affine transformation in the channel space. A subject-specific layer then projects this representation to K interpretable branches. The effective participant-specific spatial filtering matrix is defined as W(s)=WsWuC, and the corresponding branch-wise bias is b(s)=Wsbu. The branch signals before temporal filtering are computed as as(t)=W(s)xs(t)+b(s).
Refer to the framework diagram:
The front-end is designed as a collection of branches where each branch adapts to a particular neural source with specific spatial and dynamical properties. To target specific frequency ranges, the authors apply one trainable 1-D depthwise temporal filter to each of the K branch signals. Each filter has 15 samples, corresponding to 150 ms at the MEG sampling rate of 100 Hz. The temporal filters are shared across participants, whereas the preceding spatial projection is participant-specific. The output of branch k is obtained by applying its temporal filter to the spatially filtered signal:
rs,k(t)=(as,k∗hk)(t)The K branch-wise signals produced by the interpretable front-end are then passed to a non-linear temporal decoder. This decoder comprises B temporal convolutional blocks followed by a convolutional head. Each temporal block contains three one-dimensional convolutions with specific dilation factors, batch normalization, and GELU activation. The convolutional head projects the decoder channels to the 768-dimensional wav2vec feature channels.
The training objective is a one-directional MEG-to-audio contrastive cross-entropy loss. MEG-derived embeddings are compared with the unique audio targets represented in the current minibatch. Similarities are computed after L2 normalization over the feature-time dimensions and divided by a learned temperature parameter. The models are trained using the AdamW optimizer with early stopping based on validation loss.
To understand the contribution of the spatial and temporal factorization, the authors perform architectural ablations of the front-end components.
As shown in the figure below:
The ablation results indicate that the full spatial-temporal factorization performs best. The largest degradation in retrieval accuracy occurs when subject-conditioned spatial mappings are removed, highlighting the necessity of adapting the spatial projection to individual subjects due to variations in anatomy and sensor geometry. Removing the attention layer or replacing the 3D attention with a 2D version also reduces performance, supporting the use of a sensor-geometry-aware parameterization.
The authors also investigate the effect of temporal-filter support on retrieval accuracy.
As shown in the figure below:
Retrieval depends mainly on whether the filter has sufficient temporal support. The one-sample condition, which contains no temporal context, performs worse than the 150 ms default. Performance generally improves as temporal support increases up to approximately 150 ms, after which gains become less systematic, indicating that the benefit of filter length begins to saturate around this scale.
Furthermore, the authors analyze the capacity of the model by varying the number of interpretable branches K and the number of convolutional blocks in the decoder.
As shown in the figure below:
Accuracy increases sharply from very small K to approximately K=10−25, then enters a broad plateau. Larger values of K do not produce systematic gains and can mildly degrade performance. Across decoder depths, the 0-block model is consistently weaker, while models with 2 to 5 convolutional blocks form a similar high-performing regime. The main configuration with 2 convolutional blocks and K=25 branches lies on this compact high-accuracy plateau.
Experiment
A series of experiments validated an interpretable MEG-to-speech retrieval decoder trained contrastively on narrative listening data, using paired occlusion, spatial clustering, and architectural ablations. The decoder relied on a compact set of stimulus features, including silence, loudness, vowels, and acoustic onsets, with spatially organized filters concentrated over bilateral auditory, frontal, and superior temporal cortices. Performance improved monotonically with longer MEG-audio segments, and the audio target representation could be drastically compressed along the feature axis through a learned low-dimensional subspace, while temporal resolution remained critical. Subject-specific spatial mapping, 3D geometry-aware attention, and temporal filtering all contributed to retrieval, with performance resting on a broad plateau across many architectural configurations.