HyperAIHyperAI

Command Palette

Search for a command to run...

Tactile-JEPA: تعلّم تمثيلات ذاتي الإشراف مُدرِك للطوبولوجيا من أجل حسّاسات لمسية موزّعة

Elizaveta Kovtun Matvey Konovalov Andrey Sakhovskiy Semen Budennyy

الملخص

يُعدّ الاستشعار اللمسي نمطًا أساسيًا للروبوتات التي تؤدي معالجة غنية بالتلامس وذات براعة، ولا سيما في ظل الانسداد البصري. في حين تُعدّ مُرمّزات الصور المُدرَّبة مسبقًا معيارًا في خطوط تعلّم الروبوت، لا تزال مُرمّزات اللمس تُدرَّب عادةً من الصفر انطلاقًا من إشارات خام ومشوَّشة، وهو ما قد يحدّ من قدرتها التعبيرية. تركّز الأساليب الحالية للتعلّم الذاتي الإشراف (SSL) بشكل أساسي على الحسّاسات اللمسية المعتمدة على الرؤية، مما يترك الجلود الإلكترونية الموزّعة دون تناولها إلى حد كبير. غير أنّ هذه الحسّاسات تمتلك خاصية مميزة: فعناصرها الاستشعارية متناثرة ومرتّبة بشكل غير منتظم على السطح الذي تغطيه، مما يجعل إعادة استخدام أساليب SSL البصرية مباشرةً خيارًا دون المستوى الأمثل. نقدّم Tactile-JEPA، وهي طريقة فعّالة للتدريب المسبق الذاتي الإشراف تستخدم الترتيب المكاني للحسّاسات اللمسية لتعلّم تمثيلات مُدرِكة للطوبولوجيا. وبالتحديد، يتم تدريبها على التنبؤ بتضمينات العناصر الاستشعارية المُقنَّعة انطلاقًا من العناصر المتبقية غير المُقنَّعة، باستخدام رسم بياني لاتصالية الحسّاسات لتوجيه الإخفاء المكاني. يُظهر تحليلنا أن التمثيلات اللمسية الفعّالة تتطلب التقاط كلٍّ من تفاصيل التلامس المحلية والحالة العامة للسطح اللمسي، وهو ما نحققه من خلال إخفاء ثنائي المقياس. عبر ثلاث مجموعات بيانات متنوعة تشمل حسّاسات مغناطيسية وبيزومقاومة، وتجسيدات مختلفة للروبوت، وتكوينات حسّاس مفرد ومزدوج، يقلّل Tactile-JEPA خطأ تقدير القوة بنسبة 6.3% وخطأ الاتجاه داخل اليد بنسبة 20.8% مقارنةً بالأداء الأحدث السابق، مع مكاسب متسقة في تطبيقات لاحقة أخرى، بما في ذلك تعلّم السياسات. إجمالًا، تُظهر نتائجنا أن فائدة الاستشعار اللمسي تعتمد بشكل حاسم على جودة التدريب المسبق للمُرمّز، وهي مشكلة يعالجها Tactile-JEPA مباشرةً. الكود متاح على https://github.com/E-Kovtun/tactile.

One-sentence Summary

The authors propose Tactile-JEPA, a topology-aware self-supervised pre-training method for distributed tactile sensors that predicts masked sensing-element embeddings from the unmasked remainder using sensor connectivity graphs and dual-scale masking to capture local contact details and global surface state, reducing force estimation error by 6.3% and in-hand orientation error by 20.8% over prior state-of-the-art.

Key Contributions

  • Tactile-JEPA is a self-supervised pre-training method for distributed tactile sensors that predicts embeddings of masked taxels from visible ones using the sensor connectivity graph and dual-scale masking to learn topology-aware, multi-scale representations directly from multivariate time-series signals.
  • Across magnetic and piezoresistive sensors, different robot embodiments, and single- and paired-sensor configurations, Tactile-JEPA reduces force estimation error by 6.3% and in-hand orientation error by 20.8% over the prior state of the art, with consistent gains in downstream applications including policy learning.
  • An experimental study shows that masks sampled over the sensor connectivity graph outperform image-style blocks, and that mask scale controls locality: local or global masks favor particular downstream tasks and, when mixed, transfer broadly.

Introduction

Robots are increasingly expected to generalize across many tasks and environments, and Vision-Language-Action models support this goal but rely primarily on vision. Vision alone becomes limiting under heavy occlusion, dexterous contact-rich manipulation, and fragile-object handling, where tactile sensing is critical. Distributed electronic skins are a promising tactile form factor, but they output multivariate time series over sparse, irregular taxels rather than images, so they lack the large-scale pretrained representations available for vision. Prior self-supervised tactile methods often inherit pixel-grid assumptions, encode only per-taxel location without sensor connectivity, or depend on task-specific labels and simulation-only geometry. The authors address this gap with Tactile-JEPA, a self-supervised representation learning method for distributed tactile sensors that predicts masked taxel embeddings from visible ones using masks sampled over the sensor connectivity graph at local and global scales, producing topology-aware, multi-scale representations.

Dataset

The authors evaluate Tactile-JEPA on three publicly available tactile datasets, selected to cover different transduction principles, embodiments, and downstream tasks.

  • Sparshskin Source: teleoperated play data collected with a sensorized robotic hand.

  • Tactile socks Source: human locomotion recorded from wearable tactile socks.

  • DECO-50 Source: teleoperated demonstrations of contact-rich manipulation tasks on a bimanual robot. Filtering: only the Assembly task subset is used, because it is the most tactile-intensive subset.

The text does not include full numeric sizes, filtering thresholds, cropping, or metadata construction details in this section. Those characteristics are reported in Table I of the paper.

In terms of usage, the authors use these datasets to evaluate Tactile-JEPA across varied tactile sensing setups and task types. The excerpt does not describe training split ratios or mixture weights; it focuses on dataset selection for evaluation.

Method

The authors propose Tactile-JEPA, a self-supervised representation learning framework tailored for distributed tactile sensing. The method learns to predict the latent representations of hidden tactile sensors from visible ones in the embedding space, effectively bypassing the direct prediction of noisy raw signals.

As shown in the framework diagram:

The overall pipeline is divided into two main phases: self-supervised pre-training and downstream task adaptation. During pre-training, the model operates on short time windows of tactile signals to capture contact dynamics without modeling long-term temporal structures. This design allows the encoder to be invoked at the control-loop rate during inference, with longer histories composed by successive embeddings.

Input Representation and Graph Construction Let the response of a specific taxel over a time window of τ\tauτ frames be denoted as xi(k)∈Rτ×mx_i^{(k)} \in \mathbb{R}^{\tau \times m}xi(k)​∈Rτ×m. The authors map this windowed response to a ddd-dimensional embedding using a shared affine projection followed by layer normalization:

x~i(k)=LN(Wvec⁡(xi(k))+b)∈Rd\tilde{x}_i^{(k)} = \mathrm{LN} \Big (W \operatorname{vec} \left(x_i^{(k)}\right) + b \Big ) \in \mathbb{R}^dx~i(k)​=LN(Wvec(xi(k)​)+b)∈Rd

where vec⁡(⋅)\operatorname{vec}(\cdot)vec(⋅) flattens the signal, and WWW and bbb are the learnable projection parameters. Stacking these embeddings across all NNN taxels yields the input representation X~τ(k)∈RN×d\tilde{X}_\tau^{(k)} \in \mathbb{R}^{N \times d}X~τ(k)​∈RN×d.

To capture the spatial topology of the sensor, the authors construct a taxel connectivity graph G=(V,E)G = (V, E)G=(V,E). The nodes VVV represent the taxels, and the edges EEE connect physically adjacent taxels based on the known sensor layout. This graph remains fixed across time windows and requires no explicit coordinate information.

Graph-Based Mask Sampling The self-supervised task relies on dividing the taxels into visible context regions and hidden target regions. The authors sample these regions directly on the taxel connectivity graph rather than a standard pixel grid. They define two distinct types of target masks to capture multi-scale contact patterns. A local mask forms a connected subgraph covering a compact sensor region to learn localized contact cues. Conversely, a global mask consists of taxels scattered uniformly across the entire graph to reflect the overall contact state of the embodiment. Both mask types are used jointly during training. Context masks are constructed similarly to global masks, ensuring no overlap with the target taxels.

Encoding and Prediction The tactile encoder is implemented as a transformer. Before processing, each taxel embedding is summed with a learnable positional embedding. The pre-training phase utilizes two encoder instances. The context encoder Eθ\mathbf{E}_\thetaEθ​ processes only the embeddings within the context region Cj\mathcal{C}_jCj​:

ZcontextCj=Eθ({x~i(k)+ei}i∈Cj)∈R∣Cj∣×dZ_{\text{context}}^{\mathcal{C}_j} = \mathbf{E}_\theta \big (\{\tilde{x}_i^{(k)} + e_i\}_{i \in \mathcal{C}_j} \big ) \in \mathbb{R}^{|\mathcal{C}_j| \times d}ZcontextCj​​=Eθ​({x~i(k)​+ei​}i∈Cj​​)∈R∣Cj​∣×d

The target encoder Eθˉ\mathbf{E}_{\bar{\theta}}Eθˉ​ shares the same architecture but processes all taxels to generate the ground truth representations:

Ztarget=Eθˉ({x~i(k)+ei}i∈V)∈RN×dZ_{\text{target}} = \mathbf{E}_{\bar{\theta}} \big (\{\tilde{x}_i^{(k)} + e_i\}_{i \in V} \big ) \in \mathbb{R}^{N \times d}Ztarget​=Eθˉ​({x~i(k)​+ei​}i∈V​)∈RN×d

A transformer-based predictor Pϕ\mathbf{P}_\phiPϕ​ infers the hidden target representations from the observed context. For a given context-target pair, the predictor receives the context embeddings concatenated with the positional embeddings of the target taxels, each augmented with a learnable mask vector uuu:

Z^(Cj,Tl)=Pϕ(ZcontextCj∥{ei+u}i∈Tl)∈R∣Tl∣×d\hat{Z}^{(\mathcal{C}_j, \mathcal{T}_l)} = \mathbf{P}_\phi \big (Z_{\text{context}}^{\mathcal{C}_j} \parallel \{e_i + u\}_{i \in \mathcal{T}_l} \big ) \in \mathbb{R}^{|\mathcal{T}_l| \times d}Z^(Cj​,Tl​)=Pϕ​(ZcontextCj​​∥{ei​+u}i∈Tl​​)∈R∣Tl​∣×d

Attention within the predictor is computed over the concatenated sequence, retaining only the outputs at the target positions.

Self-Supervised Training Objective The training objective minimizes the mean squared error between the predicted target embeddings and the actual outputs from the target encoder. The loss is averaged over all context and target mask pairs:

LSSL=1ncnt∑j=1nc∑l=1nt1∣Tl∣∑i∈Tl∥z^i(j,l)−sg(zi(l))∥22\mathcal{L}_{\mathrm{SSL}} = \frac{1}{n_c n_t} \sum_{j=1}^{n_c} \sum_{l=1}^{n_t} \frac{1}{|\mathcal{T}_l|} \sum_{i \in \mathcal{T}_l} \left\| \hat{z}_i^{(j, l)} - \mathrm{sg}(z_i^{(l)}) \right\|_2^2LSSL​=nc​nt​1​j=1∑nc​​l=1∑nt​​∣Tl​∣1​i∈Tl​∑​​z^i(j,l)​−sg(zi(l)​)​22​

where sg(⋅)\mathrm{sg}(\cdot)sg(⋅) denotes the stop-gradient operator. Backpropagation updates only the context encoder, while the target encoder parameters are maintained as an exponential moving average of the context encoder weights.

Downstream Adaptation For downstream evaluation, the pre-trained target encoder is frozen and repurposed as the tactile encoder. It maps windowed tactile signals to per-taxel embeddings, which are then fed into a lightweight, task-specific head trained on labeled datasets. This modular design provides flexibility, allowing the same embeddings to support diverse applications such as force estimation, in-hand pose estimation, and policy learning without requiring layout information during inference.

Experiment

Tactile-JEPA is evaluated on three public tactile datasets, Sparshskin, Tactile socks, and DECO-50, which span different sensor types, embodiments, and tasks including force estimation, in-hand pose estimation, object and activity classification, full-body pose regression, and visuo-tactile manipulation policy learning. Against BYOL, MAE, DINO, and end-to-end baselines, the frozen pretrained representations improve in-hand orientation RMSE by 20.8% and force RMSE by 6.3%, while also providing stable tactile-aware policy learning on DECO-50. Ablations validate that the graph-based multi-scale masking strategy with mixed local and global targets yields a robust general-purpose prior, and Tactile-JEPA pretrains stably across all datasets where several competing self-supervised methods collapse.

The compared tactile datasets span heterogeneous sensors, platforms, and collection setups, including magnetic and piezoresistive sensors with unimanual, bipedal, and bimanual configurations. Tactile-JEPA pre-trains stably across all three datasets under a single masking configuration without per-dataset tuning, unlike several baselines that collapse on specific datasets. It is also substantially faster to pre-train than DINO while matching or exceeding downstream performance. The datasets vary widely in sensor type, taxel count, sensing axes, sampling rate, and duration, spanning magnetic and piezoresistive sensors, unimanual, bipedal, and bimanual setups. Tactile-JEPA trains stably on all heterogeneous tactile signals with the same masking configuration, whereas BYOL collapses on two datasets and MAE collapses on one. Multi-scale target masks mixing local and global sampling achieve the best force and in-hand pose accuracy and outperform graph-agnostic masking on all reported tasks. Tactile-JEPA pre-trains 65 to 75 percent faster than DINO across all datasets while matching or exceeding DINO's downstream performance.

Tactile-JEPA learns general-purpose tactile representations that transfer across embodiments and downstream tasks. It matches or exceeds strong baselines on object classification, force estimation, and in-hand pose estimation, with particularly clear gains in overall force and pose error. The method also pre-trains efficiently and remains stable on heterogeneous tactile sensors where some self-supervised baselines collapse. Tactile-JEPA matches the top object classification accuracy and achieves the lowest overall force estimation RMSE among evaluated methods. It produces the best in-hand pose estimates, with lower x and y error than end-to-end and self-supervised baselines. The default multi-scale target mix yields strong force and in-hand accuracy and outperforms graph-agnostic masking across tasks. Tactile-JEPA pre-trains substantially faster than DINO while matching or exceeding DINO downstream. Unlike BYOL and MAE, Tactile-JEPA trains stably across different tactile sensor types.

On DECO-50 policy learning, adding tactile representations reduces RMSE compared with vision-only input. Pretrained tactile encoders provide the largest improvements, with MAE achieving the lowest error and Tactile-JEPA the next lowest among stable methods. Random embeddings and end-to-end tactile encoding yield smaller gains, while BYOL-pretrained and DINO-pretrained encoders collapse. Vision-only policies show the highest RMSE, while non-collapsed tactile representations improve policy learning. The MAE-pretrained tactile encoder achieves the lowest policy error, followed by the Tactile-JEPA encoder. Random embeddings and end-to-end tactile encoders improve over vision-only but less than pretrained encoders. BYOL-pretrained and DINO-pretrained encoders collapse on DECO-50, unlike MAE and Tactile-JEPA.

Across target-mask sampling strategies, the default multi-scale mix of two local and two global masks achieves the strongest force estimation and in-hand pose accuracy while ranking closely behind the best policy learning result. Single-scale global masks perform best for policy learning, and the multi-scale setup outperforms graph-agnostic I-JEPA masking on all three downstream tasks. Replacing global context with local context improves force estimation but substantially reduces pose and policy performance. The default multi-scale target-mask mix yields the best force estimation and in-hand pose accuracy among standard context settings. Single-scale global masks achieve the lowest policy learning error, with the default multi-scale mix ranking a close second. Local context sampling improves force estimation but degrades in-hand pose and policy performance, confirming global context as the stronger general-purpose choice.

The evaluation covers heterogeneous tactile datasets across magnetic and piezoresistive sensors with unimanual, bipedal, and bimanual setups, testing representation quality on object classification, force estimation, in-hand pose estimation, and DECO-50 policy learning. Tactile-JEPA pre-trains stably with one masking configuration while BYOL and MAE collapse on some datasets, and it is substantially faster than DINO while matching or exceeding downstream performance. The default multi-scale target mask mix gives the best force and pose accuracy, whereas single-scale global masks do best for policy learning and global context remains the stronger general-purpose choice. In policy learning, pretrained tactile encoders reduce error over vision-only input, with MAE and Tactile-JEPA being the most effective stable methods.


بناء الذكاء الاصطناعي بالذكاء الاصطناعي

من الفكرة إلى الإطلاق — سرّع تطوير الذكاء الاصطناعي الخاص بك مع المساعدة البرمجية المجانية بالذكاء الاصطناعي، وبيئة جاهزة للاستخدام، وأفضل أسعار لوحدات معالجة الرسومات.

البرمجة التعاونية باستخدام الذكاء الاصطناعي
وحدات GPU جاهزة للعمل
أفضل الأسعار

HyperAI Newsletters

اشترك في آخر تحديثاتنا
سنرسل لك أحدث التحديثات الأسبوعية إلى بريدك الإلكتروني في الساعة التاسعة من صباح كل يوم اثنين
مدعوم بواسطة MailChimp