Supervised learning pays attention
Fuente:
arXiv
Saved in:
| Main Authors: | Craig, Erin, Tibshirani, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An end-to-end attention-based approach for learning on graphs
by: Buterez, David, et al.
Published: (2024)
by: Buterez, David, et al.
Published: (2024)
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024)
by: Gros, Claudius
Published: (2024)
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
DEDUCE: Multi-head attention decoupled contrastive learning to discover cancer subtypes based on multi-omics data
by: Pan, Liangrui, et al.
Published: (2023)
by: Pan, Liangrui, et al.
Published: (2023)
Multiview graph dual-attention deep learning and contrastive learning for multi-criteria recommender systems
by: Forouzandeh, Saman, et al.
Published: (2025)
by: Forouzandeh, Saman, et al.
Published: (2025)
On student-teacher deviations in distillation: does it pay to disobey?
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
by: Nagarajan, Vaishnavh, et al.
Published: (2023)
Fine-grained graph representation learning for heterogeneous mobile networks with attentive fusion and contrastive learning
by: Liu, Shengheng, et al.
Published: (2024)
by: Liu, Shengheng, et al.
Published: (2024)
Short window attention enables long-term memorization
by: Cabannes, Loïc, et al.
Published: (2025)
by: Cabannes, Loïc, et al.
Published: (2025)
Tucker Attention: A generalization of approximate attention mechanisms
by: Klein, Timon, et al.
Published: (2026)
by: Klein, Timon, et al.
Published: (2026)
CoSupFormer : A Contrastive Supervised learning approach for EEG signal Classification
by: Darankoum, D., et al.
Published: (2025)
by: Darankoum, D., et al.
Published: (2025)
Cross-attentive Cohesive Subgraph Embedding to Mitigate Oversquashing in GNNs
by: Hossain, Tanvir, et al.
Published: (2026)
by: Hossain, Tanvir, et al.
Published: (2026)
CST-AFNet: A dual attention-based deep learning framework for intrusion detection in IoT networks
by: Ishtiaq, Waqas, et al.
Published: (2025)
by: Ishtiaq, Waqas, et al.
Published: (2025)
GRC-Net: Gram Residual Co-attention Net for epilepsy prediction
by: You, Bihao, et al.
Published: (2025)
by: You, Bihao, et al.
Published: (2025)
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
by: Súkeník, Peter, et al.
Published: (2026)
by: Súkeník, Peter, et al.
Published: (2026)
GQA-μP: The maximal parameterization update for grouped query attention
by: Chickering, Kyle R., et al.
Published: (2026)
by: Chickering, Kyle R., et al.
Published: (2026)
Static and multivariate-temporal attentive fusion transformer for readmission risk prediction
by: Sun, Zhe, et al.
Published: (2024)
by: Sun, Zhe, et al.
Published: (2024)
Symbolic Autoencoding for Self-Supervised Sequence Learning
by: Amani, Mohammad Hossein, et al.
Published: (2024)
by: Amani, Mohammad Hossein, et al.
Published: (2024)
Three-dimensional attention Transformer for state evaluation in real-time strategy games
by: Ye, Yanqing, et al.
Published: (2025)
by: Ye, Yanqing, et al.
Published: (2025)
A foundation model with multi-variate parallel attention to generate neuronal activity
by: Carzaniga, Francesco, et al.
Published: (2025)
by: Carzaniga, Francesco, et al.
Published: (2025)
Fusion of Multiscale Features Via Centralized Sparse-attention Network for EEG Decoding
by: Cai, Xiangrui, et al.
Published: (2025)
by: Cai, Xiangrui, et al.
Published: (2025)
ACCORD: Autoregressive Constraint-satisfying Generation for COmbinatorial Optimization with Routing and Dynamic attention
by: Abgaryan, Henrik, et al.
Published: (2025)
by: Abgaryan, Henrik, et al.
Published: (2025)
Filter then Attend: Improving attention-based Time Series Forecasting with Spectral Filtering
by: Dayag, Elisha, et al.
Published: (2025)
by: Dayag, Elisha, et al.
Published: (2025)
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
S-JEPA: towards seamless cross-dataset transfer through dynamic spatial attention
by: Guetschel, Pierre, et al.
Published: (2024)
by: Guetschel, Pierre, et al.
Published: (2024)
CNN-TFT explained by SHAP with multi-head attention weights for time series forecasting
by: Stefenon, Stefano F., et al.
Published: (2025)
by: Stefenon, Stefano F., et al.
Published: (2025)
Graph-attention-based Casual Discovery with Trust Region-navigated Clipping Policy Optimization
by: Liu, Shixuan, et al.
Published: (2024)
by: Liu, Shixuan, et al.
Published: (2024)
Transformer autoencoder with local attention for sparse and irregular time series with application on risk estimation
by: Rodis, Panteleimon
Published: (2026)
by: Rodis, Panteleimon
Published: (2026)
Optimal message passing for molecular prediction is simple, attentive and spatial
by: Castaneda-Leautaud, Alma C., et al.
Published: (2025)
by: Castaneda-Leautaud, Alma C., et al.
Published: (2025)
Physics-integrated generative modeling using attentive planar normalizing flow based variational autoencoder
by: Akhtar, Sheikh Waqas
Published: (2024)
by: Akhtar, Sheikh Waqas
Published: (2024)
MoXGATE: Modality-aware cross-attention for multi-omic gastrointestinal cancer sub-type classification
by: Dip, Sajib Acharjee, et al.
Published: (2025)
by: Dip, Sajib Acharjee, et al.
Published: (2025)
Emergent Language Symbolic Autoencoder (ELSA) with Weak Supervision to Model Hierarchical Brain Networks
by: Latheef, Ammar Ahmed Pallikonda, et al.
Published: (2024)
by: Latheef, Ammar Ahmed Pallikonda, et al.
Published: (2024)
Sparsely Supervised Diffusion
by: Zhao, Wenshuai, et al.
Published: (2026)
by: Zhao, Wenshuai, et al.
Published: (2026)
A standard transformer and attention with linear biases for molecular conformer generation
by: Gurev, Viatcheslav, et al.
Published: (2025)
by: Gurev, Viatcheslav, et al.
Published: (2025)
TransformerFAM: Feedback attention is working memory
by: Hwang, Dongseong, et al.
Published: (2024)
by: Hwang, Dongseong, et al.
Published: (2024)
On the Universality of Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2024)
by: Qiang, Wenwen, et al.
Published: (2024)
Bitformer: An efficient Transformer with bitwise operation-based attention for Big Data Analytics at low-cost low-precision devices
by: Duan, Gaoxiang, et al.
Published: (2023)
by: Duan, Gaoxiang, et al.
Published: (2023)
LLM-ABBA: Understanding time series via symbolic approximation
by: Chen, Xinye, et al.
Published: (2024)
by: Chen, Xinye, et al.
Published: (2024)
Sinkhorn doubly stochastic attention rank decay analysis
by: Lapenna, Michela, et al.
Published: (2026)
by: Lapenna, Michela, et al.
Published: (2026)
Physics-informed GNN for medium-high voltage AC power flow with edge-aware attention and line search correction operator
by: Kim, Changhun, et al.
Published: (2025)
by: Kim, Changhun, et al.
Published: (2025)
Clustering Properties of Self-Supervised Learning
by: Weng, Xi, et al.
Published: (2025)
by: Weng, Xi, et al.
Published: (2025)
Similar Items
-
An end-to-end attention-based approach for learning on graphs
by: Buterez, David, et al.
Published: (2024) -
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024) -
Poly-attention: a general scheme for higher-order self-attention
by: Chakrabarti, Sayak, et al.
Published: (2026) -
DEDUCE: Multi-head attention decoupled contrastive learning to discover cancer subtypes based on multi-omics data
by: Pan, Liangrui, et al.
Published: (2023) -
Multiview graph dual-attention deep learning and contrastive learning for multi-criteria recommender systems
by: Forouzandeh, Saman, et al.
Published: (2025)