Enregistré dans:
| Auteur principal: | Hajra, Suvadeep |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2505.15548 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
par: Hajra, Suvadeep, et autres
Publié: (2026)
par: Hajra, Suvadeep, et autres
Publié: (2026)
Decomposable Transformer Point Processes
par: Panos, Aristeidis
Publié: (2024)
par: Panos, Aristeidis
Publié: (2024)
Transparency in Sleep Staging: Deep Learning Method for EEG Sleep Stage Classification with Model Interpretability
par: Sharma, Shivam, et autres
Publié: (2023)
par: Sharma, Shivam, et autres
Publié: (2023)
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
par: Kang, Bong Gyun, et autres
Publié: (2024)
par: Kang, Bong Gyun, et autres
Publié: (2024)
Decomposing Attention To Find Context-Sensitive Neurons
par: Gibson, Alex
Publié: (2025)
par: Gibson, Alex
Publié: (2025)
Short-Range Oversquashing
par: Mishayev, Yaaqov, et autres
Publié: (2025)
par: Mishayev, Yaaqov, et autres
Publié: (2025)
Hybrid Focal and Full-Range Attention Based Graph Transformers
par: Zhu, Minhong, et autres
Publié: (2023)
par: Zhu, Minhong, et autres
Publié: (2023)
Decomposing Global Feature Effects Based on Feature Interactions
par: Herbinger, Julia, et autres
Publié: (2023)
par: Herbinger, Julia, et autres
Publié: (2023)
The Effect of Attention Head Count on Transformer Approximation
par: Yu, Penghao, et autres
Publié: (2025)
par: Yu, Penghao, et autres
Publié: (2025)
AI Generalisation Gap In Comorbid Sleep Disorder Staging
par: Bose, Saswata, et autres
Publié: (2026)
par: Bose, Saswata, et autres
Publié: (2026)
Triple Attention Transformer Architecture for Time-Dependent Concrete Creep Prediction
par: Dokduea, Warayut, et autres
Publié: (2025)
par: Dokduea, Warayut, et autres
Publié: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
par: Lou, Chao, et autres
Publié: (2024)
par: Lou, Chao, et autres
Publié: (2024)
DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects
par: Tamano, Shu
Publié: (2025)
par: Tamano, Shu
Publié: (2025)
Integration of Mamba and Transformer -- MAT for Long-Short Range Time Series Forecasting with Application to Weather Dynamics
par: Zhang, Wenqing, et autres
Publié: (2024)
par: Zhang, Wenqing, et autres
Publié: (2024)
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
par: Bai, Xiaoyan, et autres
Publié: (2025)
par: Bai, Xiaoyan, et autres
Publié: (2025)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
par: Liao, Huanxuan, et autres
Publié: (2025)
par: Liao, Huanxuan, et autres
Publié: (2025)
Extracting Cause-Effect Pairs from a Sentence with a Dependency-Aware Transformer Model
par: Kabir, Md Ahsanul, et autres
Publié: (2025)
par: Kabir, Md Ahsanul, et autres
Publié: (2025)
Decomposable Neuro Symbolic Regression
par: Morales, Giorgio, et autres
Publié: (2025)
par: Morales, Giorgio, et autres
Publié: (2025)
Learning Long-Range Dependencies with Temporal Predictive Coding
par: Potter, Tom, et autres
Publié: (2026)
par: Potter, Tom, et autres
Publié: (2026)
Decomposing Gaussians with Unknown Covariance
par: Dharamshi, Ameer, et autres
Publié: (2024)
par: Dharamshi, Ameer, et autres
Publié: (2024)
Decomposing Prediction Mechanisms for In-Context Recall
par: Daniels, Sultan, et autres
Publié: (2025)
par: Daniels, Sultan, et autres
Publié: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
par: Engels, Joshua, et autres
Publié: (2024)
par: Engels, Joshua, et autres
Publié: (2024)
Decomposing the Depth Profile of Fine-Tuning
par: Billa, Jayadev
Publié: (2026)
par: Billa, Jayadev
Publié: (2026)
Learning Long Range Dependencies on Graphs via Random Walks
par: Chen, Dexiong, et autres
Publié: (2024)
par: Chen, Dexiong, et autres
Publié: (2024)
On the Effect of Instability on Learning Continuous-Time Linear Control Systems
par: Hafshejani, Reza Sadeghi, et autres
Publié: (2024)
par: Hafshejani, Reza Sadeghi, et autres
Publié: (2024)
Batch Normalization Decomposed
par: Nachum, Ido, et autres
Publié: (2024)
par: Nachum, Ido, et autres
Publié: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
par: Wen, Kaiyue, et autres
Publié: (2024)
par: Wen, Kaiyue, et autres
Publié: (2024)
On the Mirage of Long-Range Dependency, with an Application to Integer Multiplication
par: Wei, Zichao
Publié: (2026)
par: Wei, Zichao
Publié: (2026)
Leveraging Discrete Function Decomposability for Scientific Design
par: Bowden, James C., et autres
Publié: (2025)
par: Bowden, James C., et autres
Publié: (2025)
Decomposing Task Vectors for Refined Model Editing
par: Damirchi, Hamed, et autres
Publié: (2025)
par: Damirchi, Hamed, et autres
Publié: (2025)
DEMAU: Decompose, Explore, Model and Analyse Uncertainties
par: Hoarau, Arthur, et autres
Publié: (2024)
par: Hoarau, Arthur, et autres
Publié: (2024)
ParallelTime: Dynamically Weighting the Balance of Short- and Long-Term Temporal Dependencies
par: Katav, Itay, et autres
Publié: (2025)
par: Katav, Itay, et autres
Publié: (2025)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
par: Kim, Bum Jun, et autres
Publié: (2026)
par: Kim, Bum Jun, et autres
Publié: (2026)
Transformer Reconstructed with Dynamic Value Attention
par: Wang, Xiaowei
Publié: (2025)
par: Wang, Xiaowei
Publié: (2025)
Cottention: Linear Transformers With Cosine Attention
par: Mongaras, Gabriel, et autres
Publié: (2024)
par: Mongaras, Gabriel, et autres
Publié: (2024)
Transformers with Sparse Attention for Granger Causality
par: Mahesh, Riya, et autres
Publié: (2024)
par: Mahesh, Riya, et autres
Publié: (2024)
Preconditioned Attention: Enhancing Efficiency in Transformers
par: Saratchandran, Hemanth
Publié: (2026)
par: Saratchandran, Hemanth
Publié: (2026)
Graph External Attention Enhanced Transformer
par: Liang, Jianqing, et autres
Publié: (2024)
par: Liang, Jianqing, et autres
Publié: (2024)
Accelerating the Low-Rank Decomposed Models
par: Hajimolahoseini, Habib, et autres
Publié: (2024)
par: Hajimolahoseini, Habib, et autres
Publié: (2024)
Geometric Learning with Positively Decomposable Kernels
par: Da Costa, Nathael, et autres
Publié: (2023)
par: Da Costa, Nathael, et autres
Publié: (2023)
Documents similaires
-
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
par: Hajra, Suvadeep, et autres
Publié: (2026) -
Decomposable Transformer Point Processes
par: Panos, Aristeidis
Publié: (2024) -
Transparency in Sleep Staging: Deep Learning Method for EEG Sleep Stage Classification with Model Interpretability
par: Sharma, Shivam, et autres
Publié: (2023) -
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
par: Kang, Bong Gyun, et autres
Publié: (2024) -
Decomposing Attention To Find Context-Sensitive Neurons
par: Gibson, Alex
Publié: (2025)