Guardado en:
| Autor principal: | Hajra, Suvadeep |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2505.15548 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
por: Hajra, Suvadeep, et al.
Publicado: (2026)
por: Hajra, Suvadeep, et al.
Publicado: (2026)
Decomposable Transformer Point Processes
por: Panos, Aristeidis
Publicado: (2024)
por: Panos, Aristeidis
Publicado: (2024)
Transparency in Sleep Staging: Deep Learning Method for EEG Sleep Stage Classification with Model Interpretability
por: Sharma, Shivam, et al.
Publicado: (2023)
por: Sharma, Shivam, et al.
Publicado: (2023)
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
por: Kang, Bong Gyun, et al.
Publicado: (2024)
por: Kang, Bong Gyun, et al.
Publicado: (2024)
Decomposing Attention To Find Context-Sensitive Neurons
por: Gibson, Alex
Publicado: (2025)
por: Gibson, Alex
Publicado: (2025)
Short-Range Oversquashing
por: Mishayev, Yaaqov, et al.
Publicado: (2025)
por: Mishayev, Yaaqov, et al.
Publicado: (2025)
Hybrid Focal and Full-Range Attention Based Graph Transformers
por: Zhu, Minhong, et al.
Publicado: (2023)
por: Zhu, Minhong, et al.
Publicado: (2023)
Decomposing Global Feature Effects Based on Feature Interactions
por: Herbinger, Julia, et al.
Publicado: (2023)
por: Herbinger, Julia, et al.
Publicado: (2023)
The Effect of Attention Head Count on Transformer Approximation
por: Yu, Penghao, et al.
Publicado: (2025)
por: Yu, Penghao, et al.
Publicado: (2025)
AI Generalisation Gap In Comorbid Sleep Disorder Staging
por: Bose, Saswata, et al.
Publicado: (2026)
por: Bose, Saswata, et al.
Publicado: (2026)
Triple Attention Transformer Architecture for Time-Dependent Concrete Creep Prediction
por: Dokduea, Warayut, et al.
Publicado: (2025)
por: Dokduea, Warayut, et al.
Publicado: (2025)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
por: Lou, Chao, et al.
Publicado: (2024)
por: Lou, Chao, et al.
Publicado: (2024)
DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects
por: Tamano, Shu
Publicado: (2025)
por: Tamano, Shu
Publicado: (2025)
Integration of Mamba and Transformer -- MAT for Long-Short Range Time Series Forecasting with Application to Weather Dynamics
por: Zhang, Wenqing, et al.
Publicado: (2024)
por: Zhang, Wenqing, et al.
Publicado: (2024)
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
por: Bai, Xiaoyan, et al.
Publicado: (2025)
por: Bai, Xiaoyan, et al.
Publicado: (2025)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
por: Liao, Huanxuan, et al.
Publicado: (2025)
por: Liao, Huanxuan, et al.
Publicado: (2025)
Extracting Cause-Effect Pairs from a Sentence with a Dependency-Aware Transformer Model
por: Kabir, Md Ahsanul, et al.
Publicado: (2025)
por: Kabir, Md Ahsanul, et al.
Publicado: (2025)
Decomposable Neuro Symbolic Regression
por: Morales, Giorgio, et al.
Publicado: (2025)
por: Morales, Giorgio, et al.
Publicado: (2025)
Learning Long-Range Dependencies with Temporal Predictive Coding
por: Potter, Tom, et al.
Publicado: (2026)
por: Potter, Tom, et al.
Publicado: (2026)
Decomposing Gaussians with Unknown Covariance
por: Dharamshi, Ameer, et al.
Publicado: (2024)
por: Dharamshi, Ameer, et al.
Publicado: (2024)
Decomposing Prediction Mechanisms for In-Context Recall
por: Daniels, Sultan, et al.
Publicado: (2025)
por: Daniels, Sultan, et al.
Publicado: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
por: Engels, Joshua, et al.
Publicado: (2024)
por: Engels, Joshua, et al.
Publicado: (2024)
Decomposing the Depth Profile of Fine-Tuning
por: Billa, Jayadev
Publicado: (2026)
por: Billa, Jayadev
Publicado: (2026)
Learning Long Range Dependencies on Graphs via Random Walks
por: Chen, Dexiong, et al.
Publicado: (2024)
por: Chen, Dexiong, et al.
Publicado: (2024)
On the Effect of Instability on Learning Continuous-Time Linear Control Systems
por: Hafshejani, Reza Sadeghi, et al.
Publicado: (2024)
por: Hafshejani, Reza Sadeghi, et al.
Publicado: (2024)
Batch Normalization Decomposed
por: Nachum, Ido, et al.
Publicado: (2024)
por: Nachum, Ido, et al.
Publicado: (2024)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
por: Wen, Kaiyue, et al.
Publicado: (2024)
por: Wen, Kaiyue, et al.
Publicado: (2024)
On the Mirage of Long-Range Dependency, with an Application to Integer Multiplication
por: Wei, Zichao
Publicado: (2026)
por: Wei, Zichao
Publicado: (2026)
Leveraging Discrete Function Decomposability for Scientific Design
por: Bowden, James C., et al.
Publicado: (2025)
por: Bowden, James C., et al.
Publicado: (2025)
Decomposing Task Vectors for Refined Model Editing
por: Damirchi, Hamed, et al.
Publicado: (2025)
por: Damirchi, Hamed, et al.
Publicado: (2025)
DEMAU: Decompose, Explore, Model and Analyse Uncertainties
por: Hoarau, Arthur, et al.
Publicado: (2024)
por: Hoarau, Arthur, et al.
Publicado: (2024)
ParallelTime: Dynamically Weighting the Balance of Short- and Long-Term Temporal Dependencies
por: Katav, Itay, et al.
Publicado: (2025)
por: Katav, Itay, et al.
Publicado: (2025)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
por: Kim, Bum Jun, et al.
Publicado: (2026)
por: Kim, Bum Jun, et al.
Publicado: (2026)
Transformer Reconstructed with Dynamic Value Attention
por: Wang, Xiaowei
Publicado: (2025)
por: Wang, Xiaowei
Publicado: (2025)
Cottention: Linear Transformers With Cosine Attention
por: Mongaras, Gabriel, et al.
Publicado: (2024)
por: Mongaras, Gabriel, et al.
Publicado: (2024)
Transformers with Sparse Attention for Granger Causality
por: Mahesh, Riya, et al.
Publicado: (2024)
por: Mahesh, Riya, et al.
Publicado: (2024)
Preconditioned Attention: Enhancing Efficiency in Transformers
por: Saratchandran, Hemanth
Publicado: (2026)
por: Saratchandran, Hemanth
Publicado: (2026)
Graph External Attention Enhanced Transformer
por: Liang, Jianqing, et al.
Publicado: (2024)
por: Liang, Jianqing, et al.
Publicado: (2024)
Accelerating the Low-Rank Decomposed Models
por: Hajimolahoseini, Habib, et al.
Publicado: (2024)
por: Hajimolahoseini, Habib, et al.
Publicado: (2024)
Geometric Learning with Positively Decomposable Kernels
por: Da Costa, Nathael, et al.
Publicado: (2023)
por: Da Costa, Nathael, et al.
Publicado: (2023)
Ejemplares similares
-
Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
por: Hajra, Suvadeep, et al.
Publicado: (2026) -
Decomposable Transformer Point Processes
por: Panos, Aristeidis
Publicado: (2024) -
Transparency in Sleep Staging: Deep Learning Method for EEG Sleep Stage Classification with Model Interpretability
por: Sharma, Shivam, et al.
Publicado: (2023) -
Introducing Spectral Attention for Long-Range Dependency in Time Series Forecasting
por: Kang, Bong Gyun, et al.
Publicado: (2024) -
Decomposing Attention To Find Context-Sensitive Neurons
por: Gibson, Alex
Publicado: (2025)