Salvato in:
| Autori principali: | Shaikh, Khalid, Singh, Asmit Kumar, Dsouza, Rebecca Christopher, Shiromani, Shikhar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.13314 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
di: Mann, Logan, et al.
Pubblicazione: (2026)
di: Mann, Logan, et al.
Pubblicazione: (2026)
Focus Where It Matters: Graph Selective State Focused Attention Networks
di: Vashistha, Shikhar, et al.
Pubblicazione: (2024)
di: Vashistha, Shikhar, et al.
Pubblicazione: (2024)
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024)
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024)
MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
di: Bandyopadhyay, Asmit, et al.
Pubblicazione: (2025)
di: Bandyopadhyay, Asmit, et al.
Pubblicazione: (2025)
Improving Multimodal Large Language Models Using Continual Learning
di: Srivastava, Shikhar, et al.
Pubblicazione: (2024)
di: Srivastava, Shikhar, et al.
Pubblicazione: (2024)
CLAReSNet: When Convolution Meets Latent Attention for Hyperspectral Image Classification
di: Bandyopadhyay, Asmit, et al.
Pubblicazione: (2025)
di: Bandyopadhyay, Asmit, et al.
Pubblicazione: (2025)
Execution-Grounded Credit Assignment for GRPO in Code Generation
di: Kumar, Abhijit, et al.
Pubblicazione: (2026)
di: Kumar, Abhijit, et al.
Pubblicazione: (2026)
On the Role of Attention Heads in Large Language Model Safety
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
di: Zhou, Zhenhong, et al.
Pubblicazione: (2024)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
Superiority of Multi-Head Attention in In-Context Linear Regression
di: Cui, Yingqian, et al.
Pubblicazione: (2024)
di: Cui, Yingqian, et al.
Pubblicazione: (2024)
Attention Head Entropy of LLMs Predicts Answer Correctness
di: Ostmeier, Sophie, et al.
Pubblicazione: (2026)
di: Ostmeier, Sophie, et al.
Pubblicazione: (2026)
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
di: Shiromani, Shikhar, et al.
Pubblicazione: (2026)
di: Shiromani, Shikhar, et al.
Pubblicazione: (2026)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
Routoo: Learning to Route to Large Language Models Effectively
di: Mohammadshahi, Alireza, et al.
Pubblicazione: (2024)
di: Mohammadshahi, Alireza, et al.
Pubblicazione: (2024)
Fourier Head: Helping Large Language Models Learn Complex Probability Distributions
di: Gillman, Nate, et al.
Pubblicazione: (2024)
di: Gillman, Nate, et al.
Pubblicazione: (2024)
Interleaved Head Attention
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
di: Duvvuri, Sai Surya, et al.
Pubblicazione: (2026)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
di: Kim, Jinyeong, et al.
Pubblicazione: (2025)
di: Kim, Jinyeong, et al.
Pubblicazione: (2025)
Robust and Noise-resilient Long-Term Prediction of Spatiotemporal Data Using Variational Mode Graph Neural Networks with 3D Attention
di: Ahmad, Osama, et al.
Pubblicazione: (2025)
di: Ahmad, Osama, et al.
Pubblicazione: (2025)
Dynamic Rank Reinforcement Learning for Adaptive Low-Rank Multi-Head Self Attention in Large Language Models
di: Erden, Caner
Pubblicazione: (2025)
di: Erden, Caner
Pubblicazione: (2025)
Adaptive Head Budgeting for Efficient Multi-Head Attention
di: Faye, Bilal, et al.
Pubblicazione: (2026)
di: Faye, Bilal, et al.
Pubblicazione: (2026)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
di: He, Jianliang, et al.
Pubblicazione: (2025)
di: He, Jianliang, et al.
Pubblicazione: (2025)
Training Dynamics of In-Context Learning in Linear Attention
di: Zhang, Yedi, et al.
Pubblicazione: (2025)
di: Zhang, Yedi, et al.
Pubblicazione: (2025)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
di: Pandey, Vishal, et al.
Pubblicazione: (2026)
di: Pandey, Vishal, et al.
Pubblicazione: (2026)
Prediction of Herd Life in Dairy Cows Using Multi-Head Attention Transformers
di: Saki, Mahdi, et al.
Pubblicazione: (2025)
di: Saki, Mahdi, et al.
Pubblicazione: (2025)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
Label-free Anomaly Detection in Aerial Agricultural Images with Masked Image Modeling
di: Shikhar, Sambal, et al.
Pubblicazione: (2024)
di: Shikhar, Sambal, et al.
Pubblicazione: (2024)
DEI: Diversity in Evolutionary Inference for Quality-Diversity Search
di: Donaghy, John, et al.
Pubblicazione: (2026)
di: Donaghy, John, et al.
Pubblicazione: (2026)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
di: Ramesh, Guruprasad Viswanathan, et al.
Pubblicazione: (2026)
di: Ramesh, Guruprasad Viswanathan, et al.
Pubblicazione: (2026)
AlpaPICO: Extraction of PICO Frames from Clinical Trial Documents Using LLMs
di: Ghosh, Madhusudan, et al.
Pubblicazione: (2024)
di: Ghosh, Madhusudan, et al.
Pubblicazione: (2024)
Can Large Language Models Transform Computational Social Science?
di: Ziems, Caleb, et al.
Pubblicazione: (2023)
di: Ziems, Caleb, et al.
Pubblicazione: (2023)
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
di: Smith, James Seale, et al.
Pubblicazione: (2025)
di: Smith, James Seale, et al.
Pubblicazione: (2025)
Parallax: Parameterized Local Linear Attention for Language Modeling
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
di: Zuo, Yifei, et al.
Pubblicazione: (2026)
Multi-Head Low-Rank Attention
di: Liu, Songtao, et al.
Pubblicazione: (2026)
di: Liu, Songtao, et al.
Pubblicazione: (2026)
Benign Overfitting in Single-Head Attention
di: Magen, Roey, et al.
Pubblicazione: (2024)
di: Magen, Roey, et al.
Pubblicazione: (2024)
Large Language and Reasoning Models are Shallow Disjunctive Reasoners
di: Khalid, Irtaza, et al.
Pubblicazione: (2025)
di: Khalid, Irtaza, et al.
Pubblicazione: (2025)
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
di: Zheng, Ziwei, et al.
Pubblicazione: (2025)
di: Zheng, Ziwei, et al.
Pubblicazione: (2025)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
Large Language Models aren't all that you need
di: Holla, Kiran Voderhobli, et al.
Pubblicazione: (2024)
di: Holla, Kiran Voderhobli, et al.
Pubblicazione: (2024)
LoLCATs: On Low-Rank Linearizing of Large Language Models
di: Zhang, Michael, et al.
Pubblicazione: (2024)
di: Zhang, Michael, et al.
Pubblicazione: (2024)
Synthetic Tabular Data Generation for Imbalanced Classification: The Surprising Effectiveness of an Overlap Class
di: D'souza, Annie, et al.
Pubblicazione: (2024)
di: D'souza, Annie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
di: Mann, Logan, et al.
Pubblicazione: (2026) -
Focus Where It Matters: Graph Selective State Focused Attention Networks
di: Vashistha, Shikhar, et al.
Pubblicazione: (2024) -
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
di: Yu, Zhongzhi, et al.
Pubblicazione: (2024) -
MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
di: Bandyopadhyay, Asmit, et al.
Pubblicazione: (2025) -
Improving Multimodal Large Language Models Using Continual Learning
di: Srivastava, Shikhar, et al.
Pubblicazione: (2024)