LUCID: Attention with Preconditioned Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Duvvuri, Sai Surya, Patel, Nirmal, Gupta, Nilesh, Dhillon, Inderjit S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LASER: Attention with Exponential Transformation
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024)
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
von: Patel, Nirmal, et al.
Veröffentlicht: (2026)
Interleaved Head Attention
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026)
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026)
Towards Quantifying the Preconditioning Effect of Adam
von: Das, Rudrajit, et al.
Veröffentlicht: (2024)
von: Das, Rudrajit, et al.
Veröffentlicht: (2024)
The Art of Scaling Reinforcement Learning Compute for LLMs
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)
Dual-Encoders for Extreme Multi-Label Classification
von: Gupta, Nilesh, et al.
Veröffentlicht: (2023)
von: Gupta, Nilesh, et al.
Veröffentlicht: (2023)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
von: Yen, Jui-Nan, et al.
Veröffentlicht: (2024)
von: Yen, Jui-Nan, et al.
Veröffentlicht: (2024)
LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
von: Gupta, Nilesh, et al.
Veröffentlicht: (2025)
von: Gupta, Nilesh, et al.
Veröffentlicht: (2025)
EHI: End-to-end Learning of Hierarchical Index for Efficient Dense Retrieval
von: Kumar, Ramnath, et al.
Veröffentlicht: (2023)
von: Kumar, Ramnath, et al.
Veröffentlicht: (2023)
Scalable In-context Ranking with Generative Models
von: Gupta, Nilesh, et al.
Veröffentlicht: (2025)
von: Gupta, Nilesh, et al.
Veröffentlicht: (2025)
Geometric Median (GM) Matching for Robust Data Pruning
von: Acharya, Anish, et al.
Veröffentlicht: (2024)
von: Acharya, Anish, et al.
Veröffentlicht: (2024)
Fast and Simplex: 2-Simplicial Attention in Triton
von: Roy, Aurko, et al.
Veröffentlicht: (2025)
von: Roy, Aurko, et al.
Veröffentlicht: (2025)
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
von: Acharya, Anish, et al.
Veröffentlicht: (2025)
von: Acharya, Anish, et al.
Veröffentlicht: (2025)
Preconditioned Attention: Enhancing Efficiency in Transformers
von: Saratchandran, Hemanth
Veröffentlicht: (2026)
von: Saratchandran, Hemanth
Veröffentlicht: (2026)
Understanding Contrastive Representation Learning from Positive Unlabeled (PU) Data
von: Acharya, Anish, et al.
Veröffentlicht: (2024)
von: Acharya, Anish, et al.
Veröffentlicht: (2024)
LUCID: Learning-Enabled Uncertainty-Aware Certification of Stochastic Dynamical Systems
von: Casablanca, Ernesto, et al.
Veröffentlicht: (2025)
von: Casablanca, Ernesto, et al.
Veröffentlicht: (2025)
Compressing Many-Shots in In-Context Learning
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)
Positive Unlabeled Contrastive Learning
von: Acharya, Anish, et al.
Veröffentlicht: (2022)
von: Acharya, Anish, et al.
Veröffentlicht: (2022)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
von: Goel, Gautam, et al.
Veröffentlicht: (2026)
von: Goel, Gautam, et al.
Veröffentlicht: (2026)
Retraining with Predicted Hard Labels Provably Increases Model Accuracy
von: Das, Rudrajit, et al.
Veröffentlicht: (2024)
von: Das, Rudrajit, et al.
Veröffentlicht: (2024)
Attention Meets UAVs: A Comprehensive Evaluation of DDoS Detection in Low-Cost UAVs
von: Sharma, Ashish, et al.
Veröffentlicht: (2024)
von: Sharma, Ashish, et al.
Veröffentlicht: (2024)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
von: Zhou, Chenyu, et al.
Veröffentlicht: (2026)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2026)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
von: Wang, Yihan, et al.
Veröffentlicht: (2022)
von: Wang, Yihan, et al.
Veröffentlicht: (2022)
Universal Sequence Preconditioning
von: Marsden, Annie, et al.
Veröffentlicht: (2025)
von: Marsden, Annie, et al.
Veröffentlicht: (2025)
Multi-Head Attention Is a Multi-Player Game
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2026)
Exploring Design Choices for Building Language-Specific LLMs
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
von: Tejaswi, Atula, et al.
Veröffentlicht: (2024)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
von: Bansal, Rachit, et al.
Veröffentlicht: (2025)
von: Bansal, Rachit, et al.
Veröffentlicht: (2025)
Open-TQ-Metal: Fused Compressed-Domain Attention for Long-Context LLM Inference on Apple Silicon
von: Vegasena, Sai
Veröffentlicht: (2026)
von: Vegasena, Sai
Veröffentlicht: (2026)
On the Nystrom Approximation for Preconditioning in Kernel Machines
von: Abedsoltan, Amirhesam, et al.
Veröffentlicht: (2023)
von: Abedsoltan, Amirhesam, et al.
Veröffentlicht: (2023)
Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
Large Language Models are Interpretable Learners
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
von: Wang, Ruochen, et al.
Veröffentlicht: (2024)
Multi-Knowledge Fusion Network for Time Series Representation Learning
von: Sakhinana, Sagar Srinivas, et al.
Veröffentlicht: (2024)
von: Sakhinana, Sagar Srinivas, et al.
Veröffentlicht: (2024)
A Representation-Consistent Gated Recurrent Framework for Robust Medical Time-Series Classification
von: Sai, Maitri Krishna
Veröffentlicht: (2026)
von: Sai, Maitri Krishna
Veröffentlicht: (2026)
The Power of Second Order Methods for Sequence Preconditioning
von: Marsden, Annie, et al.
Veröffentlicht: (2026)
von: Marsden, Annie, et al.
Veröffentlicht: (2026)
Preconditioned Inexact Stochastic ADMM for Deep Model
von: Zhou, Shenglong, et al.
Veröffentlicht: (2025)
von: Zhou, Shenglong, et al.
Veröffentlicht: (2025)
Multi-Source Knowledge-Based Hybrid Neural Framework for Time Series Representation Learning
von: Sakhinana, Sagar Srinivas, et al.
Veröffentlicht: (2024)
von: Sakhinana, Sagar Srinivas, et al.
Veröffentlicht: (2024)
Matryoshka Model Learning for Improved Elastic Student Models
von: Verma, Chetan, et al.
Veröffentlicht: (2025)
von: Verma, Chetan, et al.
Veröffentlicht: (2025)
Are Anxiety Detection Models Generalizable? A Cross-Activity and Cross-Population Study Using Wearables
von: Sahu, Nilesh Kumar, et al.
Veröffentlicht: (2025)
von: Sahu, Nilesh Kumar, et al.
Veröffentlicht: (2025)
Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime
von: Ghane, Reza, et al.
Veröffentlicht: (2026)
von: Ghane, Reza, et al.
Veröffentlicht: (2026)
Gradient Preconditioning for Efficient and Reliable Reward-Guided Generation
von: Hwang, Jisung, et al.
Veröffentlicht: (2026)
von: Hwang, Jisung, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LASER: Attention with Exponential Transformation
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2024) -
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
von: Patel, Nirmal, et al.
Veröffentlicht: (2026) -
Interleaved Head Attention
von: Duvvuri, Sai Surya, et al.
Veröffentlicht: (2026) -
Towards Quantifying the Preconditioning Effect of Adam
von: Das, Rudrajit, et al.
Veröffentlicht: (2024) -
The Art of Scaling Reinforcement Learning Compute for LLMs
von: Khatri, Devvrit, et al.
Veröffentlicht: (2025)