Interleaved Head Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Duvvuri, Sai Surya, Ekbote, Chanakya, Bansal, Rachit, Tiwari, Rishabh, Khatri, Devvrit, Brandfonbrener, David, Liang, Paul, Dhillon, Inderjit, Zaheer, Manzil |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Art of Scaling Reinforcement Learning Compute for LLMs
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024)
by: Duvvuri, Sai Surya, et al.
Published: (2024)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
by: Bansal, Rachit, et al.
Published: (2025)
by: Bansal, Rachit, et al.
Published: (2025)
LUCID: Attention with Preconditioned Representations
by: Duvvuri, Sai Surya, et al.
Published: (2026)
by: Duvvuri, Sai Surya, et al.
Published: (2026)
Dual-Encoders for Extreme Multi-Label Classification
by: Gupta, Nilesh, et al.
Published: (2023)
by: Gupta, Nilesh, et al.
Published: (2023)
Fast and Simplex: 2-Simplicial Attention in Triton
by: Roy, Aurko, et al.
Published: (2025)
by: Roy, Aurko, et al.
Published: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
Understanding the Emergence of Multimodal Representation Alignment
by: Tjandrasuwita, Megan, et al.
Published: (2025)
by: Tjandrasuwita, Megan, et al.
Published: (2025)
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
by: Yen, Jui-Nan, et al.
Published: (2024)
by: Yen, Jui-Nan, et al.
Published: (2024)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
SCATR: Simple Calibrated Test-Time Ranking
by: Shyamal, Divya, et al.
Published: (2026)
by: Shyamal, Divya, et al.
Published: (2026)
Efficient Distributed Optimization under Heavy-Tailed Noise
by: Lee, Su Hyeong, et al.
Published: (2025)
by: Lee, Su Hyeong, et al.
Published: (2025)
Differentially Private Model Merging
by: Yin, Qichuan, et al.
Published: (2026)
by: Yin, Qichuan, et al.
Published: (2026)
Compressing Many-Shots in In-Context Learning
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
by: Patel, Nirmal, et al.
Published: (2026)
by: Patel, Nirmal, et al.
Published: (2026)
Geometric Median (GM) Matching for Robust Data Pruning
by: Acharya, Anish, et al.
Published: (2024)
by: Acharya, Anish, et al.
Published: (2024)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
by: Basu, Soumya, et al.
Published: (2024)
by: Basu, Soumya, et al.
Published: (2024)
Federation over Text: Insight Sharing for Multi-Agent Reasoning
by: Yao, Dixi, et al.
Published: (2026)
by: Yao, Dixi, et al.
Published: (2026)
Geometric Median Matching for Robust k-Subset Selection from Noisy Data
by: Acharya, Anish, et al.
Published: (2025)
by: Acharya, Anish, et al.
Published: (2025)
Deep Reinforcement Learning for Sequential Combinatorial Auctions
by: Ravindranath, Sai Srivatsa, et al.
Published: (2024)
by: Ravindranath, Sai Srivatsa, et al.
Published: (2024)
Local to Global: Learning Dynamics and Effect of Initialization for Transformers
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
by: Makkuva, Ashok Vardhan, et al.
Published: (2024)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
by: Ekbote, Chanakya, et al.
Published: (2025)
by: Ekbote, Chanakya, et al.
Published: (2025)
Asynchronous Heavy-Tailed Optimization
by: Sun, Junfei, et al.
Published: (2026)
by: Sun, Junfei, et al.
Published: (2026)
Towards Quantifying the Preconditioning Effect of Adam
by: Das, Rudrajit, et al.
Published: (2024)
by: Das, Rudrajit, et al.
Published: (2024)
MatFormer: Nested Transformer for Elastic Inference
by: Devvrit, et al.
Published: (2023)
by: Devvrit, et al.
Published: (2023)
A Fresh Take on Stale Embeddings: Improving Dense Retriever Training with Corrector Networks
by: Monath, Nicholas, et al.
Published: (2024)
by: Monath, Nicholas, et al.
Published: (2024)
Measures of Information Reflect Memorization Patterns
by: Bansal, Rachit, et al.
Published: (2022)
by: Bansal, Rachit, et al.
Published: (2022)
Efficient Adaptive Federated Optimization
by: Lee, Su Hyeong, et al.
Published: (2024)
by: Lee, Su Hyeong, et al.
Published: (2024)
Adaptive Retrieval and Scalable Indexing for k-NN Search with Cross-Encoders
by: Yadav, Nishant, et al.
Published: (2024)
by: Yadav, Nishant, et al.
Published: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
by: Huang, Yukun, et al.
Published: (2025)
by: Huang, Yukun, et al.
Published: (2025)
Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and Acceleration
by: Khaled, Ahmed, et al.
Published: (2025)
by: Khaled, Ahmed, et al.
Published: (2025)
ExeDec: Execution Decomposition for Compositional Generalization in Neural Program Synthesis
by: Shi, Kensen, et al.
Published: (2023)
by: Shi, Kensen, et al.
Published: (2023)
OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization
by: Ong, Keane, et al.
Published: (2026)
by: Ong, Keane, et al.
Published: (2026)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
by: Gupta, Nilesh, et al.
Published: (2025)
by: Gupta, Nilesh, et al.
Published: (2025)
EHI: End-to-end Learning of Hierarchical Index for Efficient Dense Retrieval
by: Kumar, Ramnath, et al.
Published: (2023)
by: Kumar, Ramnath, et al.
Published: (2023)
Scalable In-context Ranking with Generative Models
by: Gupta, Nilesh, et al.
Published: (2025)
by: Gupta, Nilesh, et al.
Published: (2025)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Similar Items
-
The Art of Scaling Reinforcement Learning Compute for LLMs
by: Khatri, Devvrit, et al.
Published: (2025) -
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024) -
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
by: Bansal, Rachit, et al.
Published: (2025) -
LUCID: Attention with Preconditioned Representations
by: Duvvuri, Sai Surya, et al.
Published: (2026) -
Dual-Encoders for Extreme Multi-Label Classification
by: Gupta, Nilesh, et al.
Published: (2023)