Power Law Guided Dynamic Sifting for Efficient Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Koley, Nirav, Singhania, Prajwal, Bhatele, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loki: Low-rank Keys for Efficient Sparse Attention
by: Singhania, Prajwal, et al.
Published: (2024)
by: Singhania, Prajwal, et al.
Published: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023)
by: Singh, Siddharth, et al.
Published: (2023)
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025)
by: Singhania, Prajwal, et al.
Published: (2025)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
by: Costello, Ian J., et al.
Published: (2020)
by: Costello, Ian J., et al.
Published: (2020)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination
by: Li, Yuanshuai, et al.
Published: (2026)
by: Li, Yuanshuai, et al.
Published: (2026)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
by: McLeish, Sean, et al.
Published: (2025)
by: McLeish, Sean, et al.
Published: (2025)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
by: Chaturvedi, Aman, et al.
Published: (2024)
by: Chaturvedi, Aman, et al.
Published: (2024)
Sifting out communities in large sparse networks
by: Climer, Sharlee, et al.
Published: (2024)
by: Climer, Sharlee, et al.
Published: (2024)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025)
by: Ranjan, Aditya K., et al.
Published: (2025)
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning
by: Koirala, Prajwal, et al.
Published: (2025)
by: Koirala, Prajwal, et al.
Published: (2025)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
CRAUM-Net: Contextual Recursive Attention with Uncertainty Modeling for Salient Object Detection
by: Sagar, Abhinav
Published: (2020)
by: Sagar, Abhinav
Published: (2020)
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning
by: Diwan, Nirav, et al.
Published: (2025)
by: Diwan, Nirav, et al.
Published: (2025)
Interactive Event Sifting using Bayesian Graph Neural Networks
by: Nascimento, José, et al.
Published: (2024)
by: Nascimento, José, et al.
Published: (2024)
Efficient Dilated Squeeze and Excitation Neural Operator for Differential Equations
by: Chauhan, Prajwal, et al.
Published: (2026)
by: Chauhan, Prajwal, et al.
Published: (2026)
Functional Groups are All you Need for Chemically Interpretable Molecular Property Prediction
by: Balaji, Roshan, et al.
Published: (2025)
by: Balaji, Roshan, et al.
Published: (2025)
Sifting through the Noise: A Survey of Diffusion Probabilistic Models and Their Applications to Biomolecules
by: Norton, Trevor, et al.
Published: (2024)
by: Norton, Trevor, et al.
Published: (2024)
Active Budget Allocation for Efficient Scaling Law Estimation via Surrogate-Guided Pruning
by: Schram, Viktoria, et al.
Published: (2026)
by: Schram, Viktoria, et al.
Published: (2026)
Machine Learning Hamiltonian Dynamical Systems with Sparse and Noisy Data
by: Thapar, Vedanta, et al.
Published: (2026)
by: Thapar, Vedanta, et al.
Published: (2026)
AASeg: Attention Aware Network for Real Time Semantic Segmentation
by: Sagar, Abhinav
Published: (2021)
by: Sagar, Abhinav
Published: (2021)
SoK: Privacy Preserving Machine Learning using Functional Encryption: Opportunities and Challenges
by: Panzade, Prajwal, et al.
Published: (2022)
by: Panzade, Prajwal, et al.
Published: (2022)
Solving Offline Reinforcement Learning with Decision Tree Regression
by: Koirala, Prajwal, et al.
Published: (2024)
by: Koirala, Prajwal, et al.
Published: (2024)
Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data
by: Braun, Guillaume, et al.
Published: (2025)
by: Braun, Guillaume, et al.
Published: (2025)
From Laws to Motivation: Guiding Exploration through Law-Based Reasoning and Rewards
by: Chen, Ziyu, et al.
Published: (2024)
by: Chen, Ziyu, et al.
Published: (2024)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2025)
by: Burnwal, Returaj, et al.
Published: (2025)
OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
by: Burnwal, Returaj, et al.
Published: (2026)
by: Burnwal, Returaj, et al.
Published: (2026)
When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
by: Zhang, Yizhou
Published: (2025)
by: Zhang, Yizhou
Published: (2025)
Manifold-Guided Attention Steering
by: Li, Ian, et al.
Published: (2026)
by: Li, Ian, et al.
Published: (2026)
EDiT: Efficient Diffusion Transformers with Linear Compressed Attention
by: Becker, Philipp, et al.
Published: (2025)
by: Becker, Philipp, et al.
Published: (2025)
Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data Spectra
by: Worschech, Roman, et al.
Published: (2024)
by: Worschech, Roman, et al.
Published: (2024)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
Operationalising Rawlsian Ethics for Fairness in Norm-Learning Agents
by: Woodgate, Jessica, et al.
Published: (2024)
by: Woodgate, Jessica, et al.
Published: (2024)
Attention Guided Alignment in Efficient Vision-Language Models
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
DPFAGA-Dynamic Power Flow Analysis and Fault Characteristics: A Graph Attention Neural Network
by: Le, Tan, et al.
Published: (2025)
by: Le, Tan, et al.
Published: (2025)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
by: Li, Junzhuo, et al.
Published: (2026)
by: Li, Junzhuo, et al.
Published: (2026)
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
by: Neumann, Oren, et al.
Published: (2024)
by: Neumann, Oren, et al.
Published: (2024)
Deep Learning-Based Surrogate Creep Modelling in Inconel 625: A High-Temperature Alloy Study
by: Das, Shubham, et al.
Published: (2025)
by: Das, Shubham, et al.
Published: (2025)
Similar Items
-
Loki: Low-rank Keys for Efficient Sparse Attention
by: Singhania, Prajwal, et al.
Published: (2024) -
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026) -
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023) -
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025) -
Optimizing Agentic Language Model Inference via Speculative Tool Calls
by: Nichols, Daniel, et al.
Published: (2025)