Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Muru, Mishra, Mayank, Zhou, Zhongzhu, Brandon, William, Wang, Jue, Kim, Yoon, Ragan-Kelley, Jonathan, Song, Shuaiwen Leon, Athiwaratkun, Ben, Dao, Tri |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025)
by: Nrusimha, Aniruddha, et al.
Published: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025)
by: Zhang, Zhenyu, et al.
Published: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024)
by: Brandon, William, et al.
Published: (2024)
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
by: Wang, Junlin, et al.
Published: (2025)
by: Wang, Junlin, et al.
Published: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026)
by: Jia, Jinda, et al.
Published: (2026)
Mixture-of-Agents Enhances Large Language Model Capabilities
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
by: Thapa, Rahul, et al.
Published: (2024)
by: Thapa, Rahul, et al.
Published: (2024)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
by: He, Linda, et al.
Published: (2025)
by: He, Linda, et al.
Published: (2025)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
by: Maheswaran, Monishwaran, et al.
Published: (2026)
by: Maheswaran, Monishwaran, et al.
Published: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
by: Guo, Wentao, et al.
Published: (2025)
by: Guo, Wentao, et al.
Published: (2025)
Training-Free Activation Sparsity in Large Language Models
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
by: Chandra, Kartik, et al.
Published: (2026)
by: Chandra, Kartik, et al.
Published: (2026)
When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Framework
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
Hardware-Efficient Attention for Fast Decoding
by: Zadouri, Ted, et al.
Published: (2025)
by: Zadouri, Ted, et al.
Published: (2025)
How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
Improving Estonian Text Simplification through Pretrained Language Models and Custom Datasets
by: Barbu, Eduard, et al.
Published: (2025)
by: Barbu, Eduard, et al.
Published: (2025)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
by: Zhou, Zhongzhu, et al.
Published: (2026)
by: Zhou, Zhongzhu, et al.
Published: (2026)
Energy efficiency optimization of task-parallel codes on asymmetric architectures
by: Costero, Luis, et al.
Published: (2024)
by: Costero, Luis, et al.
Published: (2024)
More efficient sifting for grid norms, and applications to multiparty communication complexity
by: Kelley, Zander, et al.
Published: (2025)
by: Kelley, Zander, et al.
Published: (2025)
Guided Optimization for Image Processing Pipelines
by: Ikarashi, Yuka, et al.
Published: (2021)
by: Ikarashi, Yuka, et al.
Published: (2021)
Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch
by: Douillard, Arthur, et al.
Published: (2025)
by: Douillard, Arthur, et al.
Published: (2025)
Visual moral inference and communication
by: Zhu, Warren, et al.
Published: (2025)
by: Zhu, Warren, et al.
Published: (2025)
Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
by: Caren, Matthew, et al.
Published: (2024)
by: Caren, Matthew, et al.
Published: (2024)
BILLNET: A Binarized Conv3D-LSTM Network with Logic-gated residual architecture for hardware-efficient video inference
by: Nguyen, Van Thien, et al.
Published: (2025)
by: Nguyen, Van Thien, et al.
Published: (2025)
Meschers: Geometry Processing of Impossible Objects
by: Dodik, Ana, et al.
Published: (2026)
by: Dodik, Ana, et al.
Published: (2026)
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
Introspective Diffusion Language Models
by: Yu, Yifan, et al.
Published: (2026)
by: Yu, Yifan, et al.
Published: (2026)
Dissociating model architectures from inference computations
by: Sajid, Noor, et al.
Published: (2025)
by: Sajid, Noor, et al.
Published: (2025)
Cost-aware simulation-based inference
by: Bharti, Ayush, et al.
Published: (2024)
by: Bharti, Ayush, et al.
Published: (2024)
Mitigating the Impact of Outlier Channels for Language Model Quantization with Activation Regularization
by: Nrusimha, Aniruddha, et al.
Published: (2024)
by: Nrusimha, Aniruddha, et al.
Published: (2024)
Measuring Sustainability Intention of ESG Fund Disclosure using Few-Shot Learning
by: Singh, Mayank, et al.
Published: (2024)
by: Singh, Mayank, et al.
Published: (2024)
Towards single-shot coherent imaging via overlap-free ptychography
by: Hoidn, Oliver, et al.
Published: (2026)
by: Hoidn, Oliver, et al.
Published: (2026)
Disentangling Reasoning and Knowledge in Medical Large Language Models
by: Thapa, Rahul, et al.
Published: (2025)
by: Thapa, Rahul, et al.
Published: (2025)
Target-aware Bayesian inference via generalized thermodynamic integration
by: Llorente, F., et al.
Published: (2025)
by: Llorente, F., et al.
Published: (2025)
Large-Scale Data Selection for Instruction Tuning
by: Ivison, Hamish, et al.
Published: (2025)
by: Ivison, Hamish, et al.
Published: (2025)
Empathy in Explanation
by: Collins, Katherine M., et al.
Published: (2025)
by: Collins, Katherine M., et al.
Published: (2025)
Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
by: Zhou, Zhongzhu, et al.
Published: (2025)
by: Zhou, Zhongzhu, et al.
Published: (2025)
A parallel implementation of reduced-order modeling of large-scale systems
by: Farcas, Ionut-Gabriel, et al.
Published: (2025)
by: Farcas, Ionut-Gabriel, et al.
Published: (2025)
Similar Items
-
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
by: Nrusimha, Aniruddha, et al.
Published: (2025) -
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
by: Zhou, Zhongzhu, et al.
Published: (2026) -
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
by: Zhang, Zhenyu, et al.
Published: (2025) -
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
by: Brandon, William, et al.
Published: (2024) -
Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
by: Wang, Junlin, et al.
Published: (2025)