Saved in:
| Main Authors: | Tyagi, Sahil, Wang, Feiyi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.18112 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Using Large-Batches in Federated Learning
by: Tyagi, Sahil
Published: (2025)
by: Tyagi, Sahil
Published: (2025)
OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
by: Tyagi, Sahil, et al.
Published: (2025)
by: Tyagi, Sahil, et al.
Published: (2025)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
Revisiting LARS for Large Batch Training Generalization of Neural Networks
by: Do, Khoi, et al.
Published: (2023)
by: Do, Khoi, et al.
Published: (2023)
Reasoning Stabilization Point: A Training-Time Signal for Stable Evidence and Shortcut Reliance
by: Dhayalkar, Sahil Rajesh
Published: (2026)
by: Dhayalkar, Sahil Rajesh
Published: (2026)
Synergizing Large Language Models and Task-specific Models for Time Series Anomaly Detection
by: Chen, Feiyi, et al.
Published: (2025)
by: Chen, Feiyi, et al.
Published: (2025)
Towards Optimizing the Costs of LLM Usage
by: Shekhar, Shivanshu, et al.
Published: (2024)
by: Shekhar, Shivanshu, et al.
Published: (2024)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
by: Markovic-Voronov, Jelena, et al.
Published: (2026)
Accelerating Large-Scale Dataset Distillation via Exploration-Exploitation Optimization
by: Alahmadi, Muhammad J., et al.
Published: (2026)
by: Alahmadi, Muhammad J., et al.
Published: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Training Greedy Policy for Proposal Batch Selection in Expensive Multi-Objective Combinatorial Optimization
by: Lee, Deokjae, et al.
Published: (2024)
by: Lee, Deokjae, et al.
Published: (2024)
GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
by: Tyagi, Sahil, et al.
Published: (2023)
by: Tyagi, Sahil, et al.
Published: (2023)
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
by: Lu, Yishun, et al.
Published: (2025)
by: Lu, Yishun, et al.
Published: (2025)
A Combinatorial Theory of Dropout: Subnetworks, Graph Geometry, and Generalization
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
Building Expressive and Tractable Probabilistic Generative Models: A Review
by: Sidheekh, Sahil, et al.
Published: (2024)
by: Sidheekh, Sahil, et al.
Published: (2024)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
A Decomposable Forward Process in Diffusion Models for Time-Series Forecasting
by: Caldas, Francisco, et al.
Published: (2026)
by: Caldas, Francisco, et al.
Published: (2026)
Scaling Up Data Parallelism in Decentralized Deep Learning
by: Xie, Bing, et al.
Published: (2025)
by: Xie, Bing, et al.
Published: (2025)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
by: Balaji, Vignesh, et al.
Published: (2025)
by: Balaji, Vignesh, et al.
Published: (2025)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Batch Acquisition Function Evaluations and Decouple Optimizer Updates for Faster Bayesian Optimization
by: Irie, Kaichi, et al.
Published: (2025)
by: Irie, Kaichi, et al.
Published: (2025)
Pareto Front-Diverse Batch Multi-Objective Bayesian Optimization
by: Ahmadianshalchi, Alaleh, et al.
Published: (2024)
by: Ahmadianshalchi, Alaleh, et al.
Published: (2024)
Edge Intelligence Optimization for Large Language Model Inference with Batching and Quantization
by: Zhang, Xinyuan, et al.
Published: (2024)
by: Zhang, Xinyuan, et al.
Published: (2024)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
Learning Multi-Pattern Normalities in the Frequency Domain for Efficient Time Series Anomaly Detection
by: Chen, Feiyi, et al.
Published: (2023)
by: Chen, Feiyi, et al.
Published: (2023)
Large-Batch, Iteration-Efficient Neural Bayesian Design Optimization
by: Ansari, Navid, et al.
Published: (2023)
by: Ansari, Navid, et al.
Published: (2023)
Accelerating Distributed ML Training via Selective Synchronization
by: Tyagi, Sahil, et al.
Published: (2023)
by: Tyagi, Sahil, et al.
Published: (2023)
Self-Certifying Primal-Dual Optimization Proxies for Large-Scale Batch Economic Dispatch
by: Klamkin, Michael, et al.
Published: (2025)
by: Klamkin, Michael, et al.
Published: (2025)
From Theory to Throughput: CUDA-Optimized APML for Large-Batch 3D Learning
by: Sharifipour, Sasan, et al.
Published: (2025)
by: Sharifipour, Sasan, et al.
Published: (2025)
Test-Time Training on Graphs with Large Language Models (LLMs)
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)
by: Egbuna, Nathan, et al.
Published: (2025)
Batch Bayesian Active Learning with Partial Batch Label Sampling
by: Hu, Kangping, et al.
Published: (2025)
by: Hu, Kangping, et al.
Published: (2025)
Generative AI in Ship Design
by: Thakur, Sahil, et al.
Published: (2024)
by: Thakur, Sahil, et al.
Published: (2024)
Federated Instrumental Variable Analysis via Federated Generalized Method of Moments
by: Geetika, et al.
Published: (2025)
by: Geetika, et al.
Published: (2025)
Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
Dynamic Context Adaptation and Information Flow Control in Transformers: Introducing the Evaluator Adjuster Unit and Gated Residual Connections
by: Dhayalkar, Sahil Rajesh
Published: (2024)
by: Dhayalkar, Sahil Rajesh
Published: (2024)
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
by: Yu, Chao, et al.
Published: (2025)
by: Yu, Chao, et al.
Published: (2025)
Sample Transform Cost-Based Training-Free Hallucination Detector for Large Language Models
by: Ding, Zeyang, et al.
Published: (2026)
by: Ding, Zeyang, et al.
Published: (2026)
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training
by: Zhu, Dingwei, et al.
Published: (2025)
by: Zhu, Dingwei, et al.
Published: (2025)
Similar Items
-
On Using Large-Batches in Federated Learning
by: Tyagi, Sahil
Published: (2025) -
OmniFed: A Modular Framework for Configurable Federated Learning from Edge to HPC
by: Tyagi, Sahil, et al.
Published: (2025) -
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026) -
Revisiting LARS for Large Batch Training Generalization of Neural Networks
by: Do, Khoi, et al.
Published: (2023) -
Reasoning Stabilization Point: A Training-Time Signal for Stable Evidence and Shortcut Reliance
by: Dhayalkar, Sahil Rajesh
Published: (2026)