Salvato in:
| Autori principali: | He, Guangxin, Cao, Yuan, He, Yutong, Bai, Tianyi, Chen, Kai, Yuan, Kun, Yuan, Binhang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.01352 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models
di: Chen, Guanduo, et al.
Pubblicazione: (2025)
di: Chen, Guanduo, et al.
Pubblicazione: (2025)
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
di: Lv, Mengtao, et al.
Pubblicazione: (2025)
di: Lv, Mengtao, et al.
Pubblicazione: (2025)
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
di: He, Yutong, et al.
Pubblicazione: (2023)
di: He, Yutong, et al.
Pubblicazione: (2023)
VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
di: Hu, Zengjie, et al.
Pubblicazione: (2025)
di: Hu, Zengjie, et al.
Pubblicazione: (2025)
Mixture-of-Channels: Exploiting Sparse FFNs for Efficient LLMs Pre-Training and Inference
di: Wu, Tong, et al.
Pubblicazione: (2025)
di: Wu, Tong, et al.
Pubblicazione: (2025)
Greedy Low-Rank Gradient Compression for Distributed Learning with Convergence Guarantees
di: Chen, Chuyan, et al.
Pubblicazione: (2025)
di: Chen, Chuyan, et al.
Pubblicazione: (2025)
Subspace Optimization for Large Language Models with Convergence Guarantees
di: He, Yutong, et al.
Pubblicazione: (2024)
di: He, Yutong, et al.
Pubblicazione: (2024)
Clapping: Removing Per-sample Storage for Pipeline Parallel Distributed Optimization with Communication Compression
di: Kong, Boao, et al.
Pubblicazione: (2025)
di: Kong, Boao, et al.
Pubblicazione: (2025)
Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression
di: He, Yutong, et al.
Pubblicazione: (2023)
di: He, Yutong, et al.
Pubblicazione: (2023)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
di: Chen, Chuyan, et al.
Pubblicazione: (2025)
di: Chen, Chuyan, et al.
Pubblicazione: (2025)
AtmosSci-Bench: Evaluating the Recent Advance of Large Language Model for Atmospheric Science
di: Li, Chenyue, et al.
Pubblicazione: (2025)
di: Li, Chenyue, et al.
Pubblicazione: (2025)
MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
di: Liu, Yuxi, et al.
Pubblicazione: (2025)
di: Liu, Yuxi, et al.
Pubblicazione: (2025)
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering with Scalable Context and Symbolic Extension
di: Qiu, Zipeng, et al.
Pubblicazione: (2024)
di: Qiu, Zipeng, et al.
Pubblicazione: (2024)
FBQuant: FeedBack Quantization for Large Language Models
di: Liu, Yijiang, et al.
Pubblicazione: (2025)
di: Liu, Yijiang, et al.
Pubblicazione: (2025)
RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers
di: Liu, Yuxi, et al.
Pubblicazione: (2026)
di: Liu, Yuxi, et al.
Pubblicazione: (2026)
Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
di: Chen, Yiming, et al.
Pubblicazione: (2024)
di: Chen, Yiming, et al.
Pubblicazione: (2024)
SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM Training
di: Li, Zhouyang, et al.
Pubblicazione: (2025)
di: Li, Zhouyang, et al.
Pubblicazione: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
di: Zhou, Wenyong, et al.
Pubblicazione: (2025)
di: Zhou, Wenyong, et al.
Pubblicazione: (2025)
Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
di: Hu, Rizhen, et al.
Pubblicazione: (2026)
di: Hu, Rizhen, et al.
Pubblicazione: (2026)
Transformers Simulate MLE for Sequence Generation in Bayesian Networks
di: Cao, Yuan, et al.
Pubblicazione: (2025)
di: Cao, Yuan, et al.
Pubblicazione: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
di: Yan, Ran, et al.
Pubblicazione: (2025)
di: Yan, Ran, et al.
Pubblicazione: (2025)
High-fidelity Multiphysics Modelling for Rapid Predictions Using Physics-informed Parallel Neural Operator
di: Yuan, Biao, et al.
Pubblicazione: (2025)
di: Yuan, Biao, et al.
Pubblicazione: (2025)
Understanding the Benefits of SimCLR Pre-Training in Two-Layer Convolutional Neural Networks
di: Zhang, Han, et al.
Pubblicazione: (2024)
di: Zhang, Han, et al.
Pubblicazione: (2024)
Unraveling Privacy Risks of Individual Fairness in Graph Neural Networks
di: Zhang, He, et al.
Pubblicazione: (2023)
di: Zhang, He, et al.
Pubblicazione: (2023)
On the Escaping Efficiency of Distributed Adversarial Training Algorithms
di: Cao, Ying, et al.
Pubblicazione: (2025)
di: Cao, Ying, et al.
Pubblicazione: (2025)
On the Opportunities of (Re)-Exploring Atmospheric Science by Foundation Models: A Case Study
di: Zhang, Lujia, et al.
Pubblicazione: (2024)
di: Zhang, Lujia, et al.
Pubblicazione: (2024)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
di: Chen, Xi, et al.
Pubblicazione: (2026)
di: Chen, Xi, et al.
Pubblicazione: (2026)
TimeRadar: A Domain-Rotatable Foundation Model for Time Series Anomaly Detection
di: He, Hui, et al.
Pubblicazione: (2026)
di: He, Hui, et al.
Pubblicazione: (2026)
Language Models as Hierarchy Encoders
di: He, Yuan, et al.
Pubblicazione: (2024)
di: He, Yuan, et al.
Pubblicazione: (2024)
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
di: Chen, Lisha, et al.
Pubblicazione: (2025)
di: Chen, Lisha, et al.
Pubblicazione: (2025)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
di: He, Guangxin, et al.
Pubblicazione: (2025)
di: He, Guangxin, et al.
Pubblicazione: (2025)
Robust Machine Unlearning for Quantized Neural Networks via Adaptive Gradient Reweighting with Similar Labels
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
di: Kim, Hyeonjae, et al.
Pubblicazione: (2025)
di: Kim, Hyeonjae, et al.
Pubblicazione: (2025)
Forget by Uncertainty: Orthogonal Entropy Unlearning for Quantized Neural Networks
di: Zhang, Tian, et al.
Pubblicazione: (2026)
di: Zhang, Tian, et al.
Pubblicazione: (2026)
A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
di: Chen, Yiming, et al.
Pubblicazione: (2025)
di: Chen, Yiming, et al.
Pubblicazione: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
di: Bai, Tianyi, et al.
Pubblicazione: (2025)
di: Bai, Tianyi, et al.
Pubblicazione: (2025)
Transformers versus the EM Algorithm in Multi-class Clustering
di: He, Yihan, et al.
Pubblicazione: (2025)
di: He, Yihan, et al.
Pubblicazione: (2025)
Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion Processes
di: Ding, Fangyu, et al.
Pubblicazione: (2026)
di: Ding, Fangyu, et al.
Pubblicazione: (2026)
Selective Prompt Anchoring for Code Generation
di: Tian, Yuan, et al.
Pubblicazione: (2024)
di: Tian, Yuan, et al.
Pubblicazione: (2024)
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs
di: He, Yufei, et al.
Pubblicazione: (2024)
di: He, Yufei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models
di: Chen, Guanduo, et al.
Pubblicazione: (2025) -
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
di: Lv, Mengtao, et al.
Pubblicazione: (2025) -
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
di: He, Yutong, et al.
Pubblicazione: (2023) -
VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
di: Hu, Zengjie, et al.
Pubblicazione: (2025) -
Mixture-of-Channels: Exploiting Sparse FFNs for Efficient LLMs Pre-Training and Inference
di: Wu, Tong, et al.
Pubblicazione: (2025)