Merging Beyond: Streaming LLM Updates via Activation-Guided Rotations
Fuente:
arXiv
Salvato in:
| Autori principali: | Yao, Yuxuan, Sheng, Haonan, Lv, Qingsong, Wu, Han, Liu, Shuqi, Liu, Zehua, Liu, Zengyan, Gao, Jiahui, Tan, Haochen, Fu, Xiaojin, Bai, Haoli, So, Hing Cheung, Guo, Zhijiang, Song, Linqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Activation-Guided Consensus Merging for Large Language Models
di: Yao, Yuxuan, et al.
Pubblicazione: (2025)
di: Yao, Yuxuan, et al.
Pubblicazione: (2025)
1bit-Merging: Dynamic Quantized Merging for Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
di: Wu, Han, et al.
Pubblicazione: (2025)
di: Wu, Han, et al.
Pubblicazione: (2025)
Sens-Merging: Sensitivity-Guided Parameter Balancing for Merging Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
Meta-Learning-Based Delayless Subband Adaptive Filter using Complex Self-Attention for Active Noise Control
di: Feng, Pengxing, et al.
Pubblicazione: (2024)
di: Feng, Pengxing, et al.
Pubblicazione: (2024)
Learning From Correctness Without Prompting Makes LLM Efficient Reasoner
di: Yao, Yuxuan, et al.
Pubblicazione: (2024)
di: Yao, Yuxuan, et al.
Pubblicazione: (2024)
Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
Gradient Shaping Beyond Clipping: A Functional Perspective on Update Magnitude Control
di: You, Haochen, et al.
Pubblicazione: (2025)
di: You, Haochen, et al.
Pubblicazione: (2025)
Bi-Chainer: Automated Large Language Models Reasoning with Bidirectional Chaining
di: Liu, Shuqi, et al.
Pubblicazione: (2024)
di: Liu, Shuqi, et al.
Pubblicazione: (2024)
REG: A Regularization Optimizer for Robust Training Dynamics
di: Liu, Zehua, et al.
Pubblicazione: (2025)
di: Liu, Zehua, et al.
Pubblicazione: (2025)
Low-Rank and Row-Sparse Decomposition for Joint DOA Estimation and Distorted Sensor Detection
di: Huang, Huiping, et al.
Pubblicazione: (2022)
di: Huang, Huiping, et al.
Pubblicazione: (2022)
Sparse array design for MIMO radar in multipath scenarios
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
Robust time-of-arrival localization via ADMM
di: Xiong, Wenxin, et al.
Pubblicazione: (2023)
di: Xiong, Wenxin, et al.
Pubblicazione: (2023)
Rotatable Antenna Aided Mixed Near-Field and Far-Field Communications in the Upper Mid-Band: Interference Analysis and Joint Optimization
di: Zhang, Yunpu, et al.
Pubblicazione: (2025)
di: Zhang, Yunpu, et al.
Pubblicazione: (2025)
Target Localization with Coprime Multistatic MIMO Radar via Coupled Canonical Polyadic Decomposition Based on Joint Eigenvalue Decomposition
di: Liao, Guo-Zhao, et al.
Pubblicazione: (2025)
di: Liao, Guo-Zhao, et al.
Pubblicazione: (2025)
Target Localization with a Coprime Multistatic MIMO Radar via Coupled Canonical Polyadic Decomposition Based on Joint EVD
di: Liao, Guo-Zhao, et al.
Pubblicazione: (2025)
di: Liao, Guo-Zhao, et al.
Pubblicazione: (2025)
LoRE-Merging: Exploring Low-Rank Estimation For Large Language Model Merging
di: Liu, Zehua, et al.
Pubblicazione: (2025)
di: Liu, Zehua, et al.
Pubblicazione: (2025)
AROMA: Autonomous Rank-one Matrix Adaptation
di: Sheng, Hao Nan, et al.
Pubblicazione: (2025)
di: Sheng, Hao Nan, et al.
Pubblicazione: (2025)
Convergence Analysis of Consensus-ADMM for General QCQP
di: Huang, Huiping, et al.
Pubblicazione: (2022)
di: Huang, Huiping, et al.
Pubblicazione: (2022)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
di: Liu, Zehua, et al.
Pubblicazione: (2026)
di: Liu, Zehua, et al.
Pubblicazione: (2026)
Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling
di: Yao, Yuxuan, et al.
Pubblicazione: (2024)
di: Yao, Yuxuan, et al.
Pubblicazione: (2024)
Technology, life‐long learning, and heutagogy
di: Hing‐yu So
Pubblicazione: (2024)
di: Hing‐yu So
Pubblicazione: (2024)
Reasoning in Action: MCTS-Driven Knowledge Retrieval for Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
Rotatable Antenna Enabled Multi-Cell Mixed Near-Field and Far-Field Communications
di: Zhang, Yunpu, et al.
Pubblicazione: (2026)
di: Zhang, Yunpu, et al.
Pubblicazione: (2026)
Sparse Fluid Antenna Arrays: Continuous Position Design Beyond Classical DOF Limits
di: Wu, Tuo, et al.
Pubblicazione: (2026)
di: Wu, Tuo, et al.
Pubblicazione: (2026)
WeightedKV: Attention Scores Weighted Key-Value Cache Merging for Large Language Models
di: Yuan, Jian, et al.
Pubblicazione: (2025)
di: Yuan, Jian, et al.
Pubblicazione: (2025)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
di: Lin, Haokun, et al.
Pubblicazione: (2024)
di: Lin, Haokun, et al.
Pubblicazione: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
di: Li, Siyuan, et al.
Pubblicazione: (2025)
di: Li, Siyuan, et al.
Pubblicazione: (2025)
PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
di: Tan, Haochen, et al.
Pubblicazione: (2024)
di: Tan, Haochen, et al.
Pubblicazione: (2024)
Representations of 3D Rotations: Mathematical Foundations and Comparative Analysis
di: Aiersilan, Aizierjiang, et al.
Pubblicazione: (2025)
di: Aiersilan, Aizierjiang, et al.
Pubblicazione: (2025)
A characterization of compactness via bilinear $T1$ theorem
di: Cao, Mingming, et al.
Pubblicazione: (2024)
di: Cao, Mingming, et al.
Pubblicazione: (2024)
Limited range extrapolation with quantitative bounds and applications
di: Cao, Mingming, et al.
Pubblicazione: (2022)
di: Cao, Mingming, et al.
Pubblicazione: (2022)
Beyond Information: Mechanisms of Risk Communication on Self‐Protective Behavior in Public Health Emergencies
di: Yunpeng Xu, et al.
Pubblicazione: (2026)
di: Yunpeng Xu, et al.
Pubblicazione: (2026)
Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
di: Liu, Shuqi, et al.
Pubblicazione: (2025)
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
di: Wan, Yingjia, et al.
Pubblicazione: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
di: Zhang, Luyuan, et al.
Pubblicazione: (2026)
di: Zhang, Luyuan, et al.
Pubblicazione: (2026)
Time-FFM: Towards LM-Empowered Federated Foundation Model for Time Series Forecasting
di: Liu, Qingxiang, et al.
Pubblicazione: (2024)
di: Liu, Qingxiang, et al.
Pubblicazione: (2024)
Topology-Aware Integrated Communication, Sensing, and Power Transfer for SAGIN
di: Yu, Han, et al.
Pubblicazione: (2026)
di: Yu, Han, et al.
Pubblicazione: (2026)
Hybrid Data-Driven SSM for Interpretable and Label-Free mmWave Channel Prediction
di: Sun, Yiyong, et al.
Pubblicazione: (2024)
di: Sun, Yiyong, et al.
Pubblicazione: (2024)
Performance Analysis and Low-Complexity Beamforming Design for Near-Field Physical Layer Security
di: Zhang, Yunpu, et al.
Pubblicazione: (2024)
di: Zhang, Yunpu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Activation-Guided Consensus Merging for Large Language Models
di: Yao, Yuxuan, et al.
Pubblicazione: (2025) -
1bit-Merging: Dynamic Quantized Merging for Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025) -
Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
di: Wu, Han, et al.
Pubblicazione: (2025) -
Sens-Merging: Sensitivity-Guided Parameter Balancing for Merging Large Language Models
di: Liu, Shuqi, et al.
Pubblicazione: (2025) -
Meta-Learning-Based Delayless Subband Adaptive Filter using Complex Self-Attention for Active Noise Control
di: Feng, Pengxing, et al.
Pubblicazione: (2024)