DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Verma, Neha, Murray, Kenton, Duh, Kevin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Merging Feed-Forward Sublayers for Compressed Transformers
von: Verma, Neha, et al.
Veröffentlicht: (2025)
von: Verma, Neha, et al.
Veröffentlicht: (2025)
Exploring Representational Disparities Between Multilingual and Bilingual Translation Models
von: Verma, Neha, et al.
Veröffentlicht: (2023)
von: Verma, Neha, et al.
Veröffentlicht: (2023)
Merging Text Transformer Models from Different Initializations
von: Verma, Neha, et al.
Veröffentlicht: (2024)
von: Verma, Neha, et al.
Veröffentlicht: (2024)
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026)
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026)
Optimal Brain Iterative Merging: Mitigating Interference in LLM Merging
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)
Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
von: Verma, Neha, et al.
Veröffentlicht: (2026)
von: Verma, Neha, et al.
Veröffentlicht: (2026)
Upsample or Upweight? Balanced Training on Heavily Imbalanced Datasets
von: Li, Tianjian, et al.
Veröffentlicht: (2024)
von: Li, Tianjian, et al.
Veröffentlicht: (2024)
LLM Enhancer: Merged Approach using Vector Embedding for Reducing Large Language Model Hallucinations with External Knowledge
von: Rayhan, Naheed, et al.
Veröffentlicht: (2025)
von: Rayhan, Naheed, et al.
Veröffentlicht: (2025)
Effective Reasoning Chains Reduce Intrinsic Dimensionality
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
von: Prasad, Archiki, et al.
Veröffentlicht: (2026)
Merging Beyond: Streaming LLM Updates via Activation-Guided Rotations
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
von: Haider, Muhammad Umair, et al.
Veröffentlicht: (2025)
von: Haider, Muhammad Umair, et al.
Veröffentlicht: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Functionality-Oriented LLM Merging on the Fisher--Rao Manifold
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
von: Pynadath, Patrick, et al.
Veröffentlicht: (2025)
von: Pynadath, Patrick, et al.
Veröffentlicht: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
von: Yang, Yushi, et al.
Veröffentlicht: (2024)
von: Yang, Yushi, et al.
Veröffentlicht: (2024)
Less Finetuning, Better Retrieval: Rethinking LLM Adaptation for Biomedical Retrievers via Synthetic Data and Model Merging
von: Khattab, Sameh, et al.
Veröffentlicht: (2026)
von: Khattab, Sameh, et al.
Veröffentlicht: (2026)
Merge to Mix: Mixing Datasets via Model Merging
von: Tao, Zhixu Silvia, et al.
Veröffentlicht: (2025)
von: Tao, Zhixu Silvia, et al.
Veröffentlicht: (2025)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
von: Tomov, Tim, et al.
Veröffentlicht: (2026)
von: Tomov, Tim, et al.
Veröffentlicht: (2026)
Data Kernel Perspective Space Performance Guarantees for Synthetic Data from Transformer Models
von: Browder, Michael, et al.
Veröffentlicht: (2026)
von: Browder, Michael, et al.
Veröffentlicht: (2026)
Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging
von: Morrison, Jacob, et al.
Veröffentlicht: (2024)
von: Morrison, Jacob, et al.
Veröffentlicht: (2024)
Training-free LLM Merging for Multi-task Learning
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
von: Zhang, Junyan, et al.
Veröffentlicht: (2025)
Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling
von: Xiao, Tim Z., et al.
Veröffentlicht: (2025)
von: Xiao, Tim Z., et al.
Veröffentlicht: (2025)
SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning
von: Yu, Qifan, et al.
Veröffentlicht: (2026)
von: Yu, Qifan, et al.
Veröffentlicht: (2026)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
von: Vicentino, Caio
Veröffentlicht: (2026)
von: Vicentino, Caio
Veröffentlicht: (2026)
SeMe: Training-Free Language Model Merging via Semantic Alignment
von: Gu, Jian, et al.
Veröffentlicht: (2025)
von: Gu, Jian, et al.
Veröffentlicht: (2025)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
von: Qin, Zongyue, et al.
Veröffentlicht: (2024)
von: Qin, Zongyue, et al.
Veröffentlicht: (2024)
Neuron Patching: Semantic-based Neuron-level Language Model Repair for Code Generation
von: Gu, Jian, et al.
Veröffentlicht: (2023)
von: Gu, Jian, et al.
Veröffentlicht: (2023)
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
von: Liu, Weiqi, et al.
Veröffentlicht: (2026)
von: Liu, Weiqi, et al.
Veröffentlicht: (2026)
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents
von: Rawat, Mrinal, et al.
Veröffentlicht: (2025)
von: Rawat, Mrinal, et al.
Veröffentlicht: (2025)
Channel Merging: Preserving Specialization for Merged Experts
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
Grow Up and Merge: Scaling Strategies for Efficient Language Adaptation
von: Glocker, Kevin, et al.
Veröffentlicht: (2025)
von: Glocker, Kevin, et al.
Veröffentlicht: (2025)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal Budget
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training
von: Thomassen, Christian Brandt
Veröffentlicht: (2026)
von: Thomassen, Christian Brandt
Veröffentlicht: (2026)
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023)
von: Tam, Derek, et al.
Veröffentlicht: (2023)
When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer
von: Verma, Lucky
Veröffentlicht: (2026)
von: Verma, Lucky
Veröffentlicht: (2026)
Ähnliche Einträge
-
Merging Feed-Forward Sublayers for Compressed Transformers
von: Verma, Neha, et al.
Veröffentlicht: (2025) -
Exploring Representational Disparities Between Multilingual and Bilingual Translation Models
von: Verma, Neha, et al.
Veröffentlicht: (2023) -
Merging Text Transformer Models from Different Initializations
von: Verma, Neha, et al.
Veröffentlicht: (2024) -
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026) -
Optimal Brain Iterative Merging: Mitigating Interference in LLM Merging
von: Wang, Zhixiang, et al.
Veröffentlicht: (2025)