Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yue, Chang, Yi, Wu, Yuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
von: Han, Xudong, et al.
Veröffentlicht: (2025)
von: Han, Xudong, et al.
Veröffentlicht: (2025)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
von: Kugler, Kai
Veröffentlicht: (2025)
von: Kugler, Kai
Veröffentlicht: (2025)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
von: Zhang, Luyan
Veröffentlicht: (2025)
von: Zhang, Luyan
Veröffentlicht: (2025)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
von: Li, Zilin, et al.
Veröffentlicht: (2026)
von: Li, Zilin, et al.
Veröffentlicht: (2026)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
von: Bertina, Abbas, et al.
Veröffentlicht: (2025)
von: Bertina, Abbas, et al.
Veröffentlicht: (2025)
Layer-Aware Embedding Fusion for LLMs in Text Classifications
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
von: Gwak, Jiho, et al.
Veröffentlicht: (2025)
On the Influence of Discourse Relations in Persuasive Texts
von: Turk, Nawar, et al.
Veröffentlicht: (2025)
von: Turk, Nawar, et al.
Veröffentlicht: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Calibrated Confidence Estimation for Tabular Question Answering
von: Voss, Lukas
Veröffentlicht: (2026)
von: Voss, Lukas
Veröffentlicht: (2026)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
von: Resck, Lucas, et al.
Veröffentlicht: (2026)
TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
von: Bian, Tingcheng, et al.
Veröffentlicht: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
von: He, Yanjin, et al.
Veröffentlicht: (2025)
von: He, Yanjin, et al.
Veröffentlicht: (2025)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
von: Naser-Moghadasi, Mahdi, et al.
Veröffentlicht: (2026)
von: Naser-Moghadasi, Mahdi, et al.
Veröffentlicht: (2026)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
von: Zhang, Yuxuan
Veröffentlicht: (2025)
von: Zhang, Yuxuan
Veröffentlicht: (2025)
Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
von: Reddy, Sandeep, et al.
Veröffentlicht: (2025)
von: Reddy, Sandeep, et al.
Veröffentlicht: (2025)
Language Models Are Implicitly Continuous
von: Marro, Samuele, et al.
Veröffentlicht: (2025)
von: Marro, Samuele, et al.
Veröffentlicht: (2025)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
von: Du, Bangde, et al.
Veröffentlicht: (2025)
von: Du, Bangde, et al.
Veröffentlicht: (2025)
Machine Unlearning for Masked Diffusion Language Models
von: Lee, Georu, et al.
Veröffentlicht: (2026)
von: Lee, Georu, et al.
Veröffentlicht: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
von: Toit, Jaco du, et al.
Veröffentlicht: (2025)
von: Toit, Jaco du, et al.
Veröffentlicht: (2025)
Truth as a Compression Artifact in Language Model Training
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
von: Krestnikov, Konstantin
Veröffentlicht: (2026)
Cost-Aware Model Selection for Text Classification: Multi-Objective Trade-offs Between Fine-Tuned Encoders and LLM Prompting in Production
von: Gonzalez, Alberto Andres Valdes
Veröffentlicht: (2026)
von: Gonzalez, Alberto Andres Valdes
Veröffentlicht: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
von: Vera, Patricio
Veröffentlicht: (2026)
von: Vera, Patricio
Veröffentlicht: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
von: Katraouras, Thomas, et al.
Veröffentlicht: (2026)
von: Katraouras, Thomas, et al.
Veröffentlicht: (2026)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
von: Young, Richard J., et al.
Veröffentlicht: (2025)
von: Young, Richard J., et al.
Veröffentlicht: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
von: Ponnock, Jesse
Veröffentlicht: (2025)
von: Ponnock, Jesse
Veröffentlicht: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
von: Shravan, Rohan
Veröffentlicht: (2026)
von: Shravan, Rohan
Veröffentlicht: (2026)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
von: Saghir, Hamidreza
Veröffentlicht: (2026)
von: Saghir, Hamidreza
Veröffentlicht: (2026)
Fine-tuning of Large Language Models for Constituency Parsing Using a Sequence to Sequence Approach
von: Delgado, Francisco Jose Cortes, et al.
Veröffentlicht: (2025)
von: Delgado, Francisco Jose Cortes, et al.
Veröffentlicht: (2025)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025) -
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025) -
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
von: Han, Xudong, et al.
Veröffentlicht: (2025) -
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
von: Ubukata, Shunsuke
Veröffentlicht: (2026)