Will it Merge? On The Causes of Model Mergeability
Fuente:
arXiv
Salvato in:
| Autori principali: | Rahamim, Adir, Yehudai, Asaf, Carmeli, Boaz, Choshen, Leshem, Mass, Yosi, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
Fast Forwarding Low-Rank Training
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
di: Carmeli, Boaz, et al.
Pubblicazione: (2024)
di: Carmeli, Boaz, et al.
Pubblicazione: (2024)
Mediocrity is the key for LLM as a Judge Anchor Selection
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2026)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2026)
Growing a Tail: Increasing Output Diversity in Large Language Models
di: Shur-Ofry, Michal, et al.
Pubblicazione: (2024)
di: Shur-Ofry, Michal, et al.
Pubblicazione: (2024)
Unsupervised Translation of Emergent Communication
di: Levy, Ido, et al.
Pubblicazione: (2025)
di: Levy, Ido, et al.
Pubblicazione: (2025)
CtD: Composition through Decomposition in Emergent Communication
di: Carmeli, Boaz, et al.
Pubblicazione: (2026)
di: Carmeli, Boaz, et al.
Pubblicazione: (2026)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
di: Habba, Eliya, et al.
Pubblicazione: (2026)
di: Habba, Eliya, et al.
Pubblicazione: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
di: Ramesh, Pratik, et al.
Pubblicazione: (2026)
di: Ramesh, Pratik, et al.
Pubblicazione: (2026)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
Can Gradient Descent Simulate Prompting?
di: Zhang, Eric, et al.
Pubblicazione: (2025)
di: Zhang, Eric, et al.
Pubblicazione: (2025)
Naturally Occurring Feedback is Common, Extractable and Useful
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2025)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Instructions Shape Production of Language, not Processing
di: Waldis, Andreas, et al.
Pubblicazione: (2026)
di: Waldis, Andreas, et al.
Pubblicazione: (2026)
Pretraining Language Models for Diachronic Linguistic Change Discovery
di: Fittschen, Elisabeth, et al.
Pubblicazione: (2025)
di: Fittschen, Elisabeth, et al.
Pubblicazione: (2025)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
Semantics and Spatiality of Emergent Communication
di: Zion, Rotem Ben, et al.
Pubblicazione: (2024)
di: Zion, Rotem Ben, et al.
Pubblicazione: (2024)
Investigating the Development of Task-Oriented Communication in Vision-Language Models
di: Carmeli, Boaz, et al.
Pubblicazione: (2026)
di: Carmeli, Boaz, et al.
Pubblicazione: (2026)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
di: Peisakhovsky, Yehonatan, et al.
Pubblicazione: (2025)
di: Peisakhovsky, Yehonatan, et al.
Pubblicazione: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
di: Din, Alexander Yom, et al.
Pubblicazione: (2023)
di: Din, Alexander Yom, et al.
Pubblicazione: (2023)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
di: Akyürek, Afra Feyza, et al.
Pubblicazione: (2024)
di: Akyürek, Afra Feyza, et al.
Pubblicazione: (2024)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
di: Hanna, Michael, et al.
Pubblicazione: (2024)
di: Hanna, Michael, et al.
Pubblicazione: (2024)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
di: Iluz, Bar, et al.
Pubblicazione: (2024)
di: Iluz, Bar, et al.
Pubblicazione: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023)
di: Arad, Dana, et al.
Pubblicazione: (2023)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
Are formal and functional linguistic mechanisms dissociated in language models?
di: Hanna, Michael, et al.
Pubblicazione: (2025)
di: Hanna, Michael, et al.
Pubblicazione: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
DEPTH: Discourse Education through Pre-Training Hierarchically
di: Bamberger, Zachary, et al.
Pubblicazione: (2024)
di: Bamberger, Zachary, et al.
Pubblicazione: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
di: Ifergan, Maxim, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
di: Yehudai, Asaf, et al.
Pubblicazione: (2024) -
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023) -
Fast Forwarding Low-Rank Training
di: Rahamim, Adir, et al.
Pubblicazione: (2024) -
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
di: Carmeli, Boaz, et al.
Pubblicazione: (2024) -
Mediocrity is the key for LLM as a Judge Anchor Selection
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2026)