MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Lu, Zhang, Tianyu, Bu, Zhiqi, Wang, Suyuchen, He, Huan, Fu, Jie, Wu, Yonghui, Bian, Jiang, Chen, Yong, Bengio, Yoshua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
di: Zhang, Tianyu, et al.
Pubblicazione: (2024)
di: Zhang, Tianyu, et al.
Pubblicazione: (2024)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
di: Zhang, Tianyu, et al.
Pubblicazione: (2025)
di: Zhang, Tianyu, et al.
Pubblicazione: (2025)
Learning Decision Trees as Amortized Structure Inference
di: Mahfoud, Mohammed, et al.
Pubblicazione: (2025)
di: Mahfoud, Mohammed, et al.
Pubblicazione: (2025)
Amortizing intractable inference in large language models
di: Hu, Edward J., et al.
Pubblicazione: (2023)
di: Hu, Edward J., et al.
Pubblicazione: (2023)
Large Language Models as Amortized Pareto-Front Generators for Constrained Bi-Objective Convex Optimization
di: Xu, Peipei, et al.
Pubblicazione: (2026)
di: Xu, Peipei, et al.
Pubblicazione: (2026)
Machine learning and information theory concepts towards an AI Mathematician
di: Bengio, Yoshua, et al.
Pubblicazione: (2024)
di: Bengio, Yoshua, et al.
Pubblicazione: (2024)
Amortized Active Generation of Pareto Sets
di: Steinberg, Daniel M., et al.
Pubblicazione: (2025)
di: Steinberg, Daniel M., et al.
Pubblicazione: (2025)
Baking Symmetry into GFlowNets
di: Ma, George, et al.
Pubblicazione: (2024)
di: Ma, George, et al.
Pubblicazione: (2024)
Amortizing intractable inference in diffusion models for vision, language, and control
di: Venkatraman, Siddarth, et al.
Pubblicazione: (2024)
di: Venkatraman, Siddarth, et al.
Pubblicazione: (2024)
Scaling depth capacity via zero/one-layer model expansion
di: Bu, Zhiqi
Pubblicazione: (2025)
di: Bu, Zhiqi
Pubblicazione: (2025)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
di: Jiralerspong, Thomas, et al.
Pubblicazione: (2026)
di: Jiralerspong, Thomas, et al.
Pubblicazione: (2026)
Fast Monte Carlo Tree Diffusion: 100x Speedup via Parallel Sparse Planning
di: Yoon, Jaesik, et al.
Pubblicazione: (2025)
di: Yoon, Jaesik, et al.
Pubblicazione: (2025)
Pareto Front Approximation for Multi-Objective Session-Based Recommender Systems
di: Wilm, Timo, et al.
Pubblicazione: (2024)
di: Wilm, Timo, et al.
Pubblicazione: (2024)
C-MORL: Multi-Objective Reinforcement Learning through Efficient Discovery of Pareto Front
di: Liu, Ruohong, et al.
Pubblicazione: (2024)
di: Liu, Ruohong, et al.
Pubblicazione: (2024)
Optimal Linear MAP Decoding of Convolutional Codes
di: Li, Yonghui, et al.
Pubblicazione: (2025)
di: Li, Yonghui, et al.
Pubblicazione: (2025)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
di: Lo, Chung-Hsiang, et al.
Pubblicazione: (2026)
di: Lo, Chung-Hsiang, et al.
Pubblicazione: (2026)
Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
di: Scimeca, Luca, et al.
Pubblicazione: (2025)
di: Scimeca, Luca, et al.
Pubblicazione: (2025)
Adaptive parameter-efficient fine-tuning via Hessian-informed subset selection
di: Xu, Shiyun, et al.
Pubblicazione: (2025)
di: Xu, Shiyun, et al.
Pubblicazione: (2025)
On Generalization for Generative Flow Networks
di: Krichel, Anas, et al.
Pubblicazione: (2024)
di: Krichel, Anas, et al.
Pubblicazione: (2024)
Interventional Causal Representation Learning
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
A Complexity-Based Theory of Compositionality
di: Elmoznino, Eric, et al.
Pubblicazione: (2024)
di: Elmoznino, Eric, et al.
Pubblicazione: (2024)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
di: Assouel, Rim, et al.
Pubblicazione: (2025)
di: Assouel, Rim, et al.
Pubblicazione: (2025)
Relative Trajectory Balance is equivalent to Trust-PCL
di: Deleu, Tristan, et al.
Pubblicazione: (2025)
di: Deleu, Tristan, et al.
Pubblicazione: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
di: Mittal, Sarthak, et al.
Pubblicazione: (2025)
di: Mittal, Sarthak, et al.
Pubblicazione: (2025)
GFlowNet Foundations
di: Bengio, Yoshua, et al.
Pubblicazione: (2021)
di: Bengio, Yoshua, et al.
Pubblicazione: (2021)
Memory Efficient Neural Processes via Constant Memory Attention Block
di: Feng, Leo, et al.
Pubblicazione: (2023)
di: Feng, Leo, et al.
Pubblicazione: (2023)
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
di: Chen, Kesheng, et al.
Pubblicazione: (2025)
di: Chen, Kesheng, et al.
Pubblicazione: (2025)
Cascaded Transformer for Robust and Scalable SLA Decomposition via Amortized Optimization
di: Hsu, Cyril Shih-Huan
Pubblicazione: (2026)
di: Hsu, Cyril Shih-Huan
Pubblicazione: (2026)
Efficient Diversity-Preserving Diffusion Alignment via Gradient-Informed GFlowNets
di: Liu, Zhen, et al.
Pubblicazione: (2024)
di: Liu, Zhen, et al.
Pubblicazione: (2024)
Expected flow networks in stochastic environments and two-player zero-sum games
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
di: Jiralerspong, Marco, et al.
Pubblicazione: (2023)
Pareto Merging: Multi-Objective Optimization for Preference-Aware Model Merging
di: Chen, Weiyu, et al.
Pubblicazione: (2024)
di: Chen, Weiyu, et al.
Pubblicazione: (2024)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
di: Williams-King, David, et al.
Pubblicazione: (2025)
di: Williams-King, David, et al.
Pubblicazione: (2025)
RL, but don't do anything I wouldn't do
di: Cohen, Michael K., et al.
Pubblicazione: (2024)
di: Cohen, Michael K., et al.
Pubblicazione: (2024)
Local Search GFlowNets
di: Kim, Minsu, et al.
Pubblicazione: (2023)
di: Kim, Minsu, et al.
Pubblicazione: (2023)
Active Attacks: Red-teaming LLMs via Adaptive Environments
di: Yun, Taeyoung, et al.
Pubblicazione: (2025)
di: Yun, Taeyoung, et al.
Pubblicazione: (2025)
Approximating Pareto Frontiers in Stochastic Multi-Objective Optimization via Hashing and Randomization
di: Li, Jinzhao, et al.
Pubblicazione: (2026)
di: Li, Jinzhao, et al.
Pubblicazione: (2026)
A Newton Method for Hausdorff Approximations of the Pareto Front within Multi-objective Evolutionary Algorithms
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Multi-Objective Bayesian Optimization with Independent Tanimoto Kernel Gaussian Processes for Diverse Pareto Front Exploration
di: Yong, Anabel
Pubblicazione: (2025)
di: Yong, Anabel
Pubblicazione: (2025)
Pareto Front Shape-Agnostic Pareto Set Learning in Multi-Objective Optimization
di: Ye, Rongguang, et al.
Pubblicazione: (2024)
di: Ye, Rongguang, et al.
Pubblicazione: (2024)
Monte Carlo Tree Diffusion for System 2 Planning
di: Yoon, Jaesik, et al.
Pubblicazione: (2025)
di: Yoon, Jaesik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
di: Zhang, Tianyu, et al.
Pubblicazione: (2024) -
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
di: Zhang, Tianyu, et al.
Pubblicazione: (2025) -
Learning Decision Trees as Amortized Structure Inference
di: Mahfoud, Mohammed, et al.
Pubblicazione: (2025) -
Amortizing intractable inference in large language models
di: Hu, Edward J., et al.
Pubblicazione: (2023) -
Large Language Models as Amortized Pareto-Front Generators for Constrained Bi-Objective Convex Optimization
di: Xu, Peipei, et al.
Pubblicazione: (2026)