Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Fan, Feng, Lan, Ye, Haotian, Liang, Weixin, Lu, Pan, Yao, Huaxiu, Alahi, Alexandre, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
Synergistic Weak-Strong Collaboration by Aligning Preferences
by: Jiao, Yizhu, et al.
Published: (2025)
by: Jiao, Yizhu, et al.
Published: (2025)
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
by: Chen, Zehao, et al.
Published: (2026)
by: Chen, Zehao, et al.
Published: (2026)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
by: Yang, Wenkai, et al.
Published: (2024)
by: Yang, Wenkai, et al.
Published: (2024)
Weak-to-Strong Reasoning
by: Yang, Yuqing, et al.
Published: (2024)
by: Yang, Yuqing, et al.
Published: (2024)
Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation
by: Zou, Deyu, et al.
Published: (2025)
by: Zou, Deyu, et al.
Published: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
by: Kale, Neil, et al.
Published: (2025)
by: Kale, Neil, et al.
Published: (2025)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
Selective Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
Mixture of Weak & Strong Experts on Graphs
by: Zeng, Hanqing, et al.
Published: (2023)
by: Zeng, Hanqing, et al.
Published: (2023)
Debate Helps Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
Quantifying the Gain in Weak-to-Strong Generalization
by: Charikar, Moses, et al.
Published: (2024)
by: Charikar, Moses, et al.
Published: (2024)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
by: Fan, Chenghao, et al.
Published: (2024)
by: Fan, Chenghao, et al.
Published: (2024)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
by: Kim, Zae Myung, et al.
Published: (2026)
by: Kim, Zae Myung, et al.
Published: (2026)
Weak-to-Strong Generalization under Distribution Shifts
by: Jeon, Myeongho, et al.
Published: (2025)
by: Jeon, Myeongho, et al.
Published: (2025)
Incentivizing Strong Reasoning from Weak Supervision
by: Yuan, Yige, et al.
Published: (2025)
by: Yuan, Yige, et al.
Published: (2025)
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization
by: Lyu, Yougang, et al.
Published: (2024)
by: Lyu, Yougang, et al.
Published: (2024)
Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
On Strong and Weak Admissibility in Non-Flat Assumption-Based Argumentation
by: Berthold, Matti, et al.
Published: (2025)
by: Berthold, Matti, et al.
Published: (2025)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Weak-to-Strong Generalization Through the Data-Centric Lens
by: Shin, Changho, et al.
Published: (2024)
by: Shin, Changho, et al.
Published: (2024)
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One
by: Song, Yiwen, et al.
Published: (2025)
by: Song, Yiwen, et al.
Published: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
by: Geng, Scott, et al.
Published: (2025)
by: Geng, Scott, et al.
Published: (2025)
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
by: Osooli, Hamid, et al.
Published: (2026)
by: Osooli, Hamid, et al.
Published: (2026)
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning
by: Ge, Haosen, et al.
Published: (2025)
by: Ge, Haosen, et al.
Published: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026)
by: Kiyani, Shayan, et al.
Published: (2026)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
by: Deng, Wei
Published: (2026)
by: Deng, Wei
Published: (2026)
Bayesian WeakS-to-Strong from Text Classification to Generation
by: Cui, Ziyun, et al.
Published: (2024)
by: Cui, Ziyun, et al.
Published: (2024)
CureAgent: A Training-Free Executor-Analyst Framework for Clinical Reasoning
by: Xie, Ting-Ting, et al.
Published: (2025)
by: Xie, Ting-Ting, et al.
Published: (2025)
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
by: Song, Feifan, et al.
Published: (2025)
by: Song, Feifan, et al.
Published: (2025)
InjectFlow: Weak Guides Strong via Orthogonal Injection for Flow Matching
by: Wang, Dayu, et al.
Published: (2026)
by: Wang, Dayu, et al.
Published: (2026)
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
by: Jin, Can, et al.
Published: (2026)
by: Jin, Can, et al.
Published: (2026)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
by: Xue, Yihao, et al.
Published: (2025)
by: Xue, Yihao, et al.
Published: (2025)
Weak to Strong: VLM-Based Pseudo-Labeling as a Weakly Supervised Training Strategy in Multimodal Video-based Hidden Emotion Understanding Tasks
by: Wang, Yufei, et al.
Published: (2026)
by: Wang, Yufei, et al.
Published: (2026)
From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging
by: Sun, Qi, et al.
Published: (2026)
by: Sun, Qi, et al.
Published: (2026)
AI-Generated Prior Authorization Letters: Strong Clinical Content, Weak Administrative Scaffolding
by: Awan, Moiz Sadiq, et al.
Published: (2026)
by: Awan, Moiz Sadiq, et al.
Published: (2026)
Similar Items
-
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024) -
Synergistic Weak-Strong Collaboration by Aligning Preferences
by: Jiao, Yizhu, et al.
Published: (2025) -
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
by: Chen, Zehao, et al.
Published: (2026) -
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
by: Yang, Wenkai, et al.
Published: (2024) -
Weak-to-Strong Reasoning
by: Yang, Yuqing, et al.
Published: (2024)