Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nie, Fan, Feng, Lan, Ye, Haotian, Liang, Weixin, Lu, Pan, Yao, Huaxiu, Alahi, Alexandre, Zou, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024)
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024)
Synergistic Weak-Strong Collaboration by Aligning Preferences
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
von: Chen, Zehao, et al.
Veröffentlicht: (2026)
von: Chen, Zehao, et al.
Veröffentlicht: (2026)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
Weak-to-Strong Reasoning
von: Yang, Yuqing, et al.
Veröffentlicht: (2024)
von: Yang, Yuqing, et al.
Veröffentlicht: (2024)
Weak-to-Strong GraphRAG: Aligning Weak Retrievers with Large Language Models for Graph-based Retrieval Augmented Generation
von: Zou, Deyu, et al.
Veröffentlicht: (2025)
von: Zou, Deyu, et al.
Veröffentlicht: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
von: Kale, Neil, et al.
Veröffentlicht: (2025)
von: Kale, Neil, et al.
Veröffentlicht: (2025)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
Selective Weak-to-Strong Generalization
von: Lang, Hao, et al.
Veröffentlicht: (2025)
von: Lang, Hao, et al.
Veröffentlicht: (2025)
Mixture of Weak & Strong Experts on Graphs
von: Zeng, Hanqing, et al.
Veröffentlicht: (2023)
von: Zeng, Hanqing, et al.
Veröffentlicht: (2023)
Debate Helps Weak-to-Strong Generalization
von: Lang, Hao, et al.
Veröffentlicht: (2025)
von: Lang, Hao, et al.
Veröffentlicht: (2025)
Quantifying the Gain in Weak-to-Strong Generalization
von: Charikar, Moses, et al.
Veröffentlicht: (2024)
von: Charikar, Moses, et al.
Veröffentlicht: (2024)
On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion
von: Fan, Chenghao, et al.
Veröffentlicht: (2024)
von: Fan, Chenghao, et al.
Veröffentlicht: (2024)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
von: Lu, Xingyu, et al.
Veröffentlicht: (2026)
von: Lu, Xingyu, et al.
Veröffentlicht: (2026)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
von: Kim, Zae Myung, et al.
Veröffentlicht: (2026)
von: Kim, Zae Myung, et al.
Veröffentlicht: (2026)
Weak-to-Strong Generalization under Distribution Shifts
von: Jeon, Myeongho, et al.
Veröffentlicht: (2025)
von: Jeon, Myeongho, et al.
Veröffentlicht: (2025)
Incentivizing Strong Reasoning from Weak Supervision
von: Yuan, Yige, et al.
Veröffentlicht: (2025)
von: Yuan, Yige, et al.
Veröffentlicht: (2025)
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
von: Wang, Songping, et al.
Veröffentlicht: (2025)
von: Wang, Songping, et al.
Veröffentlicht: (2025)
MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization
von: Lyu, Yougang, et al.
Veröffentlicht: (2024)
von: Lyu, Yougang, et al.
Veröffentlicht: (2024)
Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
On Strong and Weak Admissibility in Non-Flat Assumption-Based Argumentation
von: Berthold, Matti, et al.
Veröffentlicht: (2025)
von: Berthold, Matti, et al.
Veröffentlicht: (2025)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
von: Pawelczyk, Martin, et al.
Veröffentlicht: (2024)
Weak-to-Strong Generalization Through the Data-Centric Lens
von: Shin, Changho, et al.
Veröffentlicht: (2024)
von: Shin, Changho, et al.
Veröffentlicht: (2024)
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One
von: Song, Yiwen, et al.
Veröffentlicht: (2025)
von: Song, Yiwen, et al.
Veröffentlicht: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
von: Geng, Scott, et al.
Veröffentlicht: (2025)
von: Geng, Scott, et al.
Veröffentlicht: (2025)
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
von: Osooli, Hamid, et al.
Veröffentlicht: (2026)
von: Osooli, Hamid, et al.
Veröffentlicht: (2026)
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning
von: Ge, Haosen, et al.
Veröffentlicht: (2025)
von: Ge, Haosen, et al.
Veröffentlicht: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
von: Kiyani, Shayan, et al.
Veröffentlicht: (2026)
von: Kiyani, Shayan, et al.
Veröffentlicht: (2026)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
von: Deng, Wei
Veröffentlicht: (2026)
von: Deng, Wei
Veröffentlicht: (2026)
Bayesian WeakS-to-Strong from Text Classification to Generation
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
CureAgent: A Training-Free Executor-Analyst Framework for Clinical Reasoning
von: Xie, Ting-Ting, et al.
Veröffentlicht: (2025)
von: Xie, Ting-Ting, et al.
Veröffentlicht: (2025)
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
von: Song, Feifan, et al.
Veröffentlicht: (2025)
von: Song, Feifan, et al.
Veröffentlicht: (2025)
InjectFlow: Weak Guides Strong via Orthogonal Injection for Flow Matching
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight
von: Jin, Can, et al.
Veröffentlicht: (2026)
von: Jin, Can, et al.
Veröffentlicht: (2026)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Weak to Strong: VLM-Based Pseudo-Labeling as a Weakly Supervised Training Strategy in Multimodal Video-based Hidden Emotion Understanding Tasks
von: Wang, Yufei, et al.
Veröffentlicht: (2026)
von: Wang, Yufei, et al.
Veröffentlicht: (2026)
From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging
von: Sun, Qi, et al.
Veröffentlicht: (2026)
von: Sun, Qi, et al.
Veröffentlicht: (2026)
AI-Generated Prior Authorization Letters: Strong Clinical Content, Weak Administrative Scaffolding
von: Awan, Moiz Sadiq, et al.
Veröffentlicht: (2026)
von: Awan, Moiz Sadiq, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
von: Liu, Yuejiang, et al.
Veröffentlicht: (2024) -
Synergistic Weak-Strong Collaboration by Aligning Preferences
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025) -
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
von: Chen, Zehao, et al.
Veröffentlicht: (2026) -
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
von: Yang, Wenkai, et al.
Veröffentlicht: (2024) -
Weak-to-Strong Reasoning
von: Yang, Yuqing, et al.
Veröffentlicht: (2024)