Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Zehao, Li, Gongxun, Ai, Tianxiang, Li, Yifei, Huang, Zixuan, Zhou, Wang, Zhuang, Fuzhen, Liu, Xianglong, Li, Jianxin, Wang, Deqing, Ban, Yikun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMBoost: Make Large Language Models Stronger with Boosting
di: Chen, Zehao, et al.
Pubblicazione: (2025)
di: Chen, Zehao, et al.
Pubblicazione: (2025)
Heterogeneous Agent Collaborative Reinforcement Learning
di: Zhang, Zhixia, et al.
Pubblicazione: (2026)
di: Zhang, Zhixia, et al.
Pubblicazione: (2026)
Policy Improvement Reinforcement Learning
di: Wang, Huaiyang, et al.
Pubblicazione: (2026)
di: Wang, Huaiyang, et al.
Pubblicazione: (2026)
UniFAR: A Unified Facet-Aware Retrieval Framework for Scientific Documents
di: Dou, Zheng, et al.
Pubblicazione: (2026)
di: Dou, Zheng, et al.
Pubblicazione: (2026)
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
di: Huang, Zixuan, et al.
Pubblicazione: (2025)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
Counterfactual Credit Policy Optimization for Multi-Agent Collaboration
di: Li, Zhongyi, et al.
Pubblicazione: (2026)
di: Li, Zhongyi, et al.
Pubblicazione: (2026)
Real-Time Aligned Reward Model beyond Semantics
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment
di: Xie, Hongyan, et al.
Pubblicazione: (2026)
di: Xie, Hongyan, et al.
Pubblicazione: (2026)
Adaptive Robust Estimator for Multi-Agent Reinforcement Learning
di: Li, Zhongyi, et al.
Pubblicazione: (2026)
di: Li, Zhongyi, et al.
Pubblicazione: (2026)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
di: Huang, Xu, et al.
Pubblicazione: (2024)
di: Huang, Xu, et al.
Pubblicazione: (2024)
Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
di: Nie, Fan, et al.
Pubblicazione: (2025)
di: Nie, Fan, et al.
Pubblicazione: (2025)
CloudMatch: Weak-to-Strong Consistency Learning for Semi-Supervised Cloud Detection
di: Zhao, Jiayi, et al.
Pubblicazione: (2026)
di: Zhao, Jiayi, et al.
Pubblicazione: (2026)
Reliable Weak-to-Strong Monitoring of LLM Agents
di: Kale, Neil, et al.
Pubblicazione: (2025)
di: Kale, Neil, et al.
Pubblicazione: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
di: Lu, Xiaodong, et al.
Pubblicazione: (2026)
di: Lu, Xiaodong, et al.
Pubblicazione: (2026)
Your Group-Relative Advantage Is Biased
di: Yang, Fengkai, et al.
Pubblicazione: (2026)
di: Yang, Fengkai, et al.
Pubblicazione: (2026)
Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One
di: Song, Yiwen, et al.
Pubblicazione: (2025)
di: Song, Yiwen, et al.
Pubblicazione: (2025)
Representation Learning for Weakly Supervised Relation Extraction
di: Li, Zhuang
Pubblicazione: (2021)
di: Li, Zhuang
Pubblicazione: (2021)
Selective Weak-to-Strong Generalization
di: Lang, Hao, et al.
Pubblicazione: (2025)
di: Lang, Hao, et al.
Pubblicazione: (2025)
Debate Helps Weak-to-Strong Generalization
di: Lang, Hao, et al.
Pubblicazione: (2025)
di: Lang, Hao, et al.
Pubblicazione: (2025)
Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision?
di: Li, Zihao, et al.
Pubblicazione: (2024)
di: Li, Zihao, et al.
Pubblicazione: (2024)
WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning
di: Ge, Haosen, et al.
Pubblicazione: (2025)
di: Ge, Haosen, et al.
Pubblicazione: (2025)
FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents
di: Dou, Zheng, et al.
Pubblicazione: (2025)
di: Dou, Zheng, et al.
Pubblicazione: (2025)
On the Weak Point of the Stronger Uncertainty Relation
di: Urbanowski, K.
Pubblicazione: (2025)
di: Urbanowski, K.
Pubblicazione: (2025)
Weak to Strong: VLM-Based Pseudo-Labeling as a Weakly Supervised Training Strategy in Multimodal Video-based Hidden Emotion Understanding Tasks
di: Wang, Yufei, et al.
Pubblicazione: (2026)
di: Wang, Yufei, et al.
Pubblicazione: (2026)
KiGRAS: Kinematic-Driven Generative Model for Realistic Agent Simulation
di: Zhao, Jianbo, et al.
Pubblicazione: (2024)
di: Zhao, Jianbo, et al.
Pubblicazione: (2024)
The Amazing Agent Race: Strong Tool Users, Weak Navigators
di: Kim, Zae Myung, et al.
Pubblicazione: (2026)
di: Kim, Zae Myung, et al.
Pubblicazione: (2026)
MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization
di: Lyu, Yougang, et al.
Pubblicazione: (2024)
di: Lyu, Yougang, et al.
Pubblicazione: (2024)
The Weak Form Is Stronger Than You Think
di: Messenger, Daniel A., et al.
Pubblicazione: (2024)
di: Messenger, Daniel A., et al.
Pubblicazione: (2024)
Weak-to-Strong Knowledge Distillation Accelerates Visual Learning
di: Li, Baiang, et al.
Pubblicazione: (2026)
di: Li, Baiang, et al.
Pubblicazione: (2026)
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
di: Li, Ming, et al.
Pubblicazione: (2024)
di: Li, Ming, et al.
Pubblicazione: (2024)
Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
Weak-to-Strong Jailbreaking on Large Language Models
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
di: Zhao, Xuandong, et al.
Pubblicazione: (2024)
One for Dozens: Adaptive REcommendation for All Domains with Counterfactual Augmentation
di: Luo, Huishi, et al.
Pubblicazione: (2024)
di: Luo, Huishi, et al.
Pubblicazione: (2024)
CDC: Causal Domain Clustering for Multi-Domain Recommendation
di: Luo, Huishi, et al.
Pubblicazione: (2025)
di: Luo, Huishi, et al.
Pubblicazione: (2025)
Updateable Data-Driven Cardinality Estimator with Bounded Q-error
di: Li, Yingze, et al.
Pubblicazione: (2024)
di: Li, Yingze, et al.
Pubblicazione: (2024)
The Greater the Interaction, the Stronger the Learning Performance? Examining Pedagogical Agents' Interactive Presence in Instructional Videos
di: Changcheng Wu, et al.
Pubblicazione: (2024)
di: Changcheng Wu, et al.
Pubblicazione: (2024)
Student Guides Teacher: Weak-to-Strong Inference via Spectral Orthogonal Exploration
di: Wang, Dayu, et al.
Pubblicazione: (2026)
di: Wang, Dayu, et al.
Pubblicazione: (2026)
Debate Helps Weak Judges Reward Stronger Models
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
di: Elasky, Ethan, et al.
Pubblicazione: (2026)
Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
di: Xie, Hongyan, et al.
Pubblicazione: (2025)
di: Xie, Hongyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLMBoost: Make Large Language Models Stronger with Boosting
di: Chen, Zehao, et al.
Pubblicazione: (2025) -
Heterogeneous Agent Collaborative Reinforcement Learning
di: Zhang, Zhixia, et al.
Pubblicazione: (2026) -
Policy Improvement Reinforcement Learning
di: Wang, Huaiyang, et al.
Pubblicazione: (2026) -
UniFAR: A Unified Facet-Aware Retrieval Framework for Scientific Documents
di: Dou, Zheng, et al.
Pubblicazione: (2026) -
Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
di: Huang, Zixuan, et al.
Pubblicazione: (2025)