Language-based Trial and Error Falls Behind in the Era of Experience
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haoyu, Ma, Guozheng, Cui, Shugang, Kong, Yilun, Luo, Haotian, Shen, Li, Gao, Mengya, Wu, Yichao, Wang, Xiaogang, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
von: Kong, Yilun, et al.
Veröffentlicht: (2025)
von: Kong, Yilun, et al.
Veröffentlicht: (2025)
What Makes Value Learning Efficient in Residual Reinforcement Learning?
von: Ma, Guozheng, et al.
Veröffentlicht: (2026)
von: Ma, Guozheng, et al.
Veröffentlicht: (2026)
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2025)
von: Ma, Guozheng, et al.
Veröffentlicht: (2025)
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2025)
von: Ma, Guozheng, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Data Augmentation in Visual Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2022)
von: Ma, Guozheng, et al.
Veröffentlicht: (2022)
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
von: Ren, Zhiyao, et al.
Veröffentlicht: (2026)
von: Ren, Zhiyao, et al.
Veröffentlicht: (2026)
Concept-Guided Backdoor Attack on Vision Language Models
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
von: Zhang, Junjie, et al.
Veröffentlicht: (2025)
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
von: Ma, Guozheng, et al.
Veröffentlicht: (2023)
von: Ma, Guozheng, et al.
Veröffentlicht: (2023)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
von: Liu, Xuxu, et al.
Veröffentlicht: (2025)
Safety Reasoning with Guidelines
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning
von: Zhang, Junjie, et al.
Veröffentlicht: (2026)
von: Zhang, Junjie, et al.
Veröffentlicht: (2026)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
Decentralized Clinical Trials in the Era of Real‐World Evidence: A Critical Assessment of Recent Experiences
von: Hongwei Wang, et al.
Veröffentlicht: (2025)
von: Hongwei Wang, et al.
Veröffentlicht: (2025)
Piccolo2: General Text Embedding with Multi-task Hybrid Loss Training
von: Huang, Junqin, et al.
Veröffentlicht: (2024)
von: Huang, Junqin, et al.
Veröffentlicht: (2024)
OOP: Object-Oriented Programming Evaluation Benchmark for Large Language Models
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
The Role of miR‐124‐3p/UHRF1 in NaAsO 2 ‐Induced Apoptosis of LX‐2 Cells via DNMT1/SOCS1
von: Mengyao Zhang, et al.
Veröffentlicht: (2025)
von: Mengyao Zhang, et al.
Veröffentlicht: (2025)
Catching Up Yet Still Falling Behind: Sources, Heterogeneity, and Implications of the Modest Female Educational Disadvantage in Rural China
von: Wensong Shen
Veröffentlicht: (2025)
von: Wensong Shen
Veröffentlicht: (2025)
"Don't Fall Behind": A Unified Framework of Dynastic Survival, Two-Stage Belief Error, and the Modern Involution Trap
von: Yang, Dong
Veröffentlicht: (2025)
von: Yang, Dong
Veröffentlicht: (2025)
QPO: Query-dependent Prompt Optimization via Multi-Loop Offline Reinforcement Learning
von: Kong, Yilun, et al.
Veröffentlicht: (2024)
von: Kong, Yilun, et al.
Veröffentlicht: (2024)
Haibu Mathematical-Medical Intelligent Agent:Enhancing Large Language Model Reliability in Medical Tasks via Verifiable Reasoning Chains
von: Zhang, Yilun, et al.
Veröffentlicht: (2025)
von: Zhang, Yilun, et al.
Veröffentlicht: (2025)
Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data Scheduler
von: Hu, Zixuan, et al.
Veröffentlicht: (2025)
von: Hu, Zixuan, et al.
Veröffentlicht: (2025)
The effect of air pollution and genetic susceptibility on systemic lupus erythematosus: comment on the article by Xing et al
von: Na Wang, et al.
Veröffentlicht: (2024)
von: Na Wang, et al.
Veröffentlicht: (2024)
Little ones can do big things: Small molecule inhibitors target PTPN2/PTPN1 for tumor immunotherapy
von: Junyu Wang, et al.
Veröffentlicht: (2024)
von: Junyu Wang, et al.
Veröffentlicht: (2024)
Paternal microbiota impacts offspring: health risks and reproductive insights
von: Junyu Wang, et al.
Veröffentlicht: (2024)
von: Junyu Wang, et al.
Veröffentlicht: (2024)
The Rise and Fall of the Initial Era
von: Porter, Simon J, et al.
Veröffentlicht: (2024)
von: Porter, Simon J, et al.
Veröffentlicht: (2024)
Two nonfinitely based additively idempotent semirings of order four
von: Yue, Mengya, et al.
Veröffentlicht: (2026)
von: Yue, Mengya, et al.
Veröffentlicht: (2026)
Why Did Apple Fall: Evaluating Curiosity in Large Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
CredID: Credible Multi-Bit Watermark for Large Language Models Identification
von: Jiang, Haoyu, et al.
Veröffentlicht: (2024)
von: Jiang, Haoyu, et al.
Veröffentlicht: (2024)
SELF-EMO: Emotional Self-Evolution from Recognition to Consistent Expression
von: Zhang, Shaowei, et al.
Veröffentlicht: (2026)
von: Zhang, Shaowei, et al.
Veröffentlicht: (2026)
Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
von: Luo, Haotian, et al.
Veröffentlicht: (2025)
HiRegEx: Interactive Visual Query and Exploration of Multivariate Hierarchical Data
von: Li, Guozheng, et al.
Veröffentlicht: (2024)
von: Li, Guozheng, et al.
Veröffentlicht: (2024)
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
von: Lu, Qingyu, et al.
Veröffentlicht: (2023)
On the Evolution of Federated Post-Training Large Language Models: A Model Accessibility View
von: Guo, Tao, et al.
Veröffentlicht: (2025)
von: Guo, Tao, et al.
Veröffentlicht: (2025)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
To Save Mobile Crowdsourcing from Cheap-talk: A Game Theoretic Learning Approach
von: Hao, Shugang, et al.
Veröffentlicht: (2023)
von: Hao, Shugang, et al.
Veröffentlicht: (2023)
Algorithm Design for Continual Learning in IoT Networks
von: Hao, Shugang, et al.
Veröffentlicht: (2024)
von: Hao, Shugang, et al.
Veröffentlicht: (2024)
Online Learning from Strategic Human Feedback in LLM Fine-Tuning
von: Hao, Shugang, et al.
Veröffentlicht: (2024)
von: Hao, Shugang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
von: Kong, Yilun, et al.
Veröffentlicht: (2025) -
What Makes Value Learning Efficient in Residual Reinforcement Learning?
von: Ma, Guozheng, et al.
Veröffentlicht: (2026) -
Rethinking the Role of Dynamic Sparse Training for Scalable Deep Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2025) -
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2025) -
A Comprehensive Survey of Data Augmentation in Visual Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2022)