Gespeichert in:
| Hauptverfasser: | He, Zhitao, Liu, Zijun, Li, Peng, Fung, Yi R., Yan, Ming, Zhang, Ji, Huang, Fei, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.14496 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
von: He, Zhitao, et al.
Veröffentlicht: (2026)
von: He, Zhitao, et al.
Veröffentlicht: (2026)
ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
von: He, Zhitao, et al.
Veröffentlicht: (2025)
von: He, Zhitao, et al.
Veröffentlicht: (2025)
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2025)
MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
von: He, Zhitao, et al.
Veröffentlicht: (2025)
von: He, Zhitao, et al.
Veröffentlicht: (2025)
On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility
von: He, Zhitao, et al.
Veröffentlicht: (2026)
von: He, Zhitao, et al.
Veröffentlicht: (2026)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
von: Liu, Zijun, et al.
Veröffentlicht: (2024)
von: Liu, Zijun, et al.
Veröffentlicht: (2024)
MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
von: Liu, Jiayu, et al.
Veröffentlicht: (2025)
MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
von: Yang, Haolin, et al.
Veröffentlicht: (2025)
From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models
von: Zhang, Chenchen
Veröffentlicht: (2026)
von: Zhang, Chenchen
Veröffentlicht: (2026)
MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?
von: He, Zhitao, et al.
Veröffentlicht: (2025)
von: He, Zhitao, et al.
Veröffentlicht: (2025)
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
von: Ding, Fei, et al.
Veröffentlicht: (2026)
von: Ding, Fei, et al.
Veröffentlicht: (2026)
Re-ReST: Reflection-Reinforced Self-Training for Language Agents
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
von: Liu, Zijun, et al.
Veröffentlicht: (2023)
von: Liu, Zijun, et al.
Veröffentlicht: (2023)
Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks
von: He, Jiayi, et al.
Veröffentlicht: (2024)
von: He, Jiayi, et al.
Veröffentlicht: (2024)
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2026)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2026)
RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
von: Khandoga, Mykola, et al.
Veröffentlicht: (2026)
von: Khandoga, Mykola, et al.
Veröffentlicht: (2026)
LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
von: Chen, Junzhe, et al.
Veröffentlicht: (2024)
von: Chen, Junzhe, et al.
Veröffentlicht: (2024)
Towards Unified Alignment Between Agents, Humans, and Environment
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
von: Lei, Xuanyu, et al.
Veröffentlicht: (2025)
von: Lei, Xuanyu, et al.
Veröffentlicht: (2025)
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
von: Guo, Yiran, et al.
Veröffentlicht: (2025)
von: Guo, Yiran, et al.
Veröffentlicht: (2025)
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
von: Zhang, Naifan, et al.
Veröffentlicht: (2026)
von: Zhang, Naifan, et al.
Veröffentlicht: (2026)
Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
von: Ju, Feng, et al.
Veröffentlicht: (2025)
von: Ju, Feng, et al.
Veröffentlicht: (2025)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
von: Yang, Kailai, et al.
Veröffentlicht: (2025)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
von: Tang, Wenjie, et al.
Veröffentlicht: (2026)
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2026)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2026)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
von: Li, Ziheng, et al.
Veröffentlicht: (2026)
von: Li, Ziheng, et al.
Veröffentlicht: (2026)
Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers
von: Hu, Senkang, et al.
Veröffentlicht: (2026)
von: Hu, Senkang, et al.
Veröffentlicht: (2026)
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
von: Jiang, Xitai, et al.
Veröffentlicht: (2026)
von: Jiang, Xitai, et al.
Veröffentlicht: (2026)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
von: Xie, Guofu, et al.
Veröffentlicht: (2025)
RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents
von: Li, Mingchen, et al.
Veröffentlicht: (2026)
von: Li, Mingchen, et al.
Veröffentlicht: (2026)
APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented Generation
von: Chen, Kun, et al.
Veröffentlicht: (2026)
von: Chen, Kun, et al.
Veröffentlicht: (2026)
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
von: Yang, Matthew Y. R., et al.
Veröffentlicht: (2026)
von: Yang, Matthew Y. R., et al.
Veröffentlicht: (2026)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Exploiting Tree Structure for Credit Assignment in RL Training of LLMs
von: Tran, Hieu, et al.
Veröffentlicht: (2025)
von: Tran, Hieu, et al.
Veröffentlicht: (2025)
Reducing Distraction in Long-Context Language Models by Focused Learning
von: Wu, Zijun, et al.
Veröffentlicht: (2024)
von: Wu, Zijun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
von: Yang, Zonghan, et al.
Veröffentlicht: (2024) -
Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration
von: Liu, Zijun, et al.
Veröffentlicht: (2025) -
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
von: He, Zhitao, et al.
Veröffentlicht: (2026) -
ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
von: He, Zhitao, et al.
Veröffentlicht: (2025) -
CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2025)