Re-ReST: Reflection-Reinforced Self-Training for Language Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dou, Zi-Yi, Yang, Cheng-Fu, Wu, Xueqing, Chang, Kai-Wei, Peng, Nanyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
Medical Vision-Language Pre-Training for Brain Abnormalities
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
von: An, Yongqi, et al.
Veröffentlicht: (2026)
von: An, Yongqi, et al.
Veröffentlicht: (2026)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
Matryoshka Query Transformer for Large Vision-Language Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
Self-Routing RAG: Binding Selective Retrieval with Knowledge Verbalization
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning
von: Meng, Silin, et al.
Veröffentlicht: (2024)
von: Meng, Silin, et al.
Veröffentlicht: (2024)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
von: Yang, Zonghan, et al.
Veröffentlicht: (2024)
SNaRe: Domain-aware Data Generation for Low-Resource Event Detection
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
von: Parekh, Tanmay, et al.
Veröffentlicht: (2025)
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
von: Zhoubian, Sining, et al.
Veröffentlicht: (2025)
von: Zhoubian, Sining, et al.
Veröffentlicht: (2025)
VDebugger: Harnessing Execution Feedback for Debugging Visual Programs
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
von: Dou, Zi-Yi, et al.
Veröffentlicht: (2024)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
von: Hu, Wenbo, et al.
Veröffentlicht: (2024)
Control Large Language Models via Divide and Conquer
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
von: Wu, Xueqing, et al.
Veröffentlicht: (2025)
von: Wu, Xueqing, et al.
Veröffentlicht: (2025)
ReBeCA: Unveiling Interpretable Behavior Hierarchy behind the Iterative Self-Reflection of Language Models with Causal Analysis
von: Yan, Tianqiang, et al.
Veröffentlicht: (2026)
von: Yan, Tianqiang, et al.
Veröffentlicht: (2026)
WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
von: He, Guanzhong, et al.
Veröffentlicht: (2025)
von: He, Guanzhong, et al.
Veröffentlicht: (2025)
Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization
von: He, Zhitao, et al.
Veröffentlicht: (2025)
von: He, Zhitao, et al.
Veröffentlicht: (2025)
ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization
von: Chen, Guoxin, et al.
Veröffentlicht: (2025)
von: Chen, Guoxin, et al.
Veröffentlicht: (2025)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
von: Li, Bingxuan, et al.
Veröffentlicht: (2025)
Iterative Self-Training for Code Generation via Reinforced Re-Ranking
von: Sorokin, Nikita, et al.
Veröffentlicht: (2025)
von: Sorokin, Nikita, et al.
Veröffentlicht: (2025)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
von: Wu, Di, et al.
Veröffentlicht: (2026)
von: Wu, Di, et al.
Veröffentlicht: (2026)
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression
von: Li, Yuankai, et al.
Veröffentlicht: (2024)
von: Li, Yuankai, et al.
Veröffentlicht: (2024)
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
DeepEdit: Knowledge Editing as Decoding with Constraints
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
von: Wang, Yiwei, et al.
Veröffentlicht: (2024)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
von: Deng, Yihe, et al.
Veröffentlicht: (2025)
DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
von: Wu, Xueqing, et al.
Veröffentlicht: (2024)
ReFT: Reasoning with Reinforced Fine-Tuning
von: Luong, Trung Quoc, et al.
Veröffentlicht: (2024)
von: Luong, Trung Quoc, et al.
Veröffentlicht: (2024)
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Peng, Xiangyu, et al.
Veröffentlicht: (2024)
ReAD: Reinforcement-Guided Capability Distillation for Large Language Models
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
von: Li, Shiyu, et al.
Veröffentlicht: (2025)
von: Li, Shiyu, et al.
Veröffentlicht: (2025)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
von: Yang, Kevin, et al.
Veröffentlicht: (2023)
von: Yang, Kevin, et al.
Veröffentlicht: (2023)
QUDSELECT: Selective Decoding for Questions Under Discussion Parsing
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
Information Re-Organization Improves Reasoning in Large Language Models
von: Cheng, Xiaoxia, et al.
Veröffentlicht: (2024)
von: Cheng, Xiaoxia, et al.
Veröffentlicht: (2024)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks
von: Yao, Jiashu, et al.
Veröffentlicht: (2024)
von: Yao, Jiashu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
von: Zhang, Dan, et al.
Veröffentlicht: (2024) -
Medical Vision-Language Pre-Training for Brain Abnormalities
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024) -
ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing
von: An, Yongqi, et al.
Veröffentlicht: (2026) -
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026) -
Con-ReCall: Detecting Pre-training Data in LLMs via Contrastive Decoding
von: Wang, Cheng, et al.
Veröffentlicht: (2024)