Improving Retrospective Language Agents via Joint Policy Gradient Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Xueyang, Lan, Bo, Dai, Quanyu, Wang, Lei, Tang, Jiakai, Chen, Xu, Dong, Zhenhua, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
Process Supervision-Guided Policy Optimization for Code Generation
von: Dai, Ning, et al.
Veröffentlicht: (2024)
von: Dai, Ning, et al.
Veröffentlicht: (2024)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models
von: Bae, Ji Ho
Veröffentlicht: (2026)
von: Bae, Ji Ho
Veröffentlicht: (2026)
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
von: Gu, Yu, et al.
Veröffentlicht: (2024)
von: Gu, Yu, et al.
Veröffentlicht: (2024)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
von: Feng, Ruitao, et al.
Veröffentlicht: (2025)
von: Feng, Ruitao, et al.
Veröffentlicht: (2025)
Strategy Adaptation in Large Language Model Werewolf Agents
von: Nakamori, Fuya, et al.
Veröffentlicht: (2025)
von: Nakamori, Fuya, et al.
Veröffentlicht: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
von: Zhao, Lepeng, et al.
Veröffentlicht: (2026)
SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning
von: Zou, Run, et al.
Veröffentlicht: (2026)
von: Zou, Run, et al.
Veröffentlicht: (2026)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation
von: Yu, Jinzheng, et al.
Veröffentlicht: (2025)
von: Yu, Jinzheng, et al.
Veröffentlicht: (2025)
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
von: Chen, Jie, et al.
Veröffentlicht: (2025)
von: Chen, Jie, et al.
Veröffentlicht: (2025)
Lightweight Connective Detection Using Gradient Boosting
von: Er, Mustafa Erolcan, et al.
Veröffentlicht: (2024)
von: Er, Mustafa Erolcan, et al.
Veröffentlicht: (2024)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
von: Ge, Danying, et al.
Veröffentlicht: (2025)
von: Ge, Danying, et al.
Veröffentlicht: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Effective and Efficient Schema-aware Information Extraction Using On-Device Large Language Models
von: Wen, Zhihao, et al.
Veröffentlicht: (2025)
von: Wen, Zhihao, et al.
Veröffentlicht: (2025)
ToolGen: Unified Tool Retrieval and Calling via Generation
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
von: Dai, Xinbang, et al.
Veröffentlicht: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
von: Hong, Chunsan, et al.
Veröffentlicht: (2025)
von: Hong, Chunsan, et al.
Veröffentlicht: (2025)
Large Language Models Can Better Understand Knowledge Graphs Than We Thought
von: Dai, Xinbang, et al.
Veröffentlicht: (2024)
von: Dai, Xinbang, et al.
Veröffentlicht: (2024)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)
von: Xu, Chenjun, et al.
Veröffentlicht: (2025)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
R-Genie: Reasoning-Guided Generative Image Editing
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2025)
von: Hasan, Md Hasebul, et al.
Veröffentlicht: (2025)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
von: Costa, Rimom
Veröffentlicht: (2025)
von: Costa, Rimom
Veröffentlicht: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Low-Resource Court Judgment Summarization for Common Law Systems
von: Liu, Shuaiqi, et al.
Veröffentlicht: (2024)
von: Liu, Shuaiqi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025) -
Process Supervision-Guided Policy Optimization for Code Generation
von: Dai, Ning, et al.
Veröffentlicht: (2024) -
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026) -
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026) -
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
von: Hua, Yuncheng, et al.
Veröffentlicht: (2025)