The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Schoepp, Sheila, Jafaripour, Masoud, Cao, Yingyue, Yang, Tianpei, Abdollahi, Fatemeh, Golestan, Shadan, Sufiyan, Zahin, Zaiane, Osmar R., Taylor, Matthew E. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WINFlowNets: Warm-up Integrated Networks Training of Generative Flow Networks for Robotics and Machine Fault Adaptation
por: Sufiyan, Zahin, et al.
Publicado: (2026)
por: Sufiyan, Zahin, et al.
Publicado: (2026)
A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation
por: Sufiyan, Zahin, et al.
Publicado: (2025)
por: Sufiyan, Zahin, et al.
Publicado: (2025)
TLXML: Task-Level Explanation of Meta-Learning via Influence Functions
por: Mitsuka, Yoshihiro, et al.
Publicado: (2025)
por: Mitsuka, Yoshihiro, et al.
Publicado: (2025)
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
por: Schoepp, Sheila, et al.
Publicado: (2024)
por: Schoepp, Sheila, et al.
Publicado: (2024)
Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
por: Sheikhi, Hadi, et al.
Publicado: (2025)
por: Sheikhi, Hadi, et al.
Publicado: (2025)
SETTP: Style Extraction and Tunable Inference via Dual-level Transferable Prompt Learning
por: Jin, Chunzhen, et al.
Publicado: (2024)
por: Jin, Chunzhen, et al.
Publicado: (2024)
LaFFi: Leveraging Hybrid Natural Language Feedback for Fine-tuning Language Models
por: Li, Qianxi, et al.
Publicado: (2023)
por: Li, Qianxi, et al.
Publicado: (2023)
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning
por: Ma, Weiyu, et al.
Publicado: (2026)
por: Ma, Weiyu, et al.
Publicado: (2026)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
por: Huang, Chenyang, et al.
Publicado: (2024)
por: Huang, Chenyang, et al.
Publicado: (2024)
A Decoding Algorithm for Length-Control Summarization Based on Directed Acyclic Transformers
por: Huang, Chenyang, et al.
Publicado: (2025)
por: Huang, Chenyang, et al.
Publicado: (2025)
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
por: Huang, Chenyang, et al.
Publicado: (2025)
por: Huang, Chenyang, et al.
Publicado: (2025)
Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
por: Wang, Yinjie, et al.
Publicado: (2025)
por: Wang, Yinjie, et al.
Publicado: (2025)
Code-A1: Adversarial Evolving of Code LLM and Test LLM via Reinforcement Learning
por: Wang, Aozhe, et al.
Publicado: (2026)
por: Wang, Aozhe, et al.
Publicado: (2026)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
por: Tang, Zhenwei, et al.
Publicado: (2026)
por: Tang, Zhenwei, et al.
Publicado: (2026)
WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
por: Qi, Zehan, et al.
Publicado: (2024)
por: Qi, Zehan, et al.
Publicado: (2024)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
por: Ruan, Chi, et al.
Publicado: (2026)
por: Ruan, Chi, et al.
Publicado: (2026)
SEIF: Self-Evolving Reinforcement Learning for Instruction Following
por: Ren, Qingyu, et al.
Publicado: (2026)
por: Ren, Qingyu, et al.
Publicado: (2026)
DKG-LLM : A Framework for Medical Diagnosis and Personalized Treatment Recommendations via Dynamic Knowledge Graph and Large Language Model Integration
por: Sarabadani, Ali, et al.
Publicado: (2025)
por: Sarabadani, Ali, et al.
Publicado: (2025)
PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learning
por: Li, Bingxuan, et al.
Publicado: (2026)
por: Li, Bingxuan, et al.
Publicado: (2026)
7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement
por: Zhao, Pu, et al.
Publicado: (2024)
por: Zhao, Pu, et al.
Publicado: (2024)
LLM-Based Multi-Hop Question Answering with Knowledge Graph Integration in Evolving Environments
por: Chen, Ruirui, et al.
Publicado: (2024)
por: Chen, Ruirui, et al.
Publicado: (2024)
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
por: Zhang, Guibin, et al.
Publicado: (2025)
por: Zhang, Guibin, et al.
Publicado: (2025)
The Evolving Landscape of Generative Large Language Models and Traditional Natural Language Processing in Medicine
por: Yang, Rui, et al.
Publicado: (2025)
por: Yang, Rui, et al.
Publicado: (2025)
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
por: Yang, Ziyi, et al.
Publicado: (2025)
por: Yang, Ziyi, et al.
Publicado: (2025)
IDRL: An Individual-Aware Multimodal Depression-Related Representation Learning Framework for Depression Diagnosis
por: Wang, Chongxiao, et al.
Publicado: (2026)
por: Wang, Chongxiao, et al.
Publicado: (2026)
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
por: Li, Xiaoyuan, et al.
Publicado: (2026)
por: Li, Xiaoyuan, et al.
Publicado: (2026)
MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
por: Zhang, Shengtao, et al.
Publicado: (2026)
por: Zhang, Shengtao, et al.
Publicado: (2026)
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
por: Shao, Rulin, et al.
Publicado: (2025)
por: Shao, Rulin, et al.
Publicado: (2025)
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
por: Hu, Zhe, et al.
Publicado: (2025)
por: Hu, Zhe, et al.
Publicado: (2025)
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
por: Ling Team, et al.
Publicado: (2025)
por: Ling Team, et al.
Publicado: (2025)
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
por: Salimi, Amir, et al.
Publicado: (2025)
por: Salimi, Amir, et al.
Publicado: (2025)
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
por: Wu, Chengyue, et al.
Publicado: (2026)
por: Wu, Chengyue, et al.
Publicado: (2026)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
por: Wu, Rong, et al.
Publicado: (2025)
por: Wu, Rong, et al.
Publicado: (2025)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
por: Wang, Kangrui, et al.
Publicado: (2025)
por: Wang, Kangrui, et al.
Publicado: (2025)
FPGA Divide-and-Conquer Placement using Deep Reinforcement Learning
por: Wang, Shang, et al.
Publicado: (2024)
por: Wang, Shang, et al.
Publicado: (2024)
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation
por: Rouzegar, Hamidreza, et al.
Publicado: (2024)
por: Rouzegar, Hamidreza, et al.
Publicado: (2024)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
por: Xu, Ran, et al.
Publicado: (2025)
por: Xu, Ran, et al.
Publicado: (2025)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
por: Tamber, Manveer Singh, et al.
Publicado: (2025)
Transformative Influence of LLM and AI Tools in Student Social Media Engagement: Analyzing Personalization, Communication Efficiency, and Collaborative Learning
por: Bashiri, Masoud, et al.
Publicado: (2024)
por: Bashiri, Masoud, et al.
Publicado: (2024)
Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
por: Zhang, Yutong, et al.
Publicado: (2024)
por: Zhang, Yutong, et al.
Publicado: (2024)
Ejemplares similares
-
WINFlowNets: Warm-up Integrated Networks Training of Generative Flow Networks for Robotics and Machine Fault Adaptation
por: Sufiyan, Zahin, et al.
Publicado: (2026) -
A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation
por: Sufiyan, Zahin, et al.
Publicado: (2025) -
TLXML: Task-Level Explanation of Meta-Learning via Influence Functions
por: Mitsuka, Yoshihiro, et al.
Publicado: (2025) -
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
por: Schoepp, Sheila, et al.
Publicado: (2024) -
Improving LLM's Attachment to External Knowledge In Dialogue Generation Tasks Through Entity Anonymization
por: Sheikhi, Hadi, et al.
Publicado: (2025)