Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Zhiyuan, Hu, Yunhai, Liu, Juncheng, Li, Shuyue Stella, Wang, Yucheng, Xu, Zhen, Ng, See-Kiong, Luu, Anh Tuan, Xu, Xinxing, Hooi, Bryan, Breazeal, Cynthia, Park, Hae Won |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
par: Hu, Zhiyuan, et autres
Publié: (2026)
par: Hu, Zhiyuan, et autres
Publié: (2026)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
par: Zhao, James Xu, et autres
Publié: (2025)
par: Zhao, James Xu, et autres
Publié: (2025)
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
par: Hu, Zhiyuan, et autres
Publié: (2025)
par: Hu, Zhiyuan, et autres
Publié: (2025)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
par: Nguyen, Thong, et autres
Publié: (2025)
par: Nguyen, Thong, et autres
Publié: (2025)
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
par: Hu, Zhiyuan, et autres
Publié: (2024)
par: Hu, Zhiyuan, et autres
Publié: (2024)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
par: He, Yufei, et autres
Publié: (2025)
par: He, Yufei, et autres
Publié: (2025)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
par: Hu, Zhiyuan, et autres
Publié: (2024)
par: Hu, Zhiyuan, et autres
Publié: (2024)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
par: Goodge, Adam, et autres
Publié: (2025)
par: Goodge, Adam, et autres
Publié: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
par: Du, Mingzhe, et autres
Publié: (2024)
par: Du, Mingzhe, et autres
Publié: (2024)
How Does Response Length Affect Long-Form Factuality
par: Zhao, James Xu, et autres
Publié: (2025)
par: Zhao, James Xu, et autres
Publié: (2025)
Multi-Scale Contrastive Learning for Video Temporal Grounding
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
par: Formento, Brian, et autres
Publié: (2024)
par: Formento, Brian, et autres
Publié: (2024)
Vision-and-Language Pretraining
par: Nguyen, Thong, et autres
Publié: (2022)
par: Nguyen, Thong, et autres
Publié: (2022)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
par: Nguyen, Thong, et autres
Publié: (2023)
par: Nguyen, Thong, et autres
Publié: (2023)
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration
par: Xu, Lin, et autres
Publié: (2023)
par: Xu, Lin, et autres
Publié: (2023)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
par: Cao, Tri, et autres
Publié: (2026)
par: Cao, Tri, et autres
Publié: (2026)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
par: Nguyen, Thong, et autres
Publié: (2024)
par: Nguyen, Thong, et autres
Publié: (2024)
Does "Reasoning" with Large Language Models Improve Recognizing, Generating, and Reframing Unhelpful Thoughts?
par: Qi, Yilin, et autres
Publié: (2025)
par: Qi, Yilin, et autres
Publié: (2025)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
par: Nguyen, Thong, et autres
Publié: (2023)
par: Nguyen, Thong, et autres
Publié: (2023)
Topic Modeling as Multi-Objective Contrastive Optimization
par: Nguyen, Thong, et autres
Publié: (2024)
par: Nguyen, Thong, et autres
Publié: (2024)
Integrating Flow Theory and Adaptive Robot Roles: A Conceptual Model of Dynamic Robot Role Adaptation for the Enhanced Flow Experience in Long-term Multi-person Human-Robot Interactions
par: Chen, Huili, et autres
Publié: (2024)
par: Chen, Huili, et autres
Publié: (2024)
EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary Learning
par: He, Yufei, et autres
Publié: (2026)
par: He, Yufei, et autres
Publié: (2026)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
par: Wu, Tianyi, et autres
Publié: (2025)
par: Wu, Tianyi, et autres
Publié: (2025)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
par: Nguyen, Thong Thanh, et autres
Publié: (2024)
Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
par: Lee, Dong Won, et autres
Publié: (2025)
par: Lee, Dong Won, et autres
Publié: (2025)
Paper Espresso: From Paper Overload to Research Insight
par: Du, Mingzhe, et autres
Publié: (2026)
par: Du, Mingzhe, et autres
Publié: (2026)
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data
par: Kim, Yubin, et autres
Publié: (2024)
par: Kim, Yubin, et autres
Publié: (2024)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
par: Fu, Jinlan, et autres
Publié: (2025)
par: Fu, Jinlan, et autres
Publié: (2025)
HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs
par: Shen, Jocelyn, et autres
Publié: (2024)
par: Shen, Jocelyn, et autres
Publié: (2024)
A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling
par: Lee, Dong Won, et autres
Publié: (2026)
par: Lee, Dong Won, et autres
Publié: (2026)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
par: Lee, Dong Won, et autres
Publié: (2024)
par: Lee, Dong Won, et autres
Publié: (2024)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
par: Du, Mingzhe, et autres
Publié: (2025)
par: Du, Mingzhe, et autres
Publié: (2025)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
par: Le, Khoi, et autres
Publié: (2026)
par: Le, Khoi, et autres
Publié: (2026)
MGDCF: Distance Learning via Markov Graph Diffusion for Neural Collaborative Filtering
par: Hu, Jun, et autres
Publié: (2022)
par: Hu, Jun, et autres
Publié: (2022)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
par: Wu, Tianyi, et autres
Publié: (2025)
par: Wu, Tianyi, et autres
Publié: (2025)
CET2: Modelling Topic Transitions for Coherent and Engaging Knowledge-Grounded Conversations
par: Xu, Lin, et autres
Publié: (2024)
par: Xu, Lin, et autres
Publié: (2024)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
par: Kim, Yubin, et autres
Publié: (2024)
par: Kim, Yubin, et autres
Publié: (2024)
Social Robots as Social Proxies for Fostering Connection and Empathy Towards Humanity
par: Shen, Jocelyn, et autres
Publié: (2025)
par: Shen, Jocelyn, et autres
Publié: (2025)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
par: Shen, Jocelyn, et autres
Publié: (2025)
par: Shen, Jocelyn, et autres
Publié: (2025)
Documents similaires
-
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
par: Hu, Zhiyuan, et autres
Publié: (2026) -
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
par: Zhao, James Xu, et autres
Publié: (2025) -
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
par: Hu, Zhiyuan, et autres
Publié: (2025) -
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
par: Nguyen, Thong, et autres
Publié: (2025) -
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
par: Hu, Zhiyuan, et autres
Publié: (2024)