Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Zhiyuan, Wang, Yucheng, He, Yufei, Wu, Jiaying, Zhao, Yilun, Ng, See-Kiong, Breazeal, Cynthia, Luu, Anh Tuan, Park, Hae Won, Hooi, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025)
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
von: Lee, Dong Won, et al.
Veröffentlicht: (2025)
von: Lee, Dong Won, et al.
Veröffentlicht: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
von: Formento, Brian, et al.
Veröffentlicht: (2024)
von: Formento, Brian, et al.
Veröffentlicht: (2024)
Vision-and-Language Pretraining
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
von: Fu, Jinlan, et al.
Veröffentlicht: (2025)
von: Fu, Jinlan, et al.
Veröffentlicht: (2025)
HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs
von: Shen, Jocelyn, et al.
Veröffentlicht: (2024)
von: Shen, Jocelyn, et al.
Veröffentlicht: (2024)
How Does Response Length Affect Long-Form Factuality
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Does "Reasoning" with Large Language Models Improve Recognizing, Generating, and Reframing Unhelpful Thoughts?
von: Qi, Yilin, et al.
Veröffentlicht: (2025)
von: Qi, Yilin, et al.
Veröffentlicht: (2025)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
Topic Modeling as Multi-Objective Contrastive Optimization
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling
von: Lee, Dong Won, et al.
Veröffentlicht: (2026)
von: Lee, Dong Won, et al.
Veröffentlicht: (2026)
Integrating Flow Theory and Adaptive Robot Roles: A Conceptual Model of Dynamic Robot Role Adaptation for the Enhanced Flow Experience in Long-term Multi-person Human-Robot Interactions
von: Chen, Huili, et al.
Veröffentlicht: (2024)
von: Chen, Huili, et al.
Veröffentlicht: (2024)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
Paper Espresso: From Paper Overload to Research Insight
von: Du, Mingzhe, et al.
Veröffentlicht: (2026)
von: Du, Mingzhe, et al.
Veröffentlicht: (2026)
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge
von: Fu, Jinlan, et al.
Veröffentlicht: (2024)
von: Fu, Jinlan, et al.
Veröffentlicht: (2024)
TRACE: TRansformer-based Attribution using Contrastive Embeddings in LLMs
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Cheng, et al.
Veröffentlicht: (2024)
MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
von: Lee, Dong Won, et al.
Veröffentlicht: (2024)
von: Lee, Dong Won, et al.
Veröffentlicht: (2024)
PIED: Physics-Informed Experimental Design for Inverse Problems
von: Hemachandra, Apivich, et al.
Veröffentlicht: (2025)
von: Hemachandra, Apivich, et al.
Veröffentlicht: (2025)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
von: Du, Mingzhe, et al.
Veröffentlicht: (2025)
von: Du, Mingzhe, et al.
Veröffentlicht: (2025)
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data
von: Kim, Yubin, et al.
Veröffentlicht: (2024)
von: Kim, Yubin, et al.
Veröffentlicht: (2024)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
von: Wu, Tianyi, et al.
Veröffentlicht: (2025)
Fake News in Sheep's Clothing: Robust Fake News Detection Against LLM-Empowered Style Attacks
von: Wu, Jiaying, et al.
Veröffentlicht: (2023)
von: Wu, Jiaying, et al.
Veröffentlicht: (2023)
Social Robots as Social Proxies for Fostering Connection and Empathy Towards Humanity
von: Shen, Jocelyn, et al.
Veröffentlicht: (2025)
von: Shen, Jocelyn, et al.
Veröffentlicht: (2025)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
von: Shen, Jocelyn, et al.
Veröffentlicht: (2025)
von: Shen, Jocelyn, et al.
Veröffentlicht: (2025)
Towards accurate and reliable ICU outcome prediction: a multimodal learning framework based on belief function theory using structured EHRs and free-text notes
von: Ruan, Yucheng, et al.
Veröffentlicht: (2025)
von: Ruan, Yucheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026) -
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2025) -
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2024) -
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025) -
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
von: Goodge, Adam, et al.
Veröffentlicht: (2025)