Lessons from Training Grounded LLMs with Verifiable Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sim, Shang Hong, Pala, Tej Deep, Toh, Vernon, Chieu, Hai Leong, Zadeh, Amir, Li, Chuan, Majumder, Navonil, Poria, Soujanya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
von: Song, Maojia, et al.
Veröffentlicht: (2024)
von: Song, Maojia, et al.
Veröffentlicht: (2024)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025)
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025)
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
von: Pala, Tej Deep, et al.
Veröffentlicht: (2024)
von: Pala, Tej Deep, et al.
Veröffentlicht: (2024)
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
von: Ong, Brandon, et al.
Veröffentlicht: (2025)
von: Ong, Brandon, et al.
Veröffentlicht: (2025)
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
von: Deep, Pala Tej, et al.
Veröffentlicht: (2024)
von: Deep, Pala Tej, et al.
Veröffentlicht: (2024)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
von: Song, Maojia, et al.
Veröffentlicht: (2025)
von: Song, Maojia, et al.
Veröffentlicht: (2025)
PromptDistill: Query-based Selective Token Retention in Intermediate Layers for Efficient Large Language Model Inference
von: Jin, Weisheng, et al.
Veröffentlicht: (2025)
von: Jin, Weisheng, et al.
Veröffentlicht: (2025)
Inference Time Alignment with Reward-Guided Tree Search
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2024)
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2024)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
von: Hong, Pengfei, et al.
Veröffentlicht: (2024)
von: Hong, Pengfei, et al.
Veröffentlicht: (2024)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2024)
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2024)
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2025)
von: Toh, Vernon Y. H., et al.
Veröffentlicht: (2025)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
von: Majumder, Navonil, et al.
Veröffentlicht: (2024)
von: Majumder, Navonil, et al.
Veröffentlicht: (2024)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
Mustango: Toward Controllable Text-to-Music Generation
von: Melechovsky, Jan, et al.
Veröffentlicht: (2023)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2023)
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
von: Liu, Renhang, et al.
Veröffentlicht: (2025)
von: Liu, Renhang, et al.
Veröffentlicht: (2025)
Towards Robust Instruction Tuning on Multimodal Large Language Models
von: Han, Wei, et al.
Veröffentlicht: (2024)
von: Han, Wei, et al.
Veröffentlicht: (2024)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
von: Ghosal, Deepanway, et al.
Veröffentlicht: (2024)
von: Ghosal, Deepanway, et al.
Veröffentlicht: (2024)
PREMISE: Matching-based Prediction for Accurate Review Recommendation
von: Han, Wei, et al.
Veröffentlicht: (2025)
von: Han, Wei, et al.
Veröffentlicht: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
von: Bhardwaj, Rishabh, et al.
Veröffentlicht: (2024)
von: Bhardwaj, Rishabh, et al.
Veröffentlicht: (2024)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
von: Hazra, Rima, et al.
Veröffentlicht: (2024)
von: Hazra, Rima, et al.
Veröffentlicht: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
von: Hazra, Rima, et al.
Veröffentlicht: (2024)
von: Hazra, Rima, et al.
Veröffentlicht: (2024)
Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet Extraction
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
Stacked from One: Multi-Scale Self-Injection for Context Window Extension
von: Han, Wei, et al.
Veröffentlicht: (2026)
von: Han, Wei, et al.
Veröffentlicht: (2026)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
von: Ghimire, Mukesh, et al.
Veröffentlicht: (2026)
von: Ghimire, Mukesh, et al.
Veröffentlicht: (2026)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
Two are better than one: Context window extension with multi-grained self-injection
von: Han, Wei, et al.
Veröffentlicht: (2024)
von: Han, Wei, et al.
Veröffentlicht: (2024)
Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources
von: Li, Xingxuan, et al.
Veröffentlicht: (2023)
von: Li, Xingxuan, et al.
Veröffentlicht: (2023)
10 Open Challenges Steering the Future of Vision-Language-Action Models
von: Poria, Soujanya, et al.
Veröffentlicht: (2025)
von: Poria, Soujanya, et al.
Veröffentlicht: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
von: Lin, Han, et al.
Veröffentlicht: (2025)
von: Lin, Han, et al.
Veröffentlicht: (2025)
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
von: Han, Vernon Toh Yan, et al.
Veröffentlicht: (2023)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
von: Lei, Jingdi, et al.
Veröffentlicht: (2025)
von: Lei, Jingdi, et al.
Veröffentlicht: (2025)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
von: Li, Yingting, et al.
Veröffentlicht: (2024)
von: Li, Yingting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
von: Song, Maojia, et al.
Veröffentlicht: (2024) -
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025) -
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
von: Pala, Tej Deep, et al.
Veröffentlicht: (2024) -
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
von: Ong, Brandon, et al.
Veröffentlicht: (2025) -
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
von: Deep, Pala Tej, et al.
Veröffentlicht: (2024)