Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Song, Maojia, Sim, Shang Hong, Bhardwaj, Rishabh, Chieu, Hai Leong, Majumder, Navonil, Poria, Soujanya |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Lessons from Training Grounded LLMs with Verifiable Rewards
par: Sim, Shang Hong, et autres
Publié: (2025)
par: Sim, Shang Hong, et autres
Publié: (2025)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
par: Hong, Pengfei, et autres
Publié: (2024)
par: Hong, Pengfei, et autres
Publié: (2024)
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
par: Deep, Pala Tej, et autres
Publié: (2024)
par: Deep, Pala Tej, et autres
Publié: (2024)
Inference Time Alignment with Reward-Guided Tree Search
par: Hung, Chia-Yu, et autres
Publié: (2024)
par: Hung, Chia-Yu, et autres
Publié: (2024)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
par: Bhardwaj, Rishabh, et autres
Publié: (2024)
par: Bhardwaj, Rishabh, et autres
Publié: (2024)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
par: Han, Vernon Toh Yan, et autres
Publié: (2024)
par: Han, Vernon Toh Yan, et autres
Publié: (2024)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
par: Majumder, Navonil, et autres
Publié: (2024)
par: Majumder, Navonil, et autres
Publié: (2024)
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
par: Pala, Tej Deep, et autres
Publié: (2024)
par: Pala, Tej Deep, et autres
Publié: (2024)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
par: Li, Yingting, et autres
Publié: (2024)
par: Li, Yingting, et autres
Publié: (2024)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
par: Song, Maojia, et autres
Publié: (2025)
par: Song, Maojia, et autres
Publié: (2025)
Improving Text-To-Audio Models with Synthetic Captions
par: Kong, Zhifeng, et autres
Publié: (2024)
par: Kong, Zhifeng, et autres
Publié: (2024)
PromptDistill: Query-based Selective Token Retention in Intermediate Layers for Efficient Large Language Model Inference
par: Jin, Weisheng, et autres
Publié: (2025)
par: Jin, Weisheng, et autres
Publié: (2025)
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost
par: Xuan, Richmond Sin Jing, et autres
Publié: (2026)
par: Xuan, Richmond Sin Jing, et autres
Publié: (2026)
Mustango: Toward Controllable Text-to-Music Generation
par: Melechovsky, Jan, et autres
Publié: (2023)
par: Melechovsky, Jan, et autres
Publié: (2023)
Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
par: Song, Maojia, et autres
Publié: (2025)
par: Song, Maojia, et autres
Publié: (2025)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
par: Hung, Chia-Yu, et autres
Publié: (2024)
par: Hung, Chia-Yu, et autres
Publié: (2024)
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
par: Chia, Yew Ken, et autres
Publié: (2024)
par: Chia, Yew Ken, et autres
Publié: (2024)
Towards Robust Instruction Tuning on Multimodal Large Language Models
par: Han, Wei, et autres
Publié: (2024)
par: Han, Wei, et autres
Publié: (2024)
PREMISE: Matching-based Prediction for Accurate Review Recommendation
par: Han, Wei, et autres
Publié: (2025)
par: Han, Wei, et autres
Publié: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
par: Chia, Yew Ken, et autres
Publié: (2024)
par: Chia, Yew Ken, et autres
Publié: (2024)
Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?
par: Chawla, Divij, et autres
Publié: (2025)
par: Chawla, Divij, et autres
Publié: (2025)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
par: Hazra, Rima, et autres
Publié: (2024)
par: Hazra, Rima, et autres
Publié: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
par: Hazra, Rima, et autres
Publié: (2024)
par: Hazra, Rima, et autres
Publié: (2024)
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
par: Toh, Vernon Y. H., et autres
Publié: (2024)
par: Toh, Vernon Y. H., et autres
Publié: (2024)
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
par: Zhou, Ruiwen, et autres
Publié: (2026)
par: Zhou, Ruiwen, et autres
Publié: (2026)
Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet Extraction
par: Sun, Qi, et autres
Publié: (2024)
par: Sun, Qi, et autres
Publié: (2024)
Stacked from One: Multi-Scale Self-Injection for Context Window Extension
par: Han, Wei, et autres
Publié: (2026)
par: Han, Wei, et autres
Publié: (2026)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
par: Wang, Yuhao, et autres
Publié: (2024)
par: Wang, Yuhao, et autres
Publié: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
par: Zhang, Shaozuo, et autres
Publié: (2025)
par: Zhang, Shaozuo, et autres
Publié: (2025)
Two are better than one: Context window extension with multi-grained self-injection
par: Han, Wei, et autres
Publié: (2024)
par: Han, Wei, et autres
Publié: (2024)
Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources
par: Li, Xingxuan, et autres
Publié: (2023)
par: Li, Xingxuan, et autres
Publié: (2023)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
par: Li, Xiang, et autres
Publié: (2024)
par: Li, Xiang, et autres
Publié: (2024)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
par: Li, Yingting, et autres
Publié: (2024)
par: Li, Yingting, et autres
Publié: (2024)
DialogXpert: Driving Intelligent and Emotion-Aware Conversations through Online Value-Based Reinforcement Learning with LLM Priors
par: Rakib, Tazeek Bin Abdur, et autres
Publié: (2025)
par: Rakib, Tazeek Bin Abdur, et autres
Publié: (2025)
Toward Robust Multimodal Learning using Multimodal Foundational Models
par: Zhao, Xianbing, et autres
Publié: (2024)
par: Zhao, Xianbing, et autres
Publié: (2024)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
par: Sun, Qi, et autres
Publié: (2024)
par: Sun, Qi, et autres
Publié: (2024)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
par: Hung, Chia-Yu, et autres
Publié: (2025)
par: Hung, Chia-Yu, et autres
Publié: (2025)
Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
par: Pan, Wenbo, et autres
Publié: (2025)
par: Pan, Wenbo, et autres
Publié: (2025)
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
par: Gupta, Prannaya, et autres
Publié: (2024)
par: Gupta, Prannaya, et autres
Publié: (2024)
TrustRAG: Enhancing Robustness and Trustworthiness in Retrieval-Augmented Generation
par: Zhou, Huichi, et autres
Publié: (2025)
par: Zhou, Huichi, et autres
Publié: (2025)
Documents similaires
-
Lessons from Training Grounded LLMs with Verifiable Rewards
par: Sim, Shang Hong, et autres
Publié: (2025) -
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
par: Hong, Pengfei, et autres
Publié: (2024) -
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
par: Deep, Pala Tej, et autres
Publié: (2024) -
Inference Time Alignment with Reward-Guided Tree Search
par: Hung, Chia-Yu, et autres
Publié: (2024) -
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
par: Bhardwaj, Rishabh, et autres
Publié: (2024)