Evaluating Human Alignment and Model Faithfulness of LLM Rationale
Fuente:
arXiv
Salvato in:
| Autori principali: | Fayyaz, Mohsen, Yin, Fan, Sun, Jiao, Peng, Nanyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multilingual Routing in Mixture-of-Experts
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
Rethinking Human Preference Evaluation of LLM Rationales
di: Li, Ziang, et al.
Pubblicazione: (2025)
di: Li, Ziang, et al.
Pubblicazione: (2025)
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
di: Hu, Wenbo, et al.
Pubblicazione: (2024)
Are Akpans Trick or Treat: Unveiling Helpful Biases in Assistant Systems
di: Sun, Jiao, et al.
Pubblicazione: (2022)
di: Sun, Jiao, et al.
Pubblicazione: (2022)
SafeWorld: Geo-Diverse Safety Alignment
di: Yin, Da, et al.
Pubblicazione: (2024)
di: Yin, Da, et al.
Pubblicazione: (2024)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
di: Yang, Kevin, et al.
Pubblicazione: (2023)
di: Yang, Kevin, et al.
Pubblicazione: (2023)
Model Extrapolation Expedites Alignment
di: Zheng, Chujie, et al.
Pubblicazione: (2024)
di: Zheng, Chujie, et al.
Pubblicazione: (2024)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
di: Sun, Peng, et al.
Pubblicazione: (2026)
di: Sun, Peng, et al.
Pubblicazione: (2026)
Everything is Plausible: Investigating the Impact of LLM Rationales on Human Notions of Plausibility
di: Palta, Shramay, et al.
Pubblicazione: (2025)
di: Palta, Shramay, et al.
Pubblicazione: (2025)
Enabling Scalable Evaluation of Bias Patterns in Medical LLMs
di: Fayyaz, Hamed, et al.
Pubblicazione: (2024)
di: Fayyaz, Hamed, et al.
Pubblicazione: (2024)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
di: Chu, SeongYeub, et al.
Pubblicazione: (2024)
LLM-A*: Large Language Model Enhanced Incremental Heuristic Search on Path Planning
di: Meng, Silin, et al.
Pubblicazione: (2024)
di: Meng, Silin, et al.
Pubblicazione: (2024)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
di: Xiong, Zhen, et al.
Pubblicazione: (2025)
Reasoning Pattern Matters: Learning to Reason without Human Rationales
di: Pang, Chaoxu, et al.
Pubblicazione: (2025)
di: Pang, Chaoxu, et al.
Pubblicazione: (2025)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
di: Liang, Fangzhou, et al.
Pubblicazione: (2025)
di: Liang, Fangzhou, et al.
Pubblicazione: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
di: Zhang, Hongbin, et al.
Pubblicazione: (2025)
MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction
di: Jiang, Han, et al.
Pubblicazione: (2024)
di: Jiang, Han, et al.
Pubblicazione: (2024)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
di: Suvarna, Ashima, et al.
Pubblicazione: (2024)
di: Suvarna, Ashima, et al.
Pubblicazione: (2024)
Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
di: Wei, Bowen
Pubblicazione: (2025)
di: Wei, Bowen
Pubblicazione: (2025)
Enhanced Multimodal Aspect-Based Sentiment Analysis by LLM-Generated Rationales
di: Cao, Jun, et al.
Pubblicazione: (2025)
di: Cao, Jun, et al.
Pubblicazione: (2025)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
di: Alon, Bar, et al.
Pubblicazione: (2026)
di: Alon, Bar, et al.
Pubblicazione: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
di: Adlakha, Vaibhav, et al.
Pubblicazione: (2023)
di: Adlakha, Vaibhav, et al.
Pubblicazione: (2023)
STORYSUMM: Evaluating Faithfulness in Story Summarization
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)
di: Subbiah, Melanie, et al.
Pubblicazione: (2024)
FaithLens: Detecting and Explaining Faithfulness Hallucination
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
LLM-REVal: Can We Trust LLM Reviewers Yet?
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
Dissecting Human and LLM Preferences
di: Li, Junlong, et al.
Pubblicazione: (2024)
di: Li, Junlong, et al.
Pubblicazione: (2024)
Rationale-guided Prompting for Knowledge-based Visual Question Answering
di: Hu, Zhongjian, et al.
Pubblicazione: (2024)
di: Hu, Zhongjian, et al.
Pubblicazione: (2024)
GenAI vs. Human Fact-Checkers: Accurate Ratings, Flawed Rationales
di: Tai, Yuehong Cassandra, et al.
Pubblicazione: (2025)
di: Tai, Yuehong Cassandra, et al.
Pubblicazione: (2025)
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency
di: Li, Taiji, et al.
Pubblicazione: (2024)
di: Li, Taiji, et al.
Pubblicazione: (2024)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
di: Fan, Yang
Pubblicazione: (2025)
di: Fan, Yang
Pubblicazione: (2025)
CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models
di: Tan, Zhehao, et al.
Pubblicazione: (2026)
di: Tan, Zhehao, et al.
Pubblicazione: (2026)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
di: Huang, Xin, et al.
Pubblicazione: (2026)
di: Huang, Xin, et al.
Pubblicazione: (2026)
FaithLM: Towards Faithful Explanations for Large Language Models
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
di: Guo, Hongyi, et al.
Pubblicazione: (2024)
di: Guo, Hongyi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Multilingual Routing in Mixture-of-Experts
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025) -
Rethinking Human Preference Evaluation of LLM Rationales
di: Li, Ziang, et al.
Pubblicazione: (2025) -
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
di: Hu, Wenbo, et al.
Pubblicazione: (2024) -
Are Akpans Trick or Treat: Unveiling Helpful Biases in Assistant Systems
di: Sun, Jiao, et al.
Pubblicazione: (2022) -
SafeWorld: Geo-Diverse Safety Alignment
di: Yin, Da, et al.
Pubblicazione: (2024)