Rationale-Aware Answer Verification by Pairwise Self-Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kawabata, Akira, Sugawara, Saku |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
von: Kayano, Yoko, et al.
Veröffentlicht: (2025)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
von: Oba, Miyu, et al.
Veröffentlicht: (2026)
What Makes Language Models Good-enough?
von: Asami, Daiki, et al.
Veröffentlicht: (2024)
von: Asami, Daiki, et al.
Veröffentlicht: (2024)
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026)
von: Emura, Rei, et al.
Veröffentlicht: (2026)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2025)
From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
von: Tan, Wentao, et al.
Veröffentlicht: (2025)
von: Tan, Wentao, et al.
Veröffentlicht: (2025)
TactfulToM: Do LLMs Have the Theory of Mind Ability to Understand White Lies?
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
von: Liu, Yiwei, et al.
Veröffentlicht: (2025)
Seeing the Reasoning: How LLM Rationales Influence User Trust and Decision-Making in Factual Verification Tasks
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
MoreHopQA: More Than Multi-hop Reasoning
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
von: Schnitzler, Julian, et al.
Veröffentlicht: (2024)
Which Feedback Works for Whom? Differential Effects of LLM-Generated Feedback Elements Across Learner Profiles
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026)
von: Furuhashi, Momoka, et al.
Veröffentlicht: (2026)
Persuasiveness of Generated Free-Text Rationales in Subjective Decisions: A Case Study on Pairwise Argument Ranking
von: Elaraby, Mohamed, et al.
Veröffentlicht: (2024)
von: Elaraby, Mohamed, et al.
Veröffentlicht: (2024)
Can Language Models Induce Grammatical Knowledge from Indirect Evidence?
von: Oba, Miyu, et al.
Veröffentlicht: (2024)
von: Oba, Miyu, et al.
Veröffentlicht: (2024)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
von: Sun, Xin, et al.
Veröffentlicht: (2026)
von: Sun, Xin, et al.
Veröffentlicht: (2026)
Enhancing Relation Extraction via Supervised Rationale Verification and Feedback
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Verbosity-Aware Rationale Reduction: Effective Reduction of Redundant Rationale via Principled Criteria
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
Measuring Human Involvement in AI-Generated Text: A Case Study on Academic Writing
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
von: Guo, Yuchen, et al.
Veröffentlicht: (2025)
RORA: Robust Free-Text Rationale Evaluation
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
Self-Explaining Hate Speech Detection with Moral Rationales
von: Vargas, Francielle, et al.
Veröffentlicht: (2026)
von: Vargas, Francielle, et al.
Veröffentlicht: (2026)
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
von: Huang, Heyuan, et al.
Veröffentlicht: (2025)
von: Huang, Heyuan, et al.
Veröffentlicht: (2025)
Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
von: Yan, Jianzhi, et al.
Veröffentlicht: (2025)
von: Yan, Jianzhi, et al.
Veröffentlicht: (2025)
EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
von: Wang, Xinda, et al.
Veröffentlicht: (2025)
von: Wang, Xinda, et al.
Veröffentlicht: (2025)
A Claim Decomposition Benchmark for Long-form Answer Verification
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
von: Wu, Yuanchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuanchen, et al.
Veröffentlicht: (2025)
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
Enhancing Event Causality Identification with Rationale and Structure-Aware Causal Question Answering
von: Zhang, Baiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Baiyan, et al.
Veröffentlicht: (2024)
PairEval: Open-domain Dialogue Evaluation with Pairwise Comparison
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2026)
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2026)
Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations
von: Brandl, Stephanie, et al.
Veröffentlicht: (2024)
von: Brandl, Stephanie, et al.
Veröffentlicht: (2024)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
von: Muscato, Benedetta, et al.
Veröffentlicht: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
von: Fayyaz, Mohsen, et al.
Veröffentlicht: (2024)
MAPLE: Micro Analysis of Pairwise Language Evolution for Few-Shot Claim Verification
von: Zeng, Xia, et al.
Veröffentlicht: (2024)
von: Zeng, Xia, et al.
Veröffentlicht: (2024)
Towards Self-Contained Answers: Entity-Based Answer Rewriting in Conversational Search
von: Sekulić, Ivan, et al.
Veröffentlicht: (2024)
von: Sekulić, Ivan, et al.
Veröffentlicht: (2024)
Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
von: Hong, Pingjun, et al.
Veröffentlicht: (2026)
von: Hong, Pingjun, et al.
Veröffentlicht: (2026)
Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
von: Li, Ziheng, et al.
Veröffentlicht: (2025)
von: Li, Ziheng, et al.
Veröffentlicht: (2025)
Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
von: Waseda, Futa, et al.
Veröffentlicht: (2025)
MARE: Multi-Aspect Rationale Extractor on Unsupervised Rationale Extraction
von: Jiang, Han, et al.
Veröffentlicht: (2024)
von: Jiang, Han, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026) -
Specification-Aware Machine Translation and Evaluation for Purpose Alignment
von: Kayano, Yoko, et al.
Veröffentlicht: (2025) -
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
von: Oba, Miyu, et al.
Veröffentlicht: (2026) -
What Makes Language Models Good-enough?
von: Asami, Daiki, et al.
Veröffentlicht: (2024) -
A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language Models
von: Emura, Rei, et al.
Veröffentlicht: (2026)