ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Bae, Suyoung, Na, CheolWon, Lee, Jaehoon, Lee, Yumin, Choi, YunSeok, Lee, Jee-Hyong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
by: Na, CheolWon, et al.
Published: (2025)
by: Na, CheolWon, et al.
Published: (2025)
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
by: Bae, Suyoung, et al.
Published: (2025)
by: Bae, Suyoung, et al.
Published: (2025)
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
by: Bae, Suyoung, et al.
Published: (2025)
by: Bae, Suyoung, et al.
Published: (2025)
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
by: Lee, Jihyung, et al.
Published: (2025)
by: Lee, Jihyung, et al.
Published: (2025)
BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization
by: Rafid, Ahmed, et al.
Published: (2026)
by: Rafid, Ahmed, et al.
Published: (2026)
mFACE: Multilingual Summarization with Factual Consistency Evaluation
by: Aharoni, Roee, et al.
Published: (2022)
by: Aharoni, Roee, et al.
Published: (2022)
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
by: Flores, Lorenzo Jaime Yu, et al.
Published: (2024)
TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
by: Lee, Jin-Seop, et al.
Published: (2025)
by: Lee, Jin-Seop, et al.
Published: (2025)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
by: Kang, Jaehoon, et al.
Published: (2026)
by: Kang, Jaehoon, et al.
Published: (2026)
Improving Factual Consistency of News Summarization by Contrastive Preference Optimization
by: Feng, Huawen, et al.
Published: (2023)
by: Feng, Huawen, et al.
Published: (2023)
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
by: Thorat, Onkar, et al.
Published: (2024)
by: Thorat, Onkar, et al.
Published: (2024)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
by: Kim, Juyeon, et al.
Published: (2024)
by: Kim, Juyeon, et al.
Published: (2024)
Enhancing Factuality through Consensus and Consistency in Summarization Using Minimum Bayes Risk Decoding
by: Soetedjo, Riza Setiawan, et al.
Published: (2026)
by: Soetedjo, Riza Setiawan, et al.
Published: (2026)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
by: Mujahid, Zain Muhammad, et al.
Published: (2025)
PlainQAFact: Retrieval-augmented Factual Consistency Evaluation Metric for Biomedical Plain Language Summarization
by: You, Zhiwen, et al.
Published: (2025)
by: You, Zhiwen, et al.
Published: (2025)
Reasoning by Commented Code for Table Question Answering
by: Pyo, Seho, et al.
Published: (2026)
by: Pyo, Seho, et al.
Published: (2026)
ReGraM: Region-First Knowledge Graph Reasoning for Medical Question Answering
by: Lee, Chaerin, et al.
Published: (2026)
by: Lee, Chaerin, et al.
Published: (2026)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
by: Kim, Jaehoon, et al.
Published: (2026)
by: Kim, Jaehoon, et al.
Published: (2026)
ISQA: Informative Factuality Feedback for Scientific Summarization
by: Li, Zekai, et al.
Published: (2024)
by: Li, Zekai, et al.
Published: (2024)
Agent-as-Judge for Factual Summarization of Long Narratives
by: Jeong, Yeonseok, et al.
Published: (2025)
by: Jeong, Yeonseok, et al.
Published: (2025)
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL
by: Lee, Jimin, et al.
Published: (2025)
by: Lee, Jimin, et al.
Published: (2025)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
by: Kim, Yumin, et al.
Published: (2025)
by: Kim, Yumin, et al.
Published: (2025)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
SciZoom: A Large-scale Benchmark for Hierarchical Scientific Summarization across the LLM Era
by: Jang, Han, et al.
Published: (2026)
by: Jang, Han, et al.
Published: (2026)
MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
by: Lee, Kyungro, et al.
Published: (2025)
by: Lee, Kyungro, et al.
Published: (2025)
In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
by: Kim, Jaehoon, et al.
Published: (2025)
by: Kim, Jaehoon, et al.
Published: (2025)
SDS KoPub VDR: A Benchmark Dataset for Visual Document Retrieval in Korean Public Documents
by: Lee, Jaehoon, et al.
Published: (2025)
by: Lee, Jaehoon, et al.
Published: (2025)
From What to Respond to When to Respond: Timely Response Generation for Open-domain Dialogue Agents
by: Jang, Seongbo, et al.
Published: (2025)
by: Jang, Seongbo, et al.
Published: (2025)
Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models
by: Song, Jongyoon, et al.
Published: (2024)
by: Song, Jongyoon, et al.
Published: (2024)
Two-step Automated Cybercrime Coded Word Detection using Multi-level Representation Learning
by: Kim, Yongyeon, et al.
Published: (2024)
by: Kim, Yongyeon, et al.
Published: (2024)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
OrderSum: Semantic Sentence Ordering for Extractive Summarization
by: Kwon, Taewan, et al.
Published: (2025)
by: Kwon, Taewan, et al.
Published: (2025)
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
by: Oh, Juhyun, et al.
Published: (2026)
by: Oh, Juhyun, et al.
Published: (2026)
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
by: Yun, Taewon, et al.
Published: (2025)
by: Yun, Taewon, et al.
Published: (2025)
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
by: Joseph, Sebastian Antony, et al.
Published: (2024)
by: Joseph, Sebastian Antony, et al.
Published: (2024)
FACTS&EVIDENCE: An Interactive Tool for Transparent Fine-Grained Factual Verification of Machine-Generated Text
by: Boonsanong, Varich, et al.
Published: (2025)
by: Boonsanong, Varich, et al.
Published: (2025)
Using Similarity to Evaluate Factual Consistency in Summaries
by: Ye, Yuxuan, et al.
Published: (2024)
by: Ye, Yuxuan, et al.
Published: (2024)
Similar Items
-
Q-FAKER: Query-free Hard Black-box Attack via Controlled Generation
by: Na, CheolWon, et al.
Published: (2025) -
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
by: Bae, Suyoung, et al.
Published: (2026) -
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models
by: Bae, Suyoung, et al.
Published: (2025) -
SALAD: Improving Robustness and Generalization through Contrastive Learning with Structure-Aware and LLM-Driven Augmented Data
by: Bae, Suyoung, et al.
Published: (2025) -
DCG-SQL: Enhancing In-Context Learning for Text-to-SQL with Deep Contextual Schema Link Graph
by: Lee, Jihyung, et al.
Published: (2025)