Agent-as-Judge for Factual Summarization of Long Narratives
Fuente:
arXiv
Guardado en:
| Autores principales: | Jeong, Yeonseok, Kim, Minsoo, Hwang, Seung-won, Kim, Byung-Hak |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
por: Kim, Hyuntak, et al.
Publicado: (2025)
por: Kim, Hyuntak, et al.
Publicado: (2025)
ECoRAG: Evidentiality-guided Compression for Long Context RAG
por: Jeong, Yeonseok, et al.
Publicado: (2025)
por: Jeong, Yeonseok, et al.
Publicado: (2025)
CoEx -- Co-evolving World-model and Exploration
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
CREFT: Sequential Multi-Agent LLM for Character Relation Extraction
por: Chun, Ye Eun, et al.
Publicado: (2025)
por: Chun, Ye Eun, et al.
Publicado: (2025)
Disentangling Questions from Query Generation for Task-Adaptive Retrieval
por: Lee, Yoonsang, et al.
Publicado: (2024)
por: Lee, Yoonsang, et al.
Publicado: (2024)
Chaining Event Spans for Temporal Relation Grounding
por: Kim, Jongho, et al.
Publicado: (2025)
por: Kim, Jongho, et al.
Publicado: (2025)
R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQL
por: Han, Hojae, et al.
Publicado: (2026)
por: Han, Hojae, et al.
Publicado: (2026)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
por: Kim, Jongho, et al.
Publicado: (2025)
por: Kim, Jongho, et al.
Publicado: (2025)
DIAMOND: An LLM-Driven Agent for Context-Aware Baseball Highlight Summarization
por: Kang, Jeonghun, et al.
Publicado: (2025)
por: Kang, Jeonghun, et al.
Publicado: (2025)
Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents
por: Han, Hojae, et al.
Publicado: (2026)
por: Han, Hojae, et al.
Publicado: (2026)
Intended Target Identification for Anomia Patients with Gradient-based Selective Augmentation
por: Kim, Jongho, et al.
Publicado: (2025)
por: Kim, Jongho, et al.
Publicado: (2025)
Benchmarking Testing in Automated Theorem Proving
por: Kim, Jongyoon, et al.
Publicado: (2026)
por: Kim, Jongyoon, et al.
Publicado: (2026)
From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrieval
por: Lee, Dohyeon, et al.
Publicado: (2025)
por: Lee, Dohyeon, et al.
Publicado: (2025)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
por: Jung, Yeonjoon, et al.
Publicado: (2024)
por: Jung, Yeonjoon, et al.
Publicado: (2024)
Relevance to Utility: Process-Supervised Rewrite for RAG
por: Kim, Jaeyoung, et al.
Publicado: (2025)
por: Kim, Jaeyoung, et al.
Publicado: (2025)
HARP: Hesitation-Aware Reframing in Transformer Inference Pass
por: Storaï, Romain, et al.
Publicado: (2024)
por: Storaï, Romain, et al.
Publicado: (2024)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
por: Han, Hojae, et al.
Publicado: (2026)
por: Han, Hojae, et al.
Publicado: (2026)
SAFE: Stepwise Atomic Feedback for Error correction in Multi-hop Reasoning
por: Kwon, Daeyong, et al.
Publicado: (2026)
por: Kwon, Daeyong, et al.
Publicado: (2026)
Chain of Grounded Objectives: Bridging Process and Goal-oriented Prompting for Code Generation
por: Yeo, Sangyeop, et al.
Publicado: (2025)
por: Yeo, Sangyeop, et al.
Publicado: (2025)
OLAPH: Improving Factuality in Biomedical Long-form Question Answering
por: Jeong, Minbyul, et al.
Publicado: (2024)
por: Jeong, Minbyul, et al.
Publicado: (2024)
AcuRank: Uncertainty-Aware Adaptive Computation for Listwise Reranking
por: Yoon, Soyoung, et al.
Publicado: (2025)
por: Yoon, Soyoung, et al.
Publicado: (2025)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
por: Yoo, Jaeseok, et al.
Publicado: (2024)
por: Yoo, Jaeseok, et al.
Publicado: (2024)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
por: Han, Hojae, et al.
Publicado: (2024)
por: Han, Hojae, et al.
Publicado: (2024)
Anonpsy: A Graph-Based Framework for Structure-Preserving De-identification of Psychiatric Narratives
por: Lim, Kyung Ho, et al.
Publicado: (2026)
por: Lim, Kyung Ho, et al.
Publicado: (2026)
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
por: Zhong, Yang, et al.
Publicado: (2025)
por: Zhong, Yang, et al.
Publicado: (2025)
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
por: Samarinas, Chris, et al.
Publicado: (2025)
por: Samarinas, Chris, et al.
Publicado: (2025)
Towards Lifelong Dialogue Agents via Timeline-based Memory Management
por: Ong, Kai Tzu-iunn, et al.
Publicado: (2024)
por: Ong, Kai Tzu-iunn, et al.
Publicado: (2024)
Stress Testing Factual Consistency Metrics for Long-Document Summarization
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
por: Mujahid, Zain Muhammad, et al.
Publicado: (2025)
Ever-Evolving Memory by Blending and Refining the Past
por: Kim, Seo Hyun, et al.
Publicado: (2024)
por: Kim, Seo Hyun, et al.
Publicado: (2024)
ISQA: Informative Factuality Feedback for Scientific Summarization
por: Li, Zekai, et al.
Publicado: (2024)
por: Li, Zekai, et al.
Publicado: (2024)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
por: Han, Hojae, et al.
Publicado: (2025)
por: Han, Hojae, et al.
Publicado: (2025)
mFACE: Multilingual Summarization with Factual Consistency Evaluation
por: Aharoni, Roee, et al.
Publicado: (2022)
por: Aharoni, Roee, et al.
Publicado: (2022)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
por: Zhang, Yue, et al.
Publicado: (2024)
por: Zhang, Yue, et al.
Publicado: (2024)
Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models
por: Song, Jongyoon, et al.
Publicado: (2024)
por: Song, Jongyoon, et al.
Publicado: (2024)
Inference Scaling for Bridging Retrieval and Augmented Generation
por: Lee, Youngwon, et al.
Publicado: (2024)
por: Lee, Youngwon, et al.
Publicado: (2024)
CORD: Balancing COnsistency and Rank Distillation for Robust Retrieval-Augmented Generation
por: Lee, Youngwon, et al.
Publicado: (2024)
por: Lee, Youngwon, et al.
Publicado: (2024)
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence
por: Joseph, Sebastian Antony, et al.
Publicado: (2024)
por: Joseph, Sebastian Antony, et al.
Publicado: (2024)
Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
por: Yang, Joonho, et al.
Publicado: (2025)
por: Yang, Joonho, et al.
Publicado: (2025)
Ejemplares similares
-
Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems
por: Kim, Minsoo, et al.
Publicado: (2025) -
NexusSum: Hierarchical LLM Agents for Long-Form Narrative Summarization
por: Kim, Hyuntak, et al.
Publicado: (2025) -
ECoRAG: Evidentiality-guided Compression for Long Context RAG
por: Jeong, Yeonseok, et al.
Publicado: (2025) -
CoEx -- Co-evolving World-model and Exploration
por: Kim, Minsoo, et al.
Publicado: (2025) -
CREFT: Sequential Multi-Agent LLM for Character Relation Extraction
por: Chun, Ye Eun, et al.
Publicado: (2025)