VERISCORE: Evaluating the factuality of verifiable claims in long-form text generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Yixiao, Kim, Yekyung, Iyyer, Mohit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
One ruler to measure them all: Benchmarking multilingual long-context language models
von: Kim, Yekyung, et al.
Veröffentlicht: (2025)
von: Kim, Yekyung, et al.
Veröffentlicht: (2025)
VeriFastScore: Speeding up long-form factuality evaluation
von: Rajendhran, Rishanth, et al.
Veröffentlicht: (2025)
von: Rajendhran, Rishanth, et al.
Veröffentlicht: (2025)
Frankentext: Stitching random text fragments into long-form narratives
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025)
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025)
Localizing and Mitigating Errors in Long-form Question Answering
von: Sachdeva, Rachneet, et al.
Veröffentlicht: (2024)
von: Sachdeva, Rachneet, et al.
Veröffentlicht: (2024)
Argument Collapse: LLMs Flatten Long-Form Public Debate
von: Kim, Yekyung, et al.
Veröffentlicht: (2026)
von: Kim, Yekyung, et al.
Veröffentlicht: (2026)
Does quantization affect models' performance on long-context tasks?
von: Mekala, Anmol, et al.
Veröffentlicht: (2025)
von: Mekala, Anmol, et al.
Veröffentlicht: (2025)
CLIPPER: Compression enables long-context synthetic data generation
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025)
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025)
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
FABLES: Evaluating faithfulness and content selection in book-length summarization
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)
von: Kim, Yekyung, et al.
Veröffentlicht: (2024)
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
von: Chang, Yapei, et al.
Veröffentlicht: (2025)
von: Chang, Yapei, et al.
Veröffentlicht: (2025)
Suri: Multi-constraint Instruction Following for Long-form Text Generation
von: Pham, Chau Minh, et al.
Veröffentlicht: (2024)
von: Pham, Chau Minh, et al.
Veröffentlicht: (2024)
Literary Evidence Retrieval via Long-Context Language Models
von: Thai, Katherine, et al.
Veröffentlicht: (2025)
von: Thai, Katherine, et al.
Veröffentlicht: (2025)
Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
von: Jafari, Nazanin, et al.
Veröffentlicht: (2026)
von: Jafari, Nazanin, et al.
Veröffentlicht: (2026)
Whose story is it? Personalizing story generation by inferring author styles
von: Kumar, Nischal Ashok, et al.
Veröffentlicht: (2025)
von: Kumar, Nischal Ashok, et al.
Veröffentlicht: (2025)
CaLMQA: Exploring culturally specific long-form question answering across 23 languages
von: Arora, Shane, et al.
Veröffentlicht: (2024)
von: Arora, Shane, et al.
Veröffentlicht: (2024)
Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs
von: Samuel, Vinay, et al.
Veröffentlicht: (2026)
von: Samuel, Vinay, et al.
Veröffentlicht: (2026)
BEARCUBS: A benchmark for computer-using web agents
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
One Thousand and One Pairs: A "novel" challenge for long-context language models
von: Karpinska, Marzena, et al.
Veröffentlicht: (2024)
von: Karpinska, Marzena, et al.
Veröffentlicht: (2024)
Long-form factuality in large language models
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
Measuring short-form factuality in large language models
von: Wei, Jason, et al.
Veröffentlicht: (2024)
von: Wei, Jason, et al.
Veröffentlicht: (2024)
Measuring text summarization factuality using atomic facts entailment metrics in the context of retrieval augmented generation
von: Kriman, N. E.
Veröffentlicht: (2024)
von: Kriman, N. E.
Veröffentlicht: (2024)
EditLens: Quantifying the Extent of AI Editing in Text
von: Thai, Katherine, et al.
Veröffentlicht: (2025)
von: Thai, Katherine, et al.
Veröffentlicht: (2025)
Contextualized Evaluations: Judging Language Model Responses to Underspecified Queries
von: Malaviya, Chaitanya, et al.
Veröffentlicht: (2024)
von: Malaviya, Chaitanya, et al.
Veröffentlicht: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
von: Naseh, Ali, et al.
Veröffentlicht: (2024)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
von: Chang, Yapei, et al.
Veröffentlicht: (2023)
von: Chang, Yapei, et al.
Veröffentlicht: (2023)
TopicGPT: A Prompt-based Topic Modeling Framework
von: Pham, Chau Minh, et al.
Veröffentlicht: (2023)
von: Pham, Chau Minh, et al.
Veröffentlicht: (2023)
StoryScope: Investigating idiosyncrasies in AI fiction
von: Russell, Jenna, et al.
Veröffentlicht: (2026)
von: Russell, Jenna, et al.
Veröffentlicht: (2026)
Collaborative decoding of critical tokens for boosting factuality of large language models
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
von: Kotek, Hadas, et al.
Veröffentlicht: (2026)
von: Kotek, Hadas, et al.
Veröffentlicht: (2026)
Synthetically generated text for supervised text analysis
von: Halterman, Andrew
Veröffentlicht: (2023)
von: Halterman, Andrew
Veröffentlicht: (2023)
AI use in American newspapers is widespread, uneven, and rarely disclosed
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
von: Russell, Jenna, et al.
Veröffentlicht: (2025)
PostMark: A Robust Blackbox Watermark for Large Language Models
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
von: Chang, Yapei, et al.
Veröffentlicht: (2024)
CECOR: Correction-oriented synthetic data construction for factual error correction
von: Zhu, Lei, et al.
Veröffentlicht: (2026)
von: Zhu, Lei, et al.
Veröffentlicht: (2026)
Qwen it detect machine-generated text?
von: Marchitan, Teodor-George, et al.
Veröffentlicht: (2025)
von: Marchitan, Teodor-George, et al.
Veröffentlicht: (2025)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
von: Lee, Yukyung, et al.
Veröffentlicht: (2024)
von: Lee, Yukyung, et al.
Veröffentlicht: (2024)
Interactive Topic Models with Optimal Transport
von: Dhanania, Garima, et al.
Veröffentlicht: (2024)
von: Dhanania, Garima, et al.
Veröffentlicht: (2024)
CIDER: Context sensitive sentiment analysis for short-form text
von: Young, James C., et al.
Veröffentlicht: (2023)
von: Young, James C., et al.
Veröffentlicht: (2023)
AI-generated text boundary detection with RoFT
von: Kushnareva, Laida, et al.
Veröffentlicht: (2023)
von: Kushnareva, Laida, et al.
Veröffentlicht: (2023)
Diversify-verify-adapt: Efficient and Robust Retrieval-Augmented Ambiguous Question Answering
von: In, Yeonjun, et al.
Veröffentlicht: (2024)
von: In, Yeonjun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
One ruler to measure them all: Benchmarking multilingual long-context language models
von: Kim, Yekyung, et al.
Veröffentlicht: (2025) -
VeriFastScore: Speeding up long-form factuality evaluation
von: Rajendhran, Rishanth, et al.
Veröffentlicht: (2025) -
Frankentext: Stitching random text fragments into long-form narratives
von: Pham, Chau Minh, et al.
Veröffentlicht: (2025) -
Localizing and Mitigating Errors in Long-form Question Answering
von: Sachdeva, Rachneet, et al.
Veröffentlicht: (2024) -
Argument Collapse: LLMs Flatten Long-Form Public Debate
von: Kim, Yekyung, et al.
Veröffentlicht: (2026)