CEval: A Benchmark for Evaluating Counterfactual Text Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Van Bach, Schlötterer, Jörg, Seifert, Christin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
di: Nguyen, Van Bach, et al.
Pubblicazione: (2022)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2022)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
di: Nguyen, Van Bach, et al.
Pubblicazione: (2025)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2025)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
di: Cheng, Yinjie, et al.
Pubblicazione: (2025)
di: Cheng, Yinjie, et al.
Pubblicazione: (2025)
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
di: Wang, Qianli, et al.
Pubblicazione: (2026)
di: Wang, Qianli, et al.
Pubblicazione: (2026)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Enhancing Fact Retrieval in PLMs through Truthfulness
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Persuasion Tokens for Editing Factual Knowledge in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2026)
di: Youssef, Paul, et al.
Pubblicazione: (2026)
Investigating the Impact of Randomness on Reproducibility in Computer Vision: A Study on Applications in Civil Engineering and Medicine
di: Eryılmaz, Bahadır, et al.
Pubblicazione: (2024)
di: Eryılmaz, Bahadır, et al.
Pubblicazione: (2024)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
di: Koraş, Osman Alperen, et al.
Pubblicazione: (2024)
di: Koraş, Osman Alperen, et al.
Pubblicazione: (2024)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
di: Le, Phuong Quynh, et al.
Pubblicazione: (2024)
di: Le, Phuong Quynh, et al.
Pubblicazione: (2024)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Tracing and Reversing Edits in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2025)
di: Youssef, Paul, et al.
Pubblicazione: (2025)
Behavioral Analysis of Information Salience in Large Language Models
di: Trienes, Jan, et al.
Pubblicazione: (2025)
di: Trienes, Jan, et al.
Pubblicazione: (2025)
Position: Editing Large Language Models Poses Serious Safety Risks
di: Youssef, Paul, et al.
Pubblicazione: (2025)
di: Youssef, Paul, et al.
Pubblicazione: (2025)
Multi-Aspect Controllable Text Generation with Disentangled Counterfactual Augmentation
di: Liu, Yi, et al.
Pubblicazione: (2024)
di: Liu, Yi, et al.
Pubblicazione: (2024)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
di: Wang, Qianli, et al.
Pubblicazione: (2025)
di: Wang, Qianli, et al.
Pubblicazione: (2025)
Vietnamese AI Generated Text Detection
di: Tran, Quang-Dan, et al.
Pubblicazione: (2024)
di: Tran, Quang-Dan, et al.
Pubblicazione: (2024)
New Benchmark Dataset and Fine-Grained Cross-Modal Fusion Framework for Vietnamese Multimodal Aspect-Category Sentiment Analysis
di: Nguyen, Quy Hoang, et al.
Pubblicazione: (2024)
di: Nguyen, Quy Hoang, et al.
Pubblicazione: (2024)
Fùxì: A Benchmark for Evaluating Language Models on Ancient Chinese Text Understanding and Generation
di: Zhao, Shangqing, et al.
Pubblicazione: (2025)
di: Zhao, Shangqing, et al.
Pubblicazione: (2025)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
di: Idrissi-Yaghir, Ahmad, et al.
Pubblicazione: (2024)
di: Idrissi-Yaghir, Ahmad, et al.
Pubblicazione: (2024)
Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation
di: Ong, Keane, et al.
Pubblicazione: (2025)
di: Ong, Keane, et al.
Pubblicazione: (2025)
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
di: McAleese, Stephen, et al.
Pubblicazione: (2024)
di: McAleese, Stephen, et al.
Pubblicazione: (2024)
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
di: Trienes, Jan, et al.
Pubblicazione: (2025)
di: Trienes, Jan, et al.
Pubblicazione: (2025)
Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models
di: Lewis, Martha, et al.
Pubblicazione: (2024)
di: Lewis, Martha, et al.
Pubblicazione: (2024)
Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking
di: Nguyen, Hai Toan, et al.
Pubblicazione: (2025)
di: Nguyen, Hai Toan, et al.
Pubblicazione: (2025)
VN-MTEB: Vietnamese Massive Text Embedding Benchmark
di: Pham, Loc, et al.
Pubblicazione: (2025)
di: Pham, Loc, et al.
Pubblicazione: (2025)
Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
di: Zuo, Max, et al.
Pubblicazione: (2024)
di: Zuo, Max, et al.
Pubblicazione: (2024)
IMGTB: A Framework for Machine-Generated Text Detection Benchmarking
di: Spiegel, Michal, et al.
Pubblicazione: (2023)
di: Spiegel, Michal, et al.
Pubblicazione: (2023)
A Comparative Study of Quality Evaluation Methods for Text Summarization
di: Nguyen, Huyen, et al.
Pubblicazione: (2024)
di: Nguyen, Huyen, et al.
Pubblicazione: (2024)
Falcon: A Comprehensive Chinese Text-to-SQL Benchmark for Enterprise-Grade Evaluation
di: Luo, Wenzhen, et al.
Pubblicazione: (2025)
di: Luo, Wenzhen, et al.
Pubblicazione: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
di: Toker, Gilat, et al.
Pubblicazione: (2026)
di: Toker, Gilat, et al.
Pubblicazione: (2026)
GDGB: A Benchmark for Generative Dynamic Text-Attributed Graph Learning
di: Peng, Jie, et al.
Pubblicazione: (2025)
di: Peng, Jie, et al.
Pubblicazione: (2025)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
di: Macko, Dominik, et al.
Pubblicazione: (2024)
di: Macko, Dominik, et al.
Pubblicazione: (2024)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
di: Kashyap, Satyananda, et al.
Pubblicazione: (2025)
di: Kashyap, Satyananda, et al.
Pubblicazione: (2025)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
di: Jahan, Israt, et al.
Pubblicazione: (2023)
di: Jahan, Israt, et al.
Pubblicazione: (2023)
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
di: Zha, Yiwei, et al.
Pubblicazione: (2025)
di: Zha, Yiwei, et al.
Pubblicazione: (2025)
MILPaC: A Novel Benchmark for Evaluating Translation of Legal Text to Indian Languages
di: Mahapatra, Sayan, et al.
Pubblicazione: (2023)
di: Mahapatra, Sayan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024) -
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
di: Nguyen, Van Bach, et al.
Pubblicazione: (2022) -
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
di: Nguyen, Van Bach, et al.
Pubblicazione: (2025) -
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
di: Youssef, Paul, et al.
Pubblicazione: (2024) -
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
di: Cheng, Yinjie, et al.
Pubblicazione: (2025)