An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
Fuente:
arXiv
Saved in:
| Main Authors: | Schoenegger, Loris, Xia, Yuxi, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compact Example-Based Explanations for Language Models
by: Schoenegger, Loris, et al.
Published: (2026)
by: Schoenegger, Loris, et al.
Published: (2026)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
by: Hinterleitner, Lukas, et al.
Published: (2026)
by: Hinterleitner, Lukas, et al.
Published: (2026)
Influence-driven Curriculum Learning for Pre-training on Limited Data
by: Schoenegger, Loris, et al.
Published: (2025)
by: Schoenegger, Loris, et al.
Published: (2025)
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
by: Xia, Yuxi, et al.
Published: (2026)
by: Xia, Yuxi, et al.
Published: (2026)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods
by: Doshi, Jai, et al.
Published: (2024)
by: Doshi, Jai, et al.
Published: (2024)
RKadiyala at SemEval-2024 Task 8: Black-Box Word-Level Text Boundary Detection in Partially Machine Generated Texts
by: Kadiyala, Ram Mohan Rao
Published: (2024)
by: Kadiyala, Ram Mohan Rao
Published: (2024)
Hierarchical Text Classification Using Black Box Large Language Models
by: Yoshimura, Kosuke, et al.
Published: (2025)
by: Yoshimura, Kosuke, et al.
Published: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Exploring prompts to elicit memorization in masked language model-based named entity recognition
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
by: Sheng, Haotian, et al.
Published: (2026)
by: Sheng, Haotian, et al.
Published: (2026)
Does It Make Sense to Explain a Black Box With Another Black Box?
by: Delaunay, Julien, et al.
Published: (2024)
by: Delaunay, Julien, et al.
Published: (2024)
Transparent Neighborhood Approximation for Text Classifier Explanation
by: Cai, Yi, et al.
Published: (2024)
by: Cai, Yi, et al.
Published: (2024)
RAFT: Realistic Attacks to Fool Text Detectors
by: Wang, James, et al.
Published: (2024)
by: Wang, James, et al.
Published: (2024)
Bounded Behavioral Indistinguishability for Black-Box LLM Distillation
by: Hasan, Munawar
Published: (2026)
by: Hasan, Munawar
Published: (2026)
Detection of Machine-Generated Text: Literature Survey
by: Valiaiev, Dmytro
Published: (2024)
by: Valiaiev, Dmytro
Published: (2024)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
by: Chen, Xiwen, et al.
Published: (2026)
by: Chen, Xiwen, et al.
Published: (2026)
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
by: Chang, Shuochen, et al.
Published: (2026)
by: Chang, Shuochen, et al.
Published: (2026)
SafePassage: High-Fidelity Information Extraction with Black Box LLMs
by: Barrow, Joe, et al.
Published: (2025)
by: Barrow, Joe, et al.
Published: (2025)
Black-box Model Ensembling for Textual and Visual Question Answering via Information Fusion
by: Xia, Yuxi, et al.
Published: (2024)
by: Xia, Yuxi, et al.
Published: (2024)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Ten Words Only Still Help: Improving Black-Box AI-Generated Text Detection via Proxy-Guided Efficient Re-Sampling
by: Shi, Yuhui, et al.
Published: (2024)
by: Shi, Yuhui, et al.
Published: (2024)
Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA
by: Jawad, Hussein, et al.
Published: (2024)
by: Jawad, Hussein, et al.
Published: (2024)
A Watermark for Black-Box Language Models
by: Bahri, Dara, et al.
Published: (2024)
by: Bahri, Dara, et al.
Published: (2024)
Detecting AI Generated Text Based on NLP and Machine Learning Approaches
by: Prova, Nuzhat
Published: (2024)
by: Prova, Nuzhat
Published: (2024)
Few-Shot Detection of Machine-Generated Text using Style Representations
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
DALD: Improving Logits-based Detector without Logits from Black-box LLMs
by: Zeng, Cong, et al.
Published: (2024)
by: Zeng, Cong, et al.
Published: (2024)
Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
by: Refai, Dania, et al.
Published: (2025)
by: Refai, Dania, et al.
Published: (2025)
TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
by: Miok, Kristian, et al.
Published: (2025)
by: Miok, Kristian, et al.
Published: (2025)
Beating the Style Detector: Three Hours of Agentic Research on the AI-Text Arms Race
by: Maier, Andreas, et al.
Published: (2026)
by: Maier, Andreas, et al.
Published: (2026)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
Training Deliberative Monitors for Black-Box Scheming Detection
by: Sinha, Aditya, et al.
Published: (2026)
by: Sinha, Aditya, et al.
Published: (2026)
Methods for Generating Drift in Text Streams
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
Topic Modelling Black Box Optimization
by: Akramov, Roman, et al.
Published: (2025)
by: Akramov, Roman, et al.
Published: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Label-Efficient Model Selection for Text Generation
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
Perfect diffusion is $\mathsf{TC}^0$ -- Bad diffusion is Turing-complete
by: Liu, Yuxi
Published: (2025)
by: Liu, Yuxi
Published: (2025)
Similar Items
-
Compact Example-Based Explanations for Language Models
by: Schoenegger, Loris, et al.
Published: (2026) -
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
by: Xia, Yuxi, et al.
Published: (2026) -
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
by: Hinterleitner, Lukas, et al.
Published: (2026) -
Influence-driven Curriculum Learning for Pre-training on Limited Data
by: Schoenegger, Loris, et al.
Published: (2025) -
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
by: Xia, Yuxi, et al.
Published: (2026)