OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
Fuente:
arXiv
Saved in:
| Main Authors: | Koike, Ryuto, Kaneko, Masahiro, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
by: Koike, Ryuto, et al.
Published: (2025)
by: Koike, Ryuto, et al.
Published: (2025)
LLM Output Detectability and Task Performance Can be Jointly Optimized
by: Saito, Koshiro, et al.
Published: (2026)
by: Saito, Koshiro, et al.
Published: (2026)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024)
by: Oi, Masanari, et al.
Published: (2024)
Machine Text Detectors are Membership Inference Attacks
by: Koike, Ryuto, et al.
Published: (2025)
by: Koike, Ryuto, et al.
Published: (2025)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
by: Loem, Mengsay, et al.
Published: (2023)
by: Loem, Mengsay, et al.
Published: (2023)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026)
by: Ichinose, Tatsuya, et al.
Published: (2026)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
by: Kaneko, Masahiro, et al.
Published: (2023)
by: Kaneko, Masahiro, et al.
Published: (2023)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
LCTG Bench: LLM Controlled Text Generation Benchmark
by: Kurihara, Kentaro, et al.
Published: (2025)
by: Kurihara, Kentaro, et al.
Published: (2025)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
by: Michail, Andrianos, et al.
Published: (2025)
by: Michail, Andrianos, et al.
Published: (2025)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
by: Ma, Youmi, et al.
Published: (2024)
by: Ma, Youmi, et al.
Published: (2024)
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
by: Oi, Masanari, et al.
Published: (2026)
by: Oi, Masanari, et al.
Published: (2026)
Decoding-Free Sampling Strategies for LLM Marginalization
by: Pohl, David, et al.
Published: (2025)
by: Pohl, David, et al.
Published: (2025)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Hidding the Ghostwriters: An Adversarial Evaluation of AI-Generated Student Essay Detection
by: Peng, Xinlin, et al.
Published: (2024)
by: Peng, Xinlin, et al.
Published: (2024)
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
by: Przybyła, Piotr, et al.
Published: (2024)
by: Przybyła, Piotr, et al.
Published: (2024)
Drifting Objectives for Refining Discrete Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Two Counterexamples to Tokenization and the Noiseless Channel
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
by: Miyamoto, Sora, et al.
Published: (2026)
by: Miyamoto, Sora, et al.
Published: (2026)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
by: Kaneko, Masahiro, et al.
Published: (2026)
by: Kaneko, Masahiro, et al.
Published: (2026)
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
by: Sarukkai, Vishnu, et al.
Published: (2025)
by: Sarukkai, Vishnu, et al.
Published: (2025)
LLM-GAN: Construct Generative Adversarial Network Through Large Language Models For Explainable Fake News Detection
by: Wang, Yifeng, et al.
Published: (2024)
by: Wang, Yifeng, et al.
Published: (2024)
Rectifying Adversarial Examples Using Their Vulnerabilities
by: Morimoto, Fumiya, et al.
Published: (2026)
by: Morimoto, Fumiya, et al.
Published: (2026)
Similar Items
-
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023) -
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
by: Koike, Ryuto, et al.
Published: (2025) -
LLM Output Detectability and Task Performance Can be Jointly Optimized
by: Saito, Koshiro, et al.
Published: (2026) -
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024) -
Machine Text Detectors are Membership Inference Attacks
by: Koike, Ryuto, et al.
Published: (2025)