M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yuxia, Mansurov, Jonibek, Ivanov, Petar, Su, Jinyan, Shelmanov, Artem, Tsvigun, Akim, Afzal, Osama Mohanned, Mahmoud, Tarek, Puccetti, Giovanni, Arnold, Thomas, Aji, Alham Fikri, Habash, Nizar, Gurevych, Iryna, Nakov, Preslav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
di: Wang, Yuxia, et al.
Pubblicazione: (2024)
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
di: Mansurov, Jonibek, et al.
Pubblicazione: (2024)
di: Mansurov, Jonibek, et al.
Pubblicazione: (2024)
GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human
di: Wang, Yuxia, et al.
Pubblicazione: (2025)
di: Wang, Yuxia, et al.
Pubblicazione: (2025)
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection
di: Abassy, Mervat, et al.
Pubblicazione: (2024)
di: Abassy, Mervat, et al.
Pubblicazione: (2024)
Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI
di: Wang, Yuxia, et al.
Pubblicazione: (2025)
di: Wang, Yuxia, et al.
Pubblicazione: (2025)
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
di: Afzal, Osama Mohammed, et al.
Pubblicazione: (2025)
di: Afzal, Osama Mohammed, et al.
Pubblicazione: (2025)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
di: Elshabrawy, Ahmed, et al.
Pubblicazione: (2024)
di: Elshabrawy, Ahmed, et al.
Pubblicazione: (2024)
A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM Outputs
di: Shelmanov, Artem, et al.
Pubblicazione: (2025)
di: Shelmanov, Artem, et al.
Pubblicazione: (2025)
From Surveys to Narratives: Rethinking Cultural Value Adaptation in LLMs
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2025)
di: Adilazuarda, Muhammad Farid, et al.
Pubblicazione: (2025)
KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan
di: Togmanov, Mukhammed, et al.
Pubblicazione: (2025)
di: Togmanov, Mukhammed, et al.
Pubblicazione: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
di: Altakrori, Malik H., et al.
Pubblicazione: (2025)
di: Altakrori, Malik H., et al.
Pubblicazione: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
di: Orel, Daniil, et al.
Pubblicazione: (2026)
di: Orel, Daniil, et al.
Pubblicazione: (2026)
A Survey of Confidence Estimation and Calibration in Large Language Models
di: Geng, Jiahui, et al.
Pubblicazione: (2023)
di: Geng, Jiahui, et al.
Pubblicazione: (2023)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
di: Bates, Luke, et al.
Pubblicazione: (2025)
di: Bates, Luke, et al.
Pubblicazione: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
di: Glockner, Max, et al.
Pubblicazione: (2024)
di: Glockner, Max, et al.
Pubblicazione: (2024)
Multimodal Large Language Models to Support Real-World Fact-Checking
di: Geng, Jiahui, et al.
Pubblicazione: (2024)
di: Geng, Jiahui, et al.
Pubblicazione: (2024)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
di: Orel, Daniil, et al.
Pubblicazione: (2025)
di: Orel, Daniil, et al.
Pubblicazione: (2025)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
di: Glockner, Max, et al.
Pubblicazione: (2024)
di: Glockner, Max, et al.
Pubblicazione: (2024)
LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages
di: Aji, Alham Fikri, et al.
Pubblicazione: (2025)
di: Aji, Alham Fikri, et al.
Pubblicazione: (2025)
Improving Low-Resource Machine Translation via Round-Trip Reinforcement Learning
di: Attia, Ahmed, et al.
Pubblicazione: (2026)
di: Attia, Ahmed, et al.
Pubblicazione: (2026)
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition
di: Chevi, Rendi, et al.
Pubblicazione: (2024)
di: Chevi, Rendi, et al.
Pubblicazione: (2024)
A Template Is All You Meme
di: Bates, Luke, et al.
Pubblicazione: (2023)
di: Bates, Luke, et al.
Pubblicazione: (2023)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
di: Geng, Jiahui, et al.
Pubblicazione: (2025)
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs
di: Iqbal, Hasan, et al.
Pubblicazione: (2024)
di: Iqbal, Hasan, et al.
Pubblicazione: (2024)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
di: Sahnan, Dhruv, et al.
Pubblicazione: (2026)
Detecting Check-Worthy Claims in Political Debates, Speeches, and Interviews Using Audio Data
di: Ivanov, Petar, et al.
Pubblicazione: (2023)
di: Ivanov, Petar, et al.
Pubblicazione: (2023)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
di: Kaneko, Masahiro, et al.
Pubblicazione: (2025)
Sense Representations Are Inducible Interfaces
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
LLM Olympiad: Why Model Evaluation Needs a Sealed Exam
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2026)
Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2025)
di: Cruz, Jan Christian Blaise, et al.
Pubblicazione: (2025)
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models
di: Lyu, Chenyang, et al.
Pubblicazione: (2024)
di: Lyu, Chenyang, et al.
Pubblicazione: (2024)
FIRE: Fact-checking with Iterative Retrieval and Verification
di: Xie, Zhuohan, et al.
Pubblicazione: (2024)
di: Xie, Zhuohan, et al.
Pubblicazione: (2024)
Adapting Fake News Detection to the Era of Large Language Models
di: Su, Jinyan, et al.
Pubblicazione: (2023)
di: Su, Jinyan, et al.
Pubblicazione: (2023)
Corpus Poisoning via Approximate Greedy Gradient Descent
di: Su, Jinyan, et al.
Pubblicazione: (2024)
di: Su, Jinyan, et al.
Pubblicazione: (2024)
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
di: Vashurin, Roman, et al.
Pubblicazione: (2024)
di: Vashurin, Roman, et al.
Pubblicazione: (2024)
Rethinking STS and NLI in Large Language Models
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
di: Wang, Yuxia, et al.
Pubblicazione: (2023)
From Multiple-Choice to Extractive QA: A Case Study for English and Arabic
di: Lynn, Teresa, et al.
Pubblicazione: (2024)
di: Lynn, Teresa, et al.
Pubblicazione: (2024)
Language-Specific Latent Process Hinders Cross-Lingual Performance
di: Lim, Zheng Wei, et al.
Pubblicazione: (2025)
di: Lim, Zheng Wei, et al.
Pubblicazione: (2025)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection
di: Wang, Yuxia, et al.
Pubblicazione: (2024) -
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
di: Wang, Yuxia, et al.
Pubblicazione: (2023) -
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
di: Mansurov, Jonibek, et al.
Pubblicazione: (2024) -
GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human
di: Wang, Yuxia, et al.
Pubblicazione: (2025) -
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection
di: Abassy, Mervat, et al.
Pubblicazione: (2024)