Benchmarking of LLM Detection: Comparing Two Competing Approaches
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pröhl, Thorsten, Putzier, Erik, Zarnekow, Rüdiger |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
von: Moore, Robert J., et al.
Veröffentlicht: (2026)
von: Moore, Robert J., et al.
Veröffentlicht: (2026)
A Novel Psychometrics-Based Approach to Developing Professional Competency Benchmark for Large Language Models
von: Kardanova, Elena, et al.
Veröffentlicht: (2024)
von: Kardanova, Elena, et al.
Veröffentlicht: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
Bench4KE: Benchmarking Automated Competency Question Generation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025)
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
von: Luo, Zhimeng, et al.
Veröffentlicht: (2025)
von: Luo, Zhimeng, et al.
Veröffentlicht: (2025)
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
von: Lin, Peiqin, et al.
Veröffentlicht: (2026)
von: Lin, Peiqin, et al.
Veröffentlicht: (2026)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
von: Atasoy, I. F., et al.
Veröffentlicht: (2026)
von: Atasoy, I. F., et al.
Veröffentlicht: (2026)
Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark
von: Li, Zheqing, et al.
Veröffentlicht: (2025)
von: Li, Zheqing, et al.
Veröffentlicht: (2025)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
von: Wu, Zekun, et al.
Veröffentlicht: (2024)
von: Wu, Zekun, et al.
Veröffentlicht: (2024)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
von: Yu, Sungduk, et al.
Veröffentlicht: (2025)
von: Yu, Sungduk, et al.
Veröffentlicht: (2025)
LLM for Comparative Narrative Analysis
von: Kampen, Leo, et al.
Veröffentlicht: (2025)
von: Kampen, Leo, et al.
Veröffentlicht: (2025)
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection
von: Xu, Cheng, et al.
Veröffentlicht: (2026)
von: Xu, Cheng, et al.
Veröffentlicht: (2026)
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
von: Ayash, Lama, et al.
Veröffentlicht: (2025)
von: Ayash, Lama, et al.
Veröffentlicht: (2025)
CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation
von: Chen, Wei-Chun, et al.
Veröffentlicht: (2026)
von: Chen, Wei-Chun, et al.
Veröffentlicht: (2026)
HalluLens: LLM Hallucination Benchmark
von: Bang, Yejin, et al.
Veröffentlicht: (2025)
von: Bang, Yejin, et al.
Veröffentlicht: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
von: Ingimundarson, Finnur Ágúst, et al.
Veröffentlicht: (2026)
von: Ingimundarson, Finnur Ágúst, et al.
Veröffentlicht: (2026)
Confidence is Not Competence
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2026)
von: Hakimi, Ahmad Dawar, et al.
Veröffentlicht: (2026)
Multi-Lingual Cyber Threat Detection in Tweets/X Using ML, DL, and LLM: A Comparative Analysis
von: Murad, Saydul Akbar, et al.
Veröffentlicht: (2025)
von: Murad, Saydul Akbar, et al.
Veröffentlicht: (2025)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
von: Plaza, Irene, et al.
Veröffentlicht: (2024)
von: Plaza, Irene, et al.
Veröffentlicht: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
von: Tamber, Manveer Singh, et al.
Veröffentlicht: (2025)
Benchmark of stylistic variation in LLM-generated texts
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
Benchmarking and Improving LLM Robustness for Personalized Generation
von: Okite, Chimaobi, et al.
Veröffentlicht: (2025)
von: Okite, Chimaobi, et al.
Veröffentlicht: (2025)
Benchmarking Advanced Text Anonymisation Methods: A Comparative Study on Novel and Traditional Approaches
von: Asimopoulos, Dimitris, et al.
Veröffentlicht: (2024)
von: Asimopoulos, Dimitris, et al.
Veröffentlicht: (2024)
Benchmark Test-Time Scaling of General LLM Agents
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
Robust Planning with Compound LLM Architectures: An LLM-Modulo Approach
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
von: Luo, Wen, et al.
Veröffentlicht: (2026)
von: Luo, Wen, et al.
Veröffentlicht: (2026)
Two-dimensional early exit optimisation of LLM inference
von: Hůla, Jan, et al.
Veröffentlicht: (2026)
von: Hůla, Jan, et al.
Veröffentlicht: (2026)
keepitsimple at SemEval-2025 Task 3: LLM-Uncertainty based Approach for Multilingual Hallucination Span Detection
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
von: Noever, David, et al.
Veröffentlicht: (2025)
von: Noever, David, et al.
Veröffentlicht: (2025)
The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
von: Moore, Kyle, et al.
Veröffentlicht: (2024)
From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
von: Cunegatti, Elia, et al.
Veröffentlicht: (2026)
von: Cunegatti, Elia, et al.
Veröffentlicht: (2026)
Comparing Hallucination Detection Metrics for Multilingual Generation
von: Kang, Haoqiang, et al.
Veröffentlicht: (2024)
von: Kang, Haoqiang, et al.
Veröffentlicht: (2024)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
von: Moore, Robert J., et al.
Veröffentlicht: (2026) -
A Novel Psychometrics-Based Approach to Developing Professional Competency Benchmark for Large Language Models
von: Kardanova, Elena, et al.
Veröffentlicht: (2024) -
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024) -
Bench4KE: Benchmarking Automated Competency Question Generation
von: Lippolis, Anna Sofia, et al.
Veröffentlicht: (2025) -
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
von: Yang, Tiankai, et al.
Veröffentlicht: (2024)