Who's Your Judge? On the Detectability of LLM-Generated Judgments
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Dawei, Tan, Zhen, Zhao, Chengshuai, Jiang, Bohan, Huang, Baixiang, Ma, Pingchuan, Alnaibari, Abdullah, Shu, Kai, Liu, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
Catching Chameleons: Detecting Evolving Disinformation Generated using Large Language Models
by: Jiang, Bohan, et al.
Published: (2024)
by: Jiang, Bohan, et al.
Published: (2024)
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Are Today's LLMs Ready to Explain Well-Being Concepts?
by: Jiang, Bohan, et al.
Published: (2025)
by: Jiang, Bohan, et al.
Published: (2025)
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model
by: Zhao, Chengshuai, et al.
Published: (2026)
by: Zhao, Chengshuai, et al.
Published: (2026)
From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification
by: Ma, Pingchuan, et al.
Published: (2025)
by: Ma, Pingchuan, et al.
Published: (2025)
Exploring Large Language Models for Feature Selection: A Data-centric Perspective
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
by: Chen, Luyu, et al.
Published: (2025)
by: Chen, Luyu, et al.
Published: (2025)
Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion
by: Tan, Zhen, et al.
Published: (2026)
by: Tan, Zhen, et al.
Published: (2026)
SCALE: Towards Collaborative Content Analysis in Social Science with Large Language Model Agents and Human Intervention
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
by: Hu, Tianyu, et al.
Published: (2025)
by: Hu, Tianyu, et al.
Published: (2025)
"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Joint Detection of Fraud and Concept Drift inOnline Conversations with LLM-Assisted Judgment
by: Senol, Ali, et al.
Published: (2025)
by: Senol, Ali, et al.
Published: (2025)
Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm
by: Huang, Baixiang, et al.
Published: (2025)
by: Huang, Baixiang, et al.
Published: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
Contextualization Distillation from Large Language Model for Knowledge Graph Completion
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2025)
by: Li, Dawei, et al.
Published: (2025)
SST: Multi-Scale Hybrid Mamba-Transformer Experts for Time Series Forecasting
by: Xu, Xiongxiao, et al.
Published: (2024)
by: Xu, Xiongxiao, et al.
Published: (2024)
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
by: Li, Yuran, et al.
Published: (2025)
by: Li, Yuran, et al.
Published: (2025)
Ontology-Aware RAG for Improved Question-Answering in Cybersecurity Education
by: Zhao, Chengshuai, et al.
Published: (2024)
by: Zhao, Chengshuai, et al.
Published: (2024)
A Judge Agent Closes the Reliability Gap in AI-Generated Scientific Simulation
by: Yang, Chengshuai
Published: (2026)
by: Yang, Chengshuai
Published: (2026)
Can LLM-Generated Misinformation Be Detected?
by: Chen, Canyu, et al.
Published: (2023)
by: Chen, Canyu, et al.
Published: (2023)
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
Authorship Attribution in the Era of LLMs: Problems, Methodologies, and Challenges
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
Can Large Language Models Identify Authorship?
by: Huang, Baixiang, et al.
Published: (2024)
by: Huang, Baixiang, et al.
Published: (2024)
Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation
by: Xu, Zonghuan, et al.
Published: (2026)
by: Xu, Zonghuan, et al.
Published: (2026)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
Your One-Stop Solution for AI-Generated Video Detection
by: Ma, Long, et al.
Published: (2026)
by: Ma, Long, et al.
Published: (2026)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
Privacy-Aware Decoding: Mitigating Privacy Leakage of Large Language Models in Retrieval-Augmented Generation
by: Wang, Haoran, et al.
Published: (2025)
by: Wang, Haoran, et al.
Published: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
Intrinsic Barriers to Explaining Deep Foundation Models
by: Tan, Zhen, et al.
Published: (2025)
by: Tan, Zhen, et al.
Published: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
by: Li, Pingzhi, et al.
Published: (2025)
by: Li, Pingzhi, et al.
Published: (2025)
Similar Items
-
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
by: Li, Dawei, et al.
Published: (2024) -
Catching Chameleons: Detecting Evolving Disinformation Generated using Large Language Models
by: Jiang, Bohan, et al.
Published: (2024) -
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
by: Zhao, Chengshuai, et al.
Published: (2025) -
Are Today's LLMs Ready to Explain Well-Being Concepts?
by: Jiang, Bohan, et al.
Published: (2025) -
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model
by: Zhao, Chengshuai, et al.
Published: (2026)