Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Zonghuan, Zheng, Xiang, Wu, Yutao, Ma, Xingjun |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Order to Distribution: A Spectral Characterization of Forgetting in Continual Learning
par: Xu, Zonghuan, et autres
Publié: (2026)
par: Xu, Zonghuan, et autres
Publié: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
par: Xu, Zonghuan, et autres
Publié: (2025)
par: Xu, Zonghuan, et autres
Publié: (2025)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
par: Li, Jiayu, et autres
Publié: (2025)
par: Li, Jiayu, et autres
Publié: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
par: Wu, Yutao, et autres
Publié: (2025)
par: Wu, Yutao, et autres
Publié: (2025)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
par: Zugecova, Aneta, et autres
Publié: (2024)
par: Zugecova, Aneta, et autres
Publié: (2024)
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
par: Zheng, Han, et autres
Publié: (2026)
par: Zheng, Han, et autres
Publié: (2026)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
par: Zheng, Xiang, et autres
Publié: (2026)
par: Zheng, Xiang, et autres
Publié: (2026)
Who's Your Judge? On the Detectability of LLM-Generated Judgments
par: Li, Dawei, et autres
Publié: (2025)
par: Li, Dawei, et autres
Publié: (2025)
Constrained Intrinsic Motivation for Reinforcement Learning
par: Zheng, Xiang, et autres
Publié: (2024)
par: Zheng, Xiang, et autres
Publié: (2024)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
par: Zhao, Yunhan, et autres
Publié: (2025)
par: Zhao, Yunhan, et autres
Publié: (2025)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
par: Chen, Luyu, et autres
Publié: (2025)
par: Chen, Luyu, et autres
Publié: (2025)
Hypothesis Generation via LLM-Automated Language Bias for ILP
par: Yang, Yang, et autres
Publié: (2025)
par: Yang, Yang, et autres
Publié: (2025)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
par: Feng, Yunhao, et autres
Publié: (2026)
par: Feng, Yunhao, et autres
Publié: (2026)
Journalists' Perceptions of Artificial Intelligence and Disinformation Risks
par: Peña-Alonso, Urko, et autres
Publié: (2025)
par: Peña-Alonso, Urko, et autres
Publié: (2025)
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
par: Polo, Felipe Maia, et autres
Publié: (2025)
par: Polo, Felipe Maia, et autres
Publié: (2025)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
par: Wu, Yutao, et autres
Publié: (2025)
par: Wu, Yutao, et autres
Publié: (2025)
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
par: Wang, Yuzhe, et autres
Publié: (2026)
par: Wang, Yuzhe, et autres
Publié: (2026)
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
par: Li, Yuran, et autres
Publié: (2025)
par: Li, Yuran, et autres
Publié: (2025)
Implicit Humanization in Everyday LLM Moral Judgments
par: Ayad, Hoda, et autres
Publié: (2026)
par: Ayad, Hoda, et autres
Publié: (2026)
Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer
par: Xiang, Liangru, et autres
Publié: (2026)
par: Xiang, Liangru, et autres
Publié: (2026)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
par: Feng, Yunhao, et autres
Publié: (2026)
par: Feng, Yunhao, et autres
Publié: (2026)
Evaluating the Propensity of Generative AI for Producing Harmful Disinformation During the 2024 US Election Cycle
par: Schlicht, Erik J
Publié: (2024)
par: Schlicht, Erik J
Publié: (2024)
QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI
par: Veysi, Marjan, et autres
Publié: (2026)
par: Veysi, Marjan, et autres
Publié: (2026)
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
par: Li, Dawei, et autres
Publié: (2024)
par: Li, Dawei, et autres
Publié: (2024)
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
par: Tian, Qiuyu, et autres
Publié: (2026)
par: Tian, Qiuyu, et autres
Publié: (2026)
Evaluating Steering Techniques using Human Similarity Judgments
par: Studdiford, Zach, et autres
Publié: (2025)
par: Studdiford, Zach, et autres
Publié: (2025)
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints
par: Yin, Zhenyun, et autres
Publié: (2025)
par: Yin, Zhenyun, et autres
Publié: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
par: Ding, Jiale, et autres
Publié: (2025)
par: Ding, Jiale, et autres
Publié: (2025)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
par: Richardson, Andrew Keenan, et autres
Publié: (2025)
par: Richardson, Andrew Keenan, et autres
Publié: (2025)
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models
par: Li, Yige, et autres
Publié: (2024)
par: Li, Yige, et autres
Publié: (2024)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
par: Watanabe, Yusuke, et autres
Publié: (2026)
par: Watanabe, Yusuke, et autres
Publié: (2026)
Internal Safety Collapse in Frontier Large Language Models
par: Wu, Yutao, et autres
Publié: (2026)
par: Wu, Yutao, et autres
Publié: (2026)
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
par: Dong, Zihan, et autres
Publié: (2026)
par: Dong, Zihan, et autres
Publié: (2026)
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
par: Park, Shinwoo, et autres
Publié: (2026)
par: Park, Shinwoo, et autres
Publié: (2026)
CALM: Curiosity-Driven Auditing for Large Language Models
par: Zheng, Xiang, et autres
Publié: (2025)
par: Zheng, Xiang, et autres
Publié: (2025)
The Veln(ia)s is in the Details: Evaluating LLM Judgment on Latvian and Lithuanian Short Answer Matching
par: Kostiuk, Yevhen, et autres
Publié: (2025)
par: Kostiuk, Yevhen, et autres
Publié: (2025)
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
par: Hong, Chang, et autres
Publié: (2025)
par: Hong, Chang, et autres
Publié: (2025)
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
par: Sun, Chenkai, et autres
Publié: (2025)
par: Sun, Chenkai, et autres
Publié: (2025)
Safeguarding Marketing Research: The Generation, Identification, and Mitigation of AI-Fabricated Disinformation
par: Mukherjee, Anirban
Publié: (2024)
par: Mukherjee, Anirban
Publié: (2024)
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models
par: Zheng, Baihui, et autres
Publié: (2025)
par: Zheng, Baihui, et autres
Publié: (2025)
Documents similaires
-
From Order to Distribution: A Spectral Characterization of Forgetting in Continual Learning
par: Xu, Zonghuan, et autres
Publié: (2026) -
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
par: Xu, Zonghuan, et autres
Publié: (2025) -
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
par: Li, Jiayu, et autres
Publié: (2025) -
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
par: Wu, Yutao, et autres
Publié: (2025) -
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
par: Zugecova, Aneta, et autres
Publié: (2024)