BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Fengqing, Feng, Yichen, Li, Yuetai, Niu, Luyao, Alomair, Basel, Poovendran, Radha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024)
by: Li, Yuetai, et al.
Published: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024)
by: Xiang, Zhen, et al.
Published: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
by: Jiang, Fengqing, et al.
Published: (2026)
by: Jiang, Fengqing, et al.
Published: (2026)
Polyhedral Instability Governs Regret in Online Learning
by: Li, Yuetai, et al.
Published: (2026)
by: Li, Yuetai, et al.
Published: (2026)
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
by: Sahabandu, Dinuka, et al.
Published: (2024)
by: Sahabandu, Dinuka, et al.
Published: (2024)
Defending Against Prompt Injection with DataFilter
by: Wang, Yizhu, et al.
Published: (2025)
by: Wang, Yizhu, et al.
Published: (2025)
PromptShield: Deployable Detection for Prompt Injection Attacks
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
Preventing Prompt Injection with Type-Directed Privilege Separation
by: Jacob, Dennis, et al.
Published: (2025)
by: Jacob, Dennis, et al.
Published: (2025)
CANTXSec: A Deterministic Intrusion Detection and Prevention System for CAN Bus Monitoring ECU Activations
by: Donadel, Denis, et al.
Published: (2025)
by: Donadel, Denis, et al.
Published: (2025)
SeedAIchemy: LLM-Driven Seed Corpus Generation for Fuzzing
by: Wen, Aidan, et al.
Published: (2025)
by: Wen, Aidan, et al.
Published: (2025)
Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
by: Abdali, Sara, et al.
Published: (2024)
by: Abdali, Sara, et al.
Published: (2024)
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
by: Piet, Julien, et al.
Published: (2025)
by: Piet, Julien, et al.
Published: (2025)
Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
by: Abdelnabi, Sahar, et al.
Published: (2026)
by: Abdelnabi, Sahar, et al.
Published: (2026)
Writing a Good Security Paper for ISSCC (2025)
by: Banerjee, Utsav, et al.
Published: (2025)
by: Banerjee, Utsav, et al.
Published: (2025)
Double-Dip: Thwarting Label-Only Membership Inference Attacks with Transfer Learning and Randomization
by: Rajabi, Arezoo, et al.
Published: (2024)
by: Rajabi, Arezoo, et al.
Published: (2024)
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy
by: Jeong, Joonhyun, et al.
Published: (2025)
by: Jeong, Joonhyun, et al.
Published: (2025)
Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis
by: Thornton, Scott
Published: (2026)
by: Thornton, Scott
Published: (2026)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
by: Piet, Julien, et al.
Published: (2023)
by: Piet, Julien, et al.
Published: (2023)
StyleFool: Fooling Video Classification Systems via Style Transfer
by: Cao, Yuxin, et al.
Published: (2022)
by: Cao, Yuxin, et al.
Published: (2022)
No More Hidden Pitfalls? Exposing Smart Contract Bad Practices with LLM-Powered Hybrid Analysis
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Deep-Research Agents Can Be Poisoned via User-Generated Content
by: Zhang, Tingwei, et al.
Published: (2026)
by: Zhang, Tingwei, et al.
Published: (2026)
Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection
by: Ghimire, Sujan, et al.
Published: (2026)
by: Ghimire, Sujan, et al.
Published: (2026)
Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars
by: Tehrani, Masoud Jamshidiyan, et al.
Published: (2026)
by: Tehrani, Masoud Jamshidiyan, et al.
Published: (2026)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
by: Qraitem, Maan, et al.
Published: (2024)
by: Qraitem, Maan, et al.
Published: (2024)
When and How to Fool Explainable Models (and Humans) with Adversarial Examples
by: Vadillo, Jon, et al.
Published: (2021)
by: Vadillo, Jon, et al.
Published: (2021)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
by: Feng, Yichen, et al.
Published: (2025)
by: Feng, Yichen, et al.
Published: (2025)
Bad Crypto: Chessography and Weak Randomness of Chess Games
by: Stanek, Martin
Published: (2024)
by: Stanek, Martin
Published: (2024)
Fooling SHAP with Output Shuffling Attacks
by: Yuan, Jun, et al.
Published: (2024)
by: Yuan, Jun, et al.
Published: (2024)
Revisiting DeepFool: generalization and improvement
by: Abdollahpoorrostam, Alireza, et al.
Published: (2023)
by: Abdollahpoorrostam, Alireza, et al.
Published: (2023)
Transparency Attacks: How Imperceptible Image Layers Can Fool AI Perception
by: McKee, Forrest, et al.
Published: (2024)
by: McKee, Forrest, et al.
Published: (2024)
Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents
by: Yu, Jay, et al.
Published: (2026)
by: Yu, Jay, et al.
Published: (2026)
Your Code Secret Belongs to Me: Neural Code Completion Tools Can Memorize Hard-Coded Credentials
by: Huang, Yizhan, et al.
Published: (2023)
by: Huang, Yizhan, et al.
Published: (2023)
Similar Items
-
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
by: Li, Yuetai, et al.
Published: (2024) -
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024) -
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024) -
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024) -
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)