Generative Adversarial Reviews: When LLMs Become the Critic
Fuente:
arXiv
Saved in:
| Main Authors: | Bougie, Nicolas, Watanabe, Narimasa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation
by: Bougie, Nicolas, et al.
Published: (2025)
by: Bougie, Nicolas, et al.
Published: (2025)
SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation
by: Bougie, Nicolas, et al.
Published: (2025)
by: Bougie, Nicolas, et al.
Published: (2025)
MobileCity: An Efficient Framework for Large-Scale Urban Behavior Simulation
by: Ye, Xiaotong, et al.
Published: (2025)
by: Ye, Xiaotong, et al.
Published: (2025)
Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation
by: Bougie, Nicolas, et al.
Published: (2026)
by: Bougie, Nicolas, et al.
Published: (2026)
AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation
by: Bougie, Nicolas, et al.
Published: (2026)
by: Bougie, Nicolas, et al.
Published: (2026)
Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent
by: Zhang, Yangshijie, et al.
Published: (2025)
by: Zhang, Yangshijie, et al.
Published: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
by: Shamsi, Zafir, et al.
Published: (2026)
by: Shamsi, Zafir, et al.
Published: (2026)
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models
by: Liu, Ziyu, et al.
Published: (2026)
by: Liu, Ziyu, et al.
Published: (2026)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
by: Gan, Esther, et al.
Published: (2024)
by: Gan, Esther, et al.
Published: (2024)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
Self-Improving Customer Review Response Generation Based on LLMs
by: Azov, Guy, et al.
Published: (2024)
by: Azov, Guy, et al.
Published: (2024)
Advancing NLP Security by Leveraging LLMs as Adversarial Engines
by: Srinivasan, Sudarshan, et al.
Published: (2024)
by: Srinivasan, Sudarshan, et al.
Published: (2024)
Adversarial versification in portuguese as a jailbreak operator in LLMs
by: Queiroz, Joao
Published: (2025)
by: Queiroz, Joao
Published: (2025)
Improving Fairness in LLMs Through Testing-Time Adversaries
by: Gregio, Isabela Pereira, et al.
Published: (2025)
by: Gregio, Isabela Pereira, et al.
Published: (2025)
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
by: Ayoobi, Navid, et al.
Published: (2026)
by: Ayoobi, Navid, et al.
Published: (2026)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
by: Chai, Jianzhe, et al.
Published: (2026)
by: Chai, Jianzhe, et al.
Published: (2026)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
by: Wu, Mian, et al.
Published: (2025)
by: Wu, Mian, et al.
Published: (2025)
From Words to Collisions: LLM-Guided Evaluation and Adversarial Generation of Safety-Critical Driving Scenarios
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills
by: Wang, Yunfei, et al.
Published: (2026)
by: Wang, Yunfei, et al.
Published: (2026)
Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers
by: Li, Ruochi, et al.
Published: (2025)
by: Li, Ruochi, et al.
Published: (2025)
Surgical Feature-Space Decomposition of LLMs: Why, When and How?
by: Chavan, Arnav, et al.
Published: (2024)
by: Chavan, Arnav, et al.
Published: (2024)
When LLMs Team Up: The Emergence of Collaborative Affective Computing
by: Lai, Wenna, et al.
Published: (2025)
by: Lai, Wenna, et al.
Published: (2025)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
by: Machcha, Sravanthi, et al.
Published: (2026)
by: Machcha, Sravanthi, et al.
Published: (2026)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
by: Sun, Zhongxiang, et al.
Published: (2026)
by: Sun, Zhongxiang, et al.
Published: (2026)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
by: Kim, Minkyoung, et al.
Published: (2024)
by: Kim, Minkyoung, et al.
Published: (2024)
What Layers When: Learning to Skip Compute in LLMs with Residual Gates
by: Laitenberger, Filipe, et al.
Published: (2025)
by: Laitenberger, Filipe, et al.
Published: (2025)
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
When Facts Change: Probing LLMs on Evolving Knowledge with evolveQA
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
by: Nakshatri, Nishanth Sridhar, et al.
Published: (2025)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
by: Afzali, Amirabbas, et al.
Published: (2026)
by: Afzali, Amirabbas, et al.
Published: (2026)
When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
by: Kumar, Sandeep, et al.
Published: (2026)
by: Kumar, Sandeep, et al.
Published: (2026)
When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews
by: Watawana, Hasindri, et al.
Published: (2026)
by: Watawana, Hasindri, et al.
Published: (2026)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
by: InternAgent Team, et al.
Published: (2025)
by: InternAgent Team, et al.
Published: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
by: Amin, Hasan, et al.
Published: (2026)
by: Amin, Hasan, et al.
Published: (2026)
Adversarial Math Word Problem Generation
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
by: Seleznyov, Mikhail, et al.
Published: (2025)
by: Seleznyov, Mikhail, et al.
Published: (2025)
Similar Items
-
CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation
by: Bougie, Nicolas, et al.
Published: (2025) -
SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation
by: Bougie, Nicolas, et al.
Published: (2025) -
MobileCity: An Efficient Framework for Large-Scale Urban Behavior Simulation
by: Ye, Xiaotong, et al.
Published: (2025) -
Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation
by: Bougie, Nicolas, et al.
Published: (2026) -
AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation
by: Bougie, Nicolas, et al.
Published: (2026)