Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Lizhi, Mu, Honglin, Zhai, Zenan, Wang, Minghan, Wang, Yuxia, Wang, Renxi, Gao, Junjie, Zhang, Yixuan, Che, Wanxiang, Baldwin, Timothy, Han, Xudong, Li, Haonan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
von: Wang, Renxi, et al.
Veröffentlicht: (2025)
von: Wang, Renxi, et al.
Veröffentlicht: (2025)
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
Loki: An Open-Source Tool for Fact Verification
von: Li, Haonan, et al.
Veröffentlicht: (2024)
von: Li, Haonan, et al.
Veröffentlicht: (2024)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
ToolGen: Unified Tool Retrieval and Calling via Generation
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
von: Wang, Renxi, et al.
Veröffentlicht: (2024)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
von: Zhai, Zenan, et al.
Veröffentlicht: (2025)
von: Zhai, Zenan, et al.
Veröffentlicht: (2025)
Achilles Heel of Distributed Multi-Agent Systems
von: Zhang, Yiting, et al.
Veröffentlicht: (2025)
von: Zhang, Yiting, et al.
Veröffentlicht: (2025)
AI Security Beyond Core Domains: Resume Screening as a Case Study of Adversarial Vulnerabilities in Specialized LLM Applications
von: Mu, Honglin, et al.
Veröffentlicht: (2025)
von: Mu, Honglin, et al.
Veröffentlicht: (2025)
Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
von: Wang, Songping, et al.
Veröffentlicht: (2026)
von: Wang, Songping, et al.
Veröffentlicht: (2026)
Extramedullary Disease—Achilles Heel in Myeloma?
von: Shaji Kumar, et al.
Veröffentlicht: (2025)
von: Shaji Kumar, et al.
Veröffentlicht: (2025)
Centromere Protein F in Tumor Biology: Cancer's Achilles Heel
von: Zitong Wan, et al.
Veröffentlicht: (2025)
von: Zitong Wan, et al.
Veröffentlicht: (2025)
Weak Pareto Boundary: The Achilles' Heel of Evolutionary Multi-Objective Optimization
von: Zheng, Ruihao, et al.
Veröffentlicht: (2025)
von: Zheng, Ruihao, et al.
Veröffentlicht: (2025)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
von: Rao, Abhinav, et al.
Veröffentlicht: (2024)
von: Rao, Abhinav, et al.
Veröffentlicht: (2024)
User Profiles: The Achilles' Heel of Web Browsers
von: Somé, Dolière Francis, et al.
Veröffentlicht: (2025)
von: Somé, Dolière Francis, et al.
Veröffentlicht: (2025)
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
von: Geng, Yilin, et al.
Veröffentlicht: (2025)
von: Geng, Yilin, et al.
Veröffentlicht: (2025)
The Achilles' Heel of Angular Margins: A Chebyshev Polynomial Fix for Speaker Verification
von: Wang, Yang, et al.
Veröffentlicht: (2026)
von: Wang, Yang, et al.
Veröffentlicht: (2026)
Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)
von: Chen, Tianyi, et al.
Veröffentlicht: (2025)
The Sword, Shield, and Achilles' Heel: Characterizing the Linguistic Inductive Bias of Large Language Models for Spatial Reasoning in Navigation Planning
von: Zhang, Xudong, et al.
Veröffentlicht: (2026)
von: Zhang, Xudong, et al.
Veröffentlicht: (2026)
Achilles' Heels: Vulnerable Record Identification in Synthetic Data Publishing
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
von: Meeus, Matthieu, et al.
Veröffentlicht: (2023)
Communication in Public Relations: The Achilles Heel of Quality Public Relations
von: Betteke van Ruler
Veröffentlicht: (2007)
von: Betteke van Ruler
Veröffentlicht: (2007)
Rethinking STS and NLI in Large Language Models
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
SimuScene: Training and Benchmarking Code Generation to Simulate Physical Scenarios
von: Wang, Yanan, et al.
Veröffentlicht: (2026)
von: Wang, Yanan, et al.
Veröffentlicht: (2026)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2024)
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
von: Puerto, Haritz, et al.
Veröffentlicht: (2026)
von: Puerto, Haritz, et al.
Veröffentlicht: (2026)
AI's Achilles Heel: The Critical Role of Data Validation in Robust and Trustworthy Systems
von: Revista, Zen, et al.
Veröffentlicht: (2025)
von: Revista, Zen, et al.
Veröffentlicht: (2025)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
LM-Combiner: A Contextual Rewriting Model for Chinese Grammatical Error Correction
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
Beyond Static Evaluation: A Dynamic Approach to Assessing AI Assistants' API Invocation Capabilities
von: Mu, Honglin, et al.
Veröffentlicht: (2024)
von: Mu, Honglin, et al.
Veröffentlicht: (2024)
The Achilles Heel of AI: Fundamentals of Risk-Aware Training Data for High-Consequence Models
von: Cook, Dave, et al.
Veröffentlicht: (2025)
von: Cook, Dave, et al.
Veröffentlicht: (2025)
AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
von: Wang, Renxi, et al.
Veröffentlicht: (2025)
von: Wang, Renxi, et al.
Veröffentlicht: (2025)
A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
von: Wan, Kaiyang, et al.
Veröffentlicht: (2025)
von: Wan, Kaiyang, et al.
Veröffentlicht: (2025)
Benchmarking Gender and Political Bias in Large Language Models
von: Yang, Jinrui, et al.
Veröffentlicht: (2025)
von: Yang, Jinrui, et al.
Veröffentlicht: (2025)
RE$^2$: Improving Chinese Grammatical Error Correction via Retrieving Appropriate Examples with Explanation
von: Wang, Baoxin, et al.
Veröffentlicht: (2025)
von: Wang, Baoxin, et al.
Veröffentlicht: (2025)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
von: Qin, Zixuan, et al.
Veröffentlicht: (2025)
von: Qin, Zixuan, et al.
Veröffentlicht: (2025)
Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time Scaling
von: Ji, Shiyu, et al.
Veröffentlicht: (2025)
von: Ji, Shiyu, et al.
Veröffentlicht: (2025)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
von: Wang, Renxi, et al.
Veröffentlicht: (2025) -
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
von: Wang, Renxi, et al.
Veröffentlicht: (2024) -
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
von: Wang, Yuxia, et al.
Veröffentlicht: (2024) -
Loki: An Open-Source Tool for Fact Verification
von: Li, Haonan, et al.
Veröffentlicht: (2024) -
Demystifying Instruction Mixing for Fine-tuning Large Language Models
von: Wang, Renxi, et al.
Veröffentlicht: (2023)