Saved in:
| Main Authors: | Kim, Euntae, Han, Soomin, Chang, Buru |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.19274 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
by: Lee, Kyuho, et al.
Published: (2025)
by: Lee, Kyuho, et al.
Published: (2025)
SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script
by: Kim, Eunwon, et al.
Published: (2024)
by: Kim, Eunwon, et al.
Published: (2024)
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
by: Lee, Seohyeong, et al.
Published: (2025)
by: Lee, Seohyeong, et al.
Published: (2025)
CoAuthorAI: A Human in the Loop System For Scientific Book Writing
by: Tian, Yangjie, et al.
Published: (2026)
by: Tian, Yangjie, et al.
Published: (2026)
In-Context Learning with Noisy Labels
by: Kang, Junyong, et al.
Published: (2024)
by: Kang, Junyong, et al.
Published: (2024)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
by: Dong, Yiting, et al.
Published: (2024)
by: Dong, Yiting, et al.
Published: (2024)
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Models
by: Xu, Zhi, et al.
Published: (2026)
by: Xu, Zhi, et al.
Published: (2026)
Hermit Kingdom Through the Lens of Multiple Perspectives: A Case Study of LLM Hallucination on North Korea
by: Cho, Eunjung, et al.
Published: (2025)
by: Cho, Eunjung, et al.
Published: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
by: Xu, Zhao, et al.
Published: (2024)
by: Xu, Zhao, et al.
Published: (2024)
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
by: Mysore, Sheshera, et al.
Published: (2025)
by: Mysore, Sheshera, et al.
Published: (2025)
Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
by: Kim, Jieyong, et al.
Published: (2024)
by: Kim, Jieyong, et al.
Published: (2024)
Chain of Draft: Thinking Faster by Writing Less
by: Xu, Silei, et al.
Published: (2025)
by: Xu, Silei, et al.
Published: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
by: Wei, Zhipeng, et al.
Published: (2024)
by: Wei, Zhipeng, et al.
Published: (2024)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
by: Huang, Caishuang, et al.
Published: (2024)
by: Huang, Caishuang, et al.
Published: (2024)
SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis
by: Wong, Aidan, et al.
Published: (2024)
by: Wong, Aidan, et al.
Published: (2024)
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Cascade Speculative Drafting for Even Faster LLM Inference
by: Chen, Ziyi, et al.
Published: (2023)
by: Chen, Ziyi, et al.
Published: (2023)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
Rotate, Clip, and Partition: Towards W2A4KV4 Quantization by Integrating Rotation and Learnable Non-uniform Quantizer
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
SafeInt: Shielding Large Language Models from Jailbreak Attacks via Safety-Aware Representation Intervention
by: Wu, Jiaqi, et al.
Published: (2025)
by: Wu, Jiaqi, et al.
Published: (2025)
Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
by: Lin, Yuping, et al.
Published: (2024)
by: Lin, Yuping, et al.
Published: (2024)
MEGAnno+: A Human-LLM Collaborative Annotation System
by: Kim, Hannah, et al.
Published: (2024)
by: Kim, Hannah, et al.
Published: (2024)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
by: Li, Xiaoxia, et al.
Published: (2024)
by: Li, Xiaoxia, et al.
Published: (2024)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
by: Shen, Guobin, et al.
Published: (2025)
by: Shen, Guobin, et al.
Published: (2025)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
by: Lee, Minhwa, et al.
Published: (2024)
by: Lee, Minhwa, et al.
Published: (2024)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
by: Chen, Shuo, et al.
Published: (2024)
by: Chen, Shuo, et al.
Published: (2024)
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free
by: Choi, Euntae, et al.
Published: (2025)
by: Choi, Euntae, et al.
Published: (2025)
Enhancing Jailbreak Attacks with Diversity Guidance
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
by: Yu, Miao, et al.
Published: (2024)
by: Yu, Miao, et al.
Published: (2024)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
by: Fang, Zhicheng, et al.
Published: (2026)
by: Fang, Zhicheng, et al.
Published: (2026)
RICoTA: Red-teaming of In-the-wild Conversation with Test Attempts
by: Choi, Eujeong, et al.
Published: (2025)
by: Choi, Eujeong, et al.
Published: (2025)
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
by: Cui, Chenhang, et al.
Published: (2024)
by: Cui, Chenhang, et al.
Published: (2024)
PATENTWRITER: A Benchmarking Study for Patent Drafting with LLMs
by: Shomee, Homaira Huda, et al.
Published: (2025)
by: Shomee, Homaira Huda, et al.
Published: (2025)
LLM-as-a-tutor in EFL Writing Education: Focusing on Evaluation of Student-LLM Interaction
by: Han, Jieun, et al.
Published: (2023)
by: Han, Jieun, et al.
Published: (2023)
Similar Items
-
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
by: Lee, Kyuho, et al.
Published: (2025) -
SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script
by: Kim, Eunwon, et al.
Published: (2024) -
Alignment Data Map for Efficient Preference Data Selection and Diagnosis
by: Lee, Seohyeong, et al.
Published: (2025) -
CoAuthorAI: A Human in the Loop System For Scientific Book Writing
by: Tian, Yangjie, et al.
Published: (2026) -
In-Context Learning with Noisy Labels
by: Kang, Junyong, et al.
Published: (2024)