ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Xiao, Li, Philip, Zeng, Zhichen, Li, Tingwei, Wei, Tianxin, Ning, Xuying, Li, Gaotang, Chen, Yuzhong, Tong, Hanghang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FeDecider: An LLM-Based Framework for Federated Cross-Domain Recommendation
by: He, Xinrui, et al.
Published: (2026)
by: He, Xinrui, et al.
Published: (2026)
i$^2$VAE: Interest Information Augmentation with Variational Regularizers for Cross-Domain Sequential Recommendation
by: Ning, Xuying, et al.
Published: (2024)
by: Ning, Xuying, et al.
Published: (2024)
Taming Knowledge Conflicts in Language Models
by: Li, Gaotang, et al.
Published: (2025)
by: Li, Gaotang, et al.
Published: (2025)
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
by: Zhao, Yanjun, et al.
Published: (2026)
by: Zhao, Yanjun, et al.
Published: (2026)
CATS: Mitigating Correlation Shift for Multivariate Time Series Classification
by: Lin, Xiao, et al.
Published: (2025)
by: Lin, Xiao, et al.
Published: (2025)
THeGCN: Temporal Heterophilic Graph Convolutional Network
by: Yan, Yuchen, et al.
Published: (2024)
by: Yan, Yuchen, et al.
Published: (2024)
Harnessing Consistency for Robust Test-Time LLM Ensemble
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
Mixture of Sequence: Theme-Aware Mixture-of-Experts for Long-Sequence Recommendation
by: Lin, Xiao, et al.
Published: (2026)
by: Lin, Xiao, et al.
Published: (2026)
A Zero-shot Explainable Doctor Ranking Framework with Large Language Models
by: Zeng, Ziyang, et al.
Published: (2025)
by: Zeng, Ziyang, et al.
Published: (2025)
Saffron-1: Safety Inference Scaling
by: Qiu, Ruizhong, et al.
Published: (2025)
by: Qiu, Ruizhong, et al.
Published: (2025)
Hierarchical LoRA MoE for Efficient CTR Model Scaling
by: Zeng, Zhichen, et al.
Published: (2025)
by: Zeng, Zhichen, et al.
Published: (2025)
RALLM-POI: Retrieval-Augmented LLM for Zero-shot Next POI Recommendation with Geographical Reranking
by: Li, Kunrong, et al.
Published: (2025)
by: Li, Kunrong, et al.
Published: (2025)
CoFiRec: Coarse-to-Fine Tokenization for Generative Recommendation
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
Continual Low-Rank Adapters for LLM-based Generative Recommender Systems
by: Yoo, Hyunsik, et al.
Published: (2025)
by: Yoo, Hyunsik, et al.
Published: (2025)
RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
by: Li, Shijun, et al.
Published: (2026)
by: Li, Shijun, et al.
Published: (2026)
An Investigation of Prompt Variations for Zero-shot LLM-based Rankers
by: Sun, Shuoqi, et al.
Published: (2024)
by: Sun, Shuoqi, et al.
Published: (2024)
Zero-shot Cross-domain Knowledge Distillation: A Case study on YouTube Music
by: Ranganathan, Srivaths, et al.
Published: (2026)
by: Ranganathan, Srivaths, et al.
Published: (2026)
Unleashing the Power of Large Language Models in Zero-shot Relation Extraction via Self-Prompting
by: Liu, Siyi, et al.
Published: (2024)
by: Liu, Siyi, et al.
Published: (2024)
Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation
by: Mu, Lingyu, et al.
Published: (2025)
by: Mu, Lingyu, et al.
Published: (2025)
Ensuring User-side Fairness in Dynamic Recommender Systems
by: Yoo, Hyunsik, et al.
Published: (2023)
by: Yoo, Hyunsik, et al.
Published: (2023)
Beyond Reproducibility: Advancing Zero-shot LLM Reranking Efficiency with Setwise Insertion
by: Podolak, Jakub, et al.
Published: (2025)
by: Podolak, Jakub, et al.
Published: (2025)
Internalizing Multi-Agent Reasoning for Accurate and Efficient LLM-based Recommendation
by: Wu, Yang, et al.
Published: (2026)
by: Wu, Yang, et al.
Published: (2026)
Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
by: Li, Haiwen, et al.
Published: (2024)
by: Li, Haiwen, et al.
Published: (2024)
Continual Recommender Systems
by: Yoo, Hyunsik, et al.
Published: (2025)
by: Yoo, Hyunsik, et al.
Published: (2025)
Graph Homophily Booster: Rethinking the Role of Discrete Features on Heterophilic Graphs
by: Qiu, Ruizhong, et al.
Published: (2025)
by: Qiu, Ruizhong, et al.
Published: (2025)
Graph homophily booster: Reimagining the role of discrete features in heterophilic graph learning
by: Qiu, Ruizhong, et al.
Published: (2026)
by: Qiu, Ruizhong, et al.
Published: (2026)
Enhancing Complex Question Answering over Knowledge Graphs through Evidence Pattern Retrieval
by: Ding, Wentao, et al.
Published: (2024)
by: Ding, Wentao, et al.
Published: (2024)
InRanker: Distilled Rankers for Zero-shot Information Retrieval
by: Laitz, Thiago, et al.
Published: (2024)
by: Laitz, Thiago, et al.
Published: (2024)
Flow Matching Meets Biology and Life Science: A Survey
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Subspace Alignment for Vision-Language Model Test-time Adaptation
by: Zeng, Zhichen, et al.
Published: (2026)
by: Zeng, Zhichen, et al.
Published: (2026)
Decoding Matters: Addressing Amplification Bias and Homogeneity Issue for LLM-based Recommendation
by: Bao, Keqin, et al.
Published: (2024)
by: Bao, Keqin, et al.
Published: (2024)
iAgent: LLM Agent as a Shield between User and Recommender Systems
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals
by: Zeng, Linda, et al.
Published: (2025)
by: Zeng, Linda, et al.
Published: (2025)
A Collaborative Ensemble Framework for CTR Prediction
by: Liu, Xiaolong, et al.
Published: (2024)
by: Liu, Xiaolong, et al.
Published: (2024)
ListT5: Listwise Reranking with Fusion-in-Decoder Improves Zero-shot Retrieval
by: Yoon, Soyoung, et al.
Published: (2024)
by: Yoon, Soyoung, et al.
Published: (2024)
Bias Amplification Enhances Minority Group Performance
by: Li, Gaotang, et al.
Published: (2023)
by: Li, Gaotang, et al.
Published: (2023)
Retrieval-Augmented Purifier for Robust LLM-Empowered Recommendation
by: Ning, Liangbo, et al.
Published: (2025)
by: Ning, Liangbo, et al.
Published: (2025)
Fishing for Answers: Exploring One-shot vs. Iterative Retrieval Strategies for Retrieval Augmented Generation
by: Lin, Huifeng, et al.
Published: (2025)
by: Lin, Huifeng, et al.
Published: (2025)
A Pre-trained Sequential Recommendation Framework: Popularity Dynamics for Zero-shot Transfer
by: Wang, Junting, et al.
Published: (2024)
by: Wang, Junting, et al.
Published: (2024)
Soft Filtering: Guiding Zero-shot Composed Image Retrieval with Prescriptive and Proscriptive Constraints
by: Jung, Youjin, et al.
Published: (2025)
by: Jung, Youjin, et al.
Published: (2025)
Similar Items
-
FeDecider: An LLM-Based Framework for Federated Cross-Domain Recommendation
by: He, Xinrui, et al.
Published: (2026) -
i$^2$VAE: Interest Information Augmentation with Variational Regularizers for Cross-Domain Sequential Recommendation
by: Ning, Xuying, et al.
Published: (2024) -
Taming Knowledge Conflicts in Language Models
by: Li, Gaotang, et al.
Published: (2025) -
PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
by: Zhao, Yanjun, et al.
Published: (2026) -
CATS: Mitigating Correlation Shift for Multivariate Time Series Classification
by: Lin, Xiao, et al.
Published: (2025)