Saved in:
| Main Authors: | Yung, Canaan, Huang, Hanxun, Leckie, Christopher, Erfani, Sarah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.03502 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
by: Yung, Canaan, et al.
Published: (2024)
by: Yung, Canaan, et al.
Published: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025)
by: Huang, Hanxun, et al.
Published: (2025)
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
by: Wang, Qizhou, et al.
Published: (2025)
by: Wang, Qizhou, et al.
Published: (2025)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026)
by: Cui, Kaiyuan, et al.
Published: (2026)
AudioMosaic: Contrastive Masked Audio Representation Learning
by: Huang, Hanxun, et al.
Published: (2026)
by: Huang, Hanxun, et al.
Published: (2026)
Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
by: Hasanebrahimi, Afsaneh, et al.
Published: (2026)
Intention-aware Hierarchical Diffusion Model for Long-term Trajectory Anomaly Detection
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
OIL-AD: An Anomaly Detection Framework for Sequential Decision Sequences
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
TRACER: Persistent Regularization for Robust Multimodal Finetuning
by: Asadollahzadeh, Hesam, et al.
Published: (2026)
by: Asadollahzadeh, Hesam, et al.
Published: (2026)
LDReg: Local Dimensionality Regularized Self-Supervised Learning
by: Huang, Hanxun, et al.
Published: (2024)
by: Huang, Hanxun, et al.
Published: (2024)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
by: Zhang, Yechen, et al.
Published: (2026)
by: Zhang, Yechen, et al.
Published: (2026)
Prompting Implicit Discourse Relation Annotation
by: Yung, Frances, et al.
Published: (2024)
by: Yung, Frances, et al.
Published: (2024)
DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection
by: Yu, Xiao, et al.
Published: (2023)
by: Yu, Xiao, et al.
Published: (2023)
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
by: Li, Wenyun, et al.
Published: (2025)
by: Li, Wenyun, et al.
Published: (2025)
Fighting Fire with Fire: Adversarial Prompting to Generate a Misinformation Detection Dataset
by: Satapara, Shrey, et al.
Published: (2024)
by: Satapara, Shrey, et al.
Published: (2024)
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection
by: Marinelli, Ryan, et al.
Published: (2025)
by: Marinelli, Ryan, et al.
Published: (2025)
SelfPrompt: Autonomously Evaluating LLM Robustness via Domain-Constrained Knowledge Guidelines and Refined Adversarial Prompts
by: Pei, Aihua, et al.
Published: (2024)
by: Pei, Aihua, et al.
Published: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026)
by: Aneja, Krishak, et al.
Published: (2026)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
by: Lee, Yu-Ting, et al.
Published: (2025)
by: Lee, Yu-Ting, et al.
Published: (2025)
Label-Guided Prompt for Multi-label Few-shot Aspect Category Detection
by: Guan, ChaoFeng, et al.
Published: (2024)
by: Guan, ChaoFeng, et al.
Published: (2024)
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
by: Pedashenko, Vladislav, et al.
Published: (2025)
by: Pedashenko, Vladislav, et al.
Published: (2025)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
by: Tuck, Bryan E., et al.
Published: (2025)
by: Tuck, Bryan E., et al.
Published: (2025)
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
by: Ghasemabadi, Amirhosein, et al.
Published: (2025)
by: Ghasemabadi, Amirhosein, et al.
Published: (2025)
Ontology-Guided, Hybrid Prompt Learning for Generalization in Knowledge Graph Question Answering
by: Jiang, Longquan, et al.
Published: (2025)
by: Jiang, Longquan, et al.
Published: (2025)
BeSimulator: A Large Language Model Powered Text-based Behavior Simulator
by: Wang, Jianan, et al.
Published: (2024)
by: Wang, Jianan, et al.
Published: (2024)
Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
by: Tulchinskii, Eduard, et al.
Published: (2023)
by: Tulchinskii, Eduard, et al.
Published: (2023)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
by: Zheng, Yujia, et al.
Published: (2025)
by: Zheng, Yujia, et al.
Published: (2025)
DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
by: Borah, Abhilekh, et al.
Published: (2025)
by: Borah, Abhilekh, et al.
Published: (2025)
Parse Trees Guided LLM Prompt Compression
by: Mao, Wenhao, et al.
Published: (2024)
by: Mao, Wenhao, et al.
Published: (2024)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
by: Gogoulou, Evangelia, et al.
Published: (2025)
by: Gogoulou, Evangelia, et al.
Published: (2025)
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
by: Zhang, Xiaomei, et al.
Published: (2025)
by: Zhang, Xiaomei, et al.
Published: (2025)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
by: Wang, Xilong, et al.
Published: (2026)
by: Wang, Xilong, et al.
Published: (2026)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024)
by: Das, Nilanjana, et al.
Published: (2024)
Local Prompt Optimization
by: Jain, Yash, et al.
Published: (2025)
by: Jain, Yash, et al.
Published: (2025)
Fight Back Against Jailbreaking via Prompt Adversarial Tuning
by: Mo, Yichuan, et al.
Published: (2024)
by: Mo, Yichuan, et al.
Published: (2024)
Similar Items
-
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
by: Yung, Canaan, et al.
Published: (2024) -
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
by: Huang, Hanxun, et al.
Published: (2025) -
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
by: Wang, Qizhou, et al.
Published: (2025) -
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
by: Cui, Kaiyuan, et al.
Published: (2026) -
AudioMosaic: Contrastive Masked Audio Representation Learning
by: Huang, Hanxun, et al.
Published: (2026)