Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Jinqi, Yang, Jinyu, Neiman, Tal, Fan, Lei, Yin, Bing, Tran, Son, Shah, Mubarak, Vidal, René |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CompLLM: Compression for Long Context Q&A
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024)
by: Swetha, Sirnam, et al.
Published: (2024)
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025)
by: Luo, Jinqi, et al.
Published: (2025)
Contextual Knowledge Pursuit for Faithful Visual Synthesis
by: Luo, Jinqi, et al.
Published: (2023)
by: Luo, Jinqi, et al.
Published: (2023)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
PaCE: Parsimonious Concept Engineering for Large Language Models
by: Luo, Jinqi, et al.
Published: (2024)
by: Luo, Jinqi, et al.
Published: (2024)
M-LLM Based Video Frame Selection for Efficient Video Understanding
by: Hu, Kai, et al.
Published: (2025)
by: Hu, Kai, et al.
Published: (2025)
VidLA: Video-Language Alignment at Scale
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
by: Rizve, Mamshad Nayeem, et al.
Published: (2024)
Evaluating and Aligning CodeLLMs on Human Preference
by: Yang, Jian, et al.
Published: (2024)
by: Yang, Jian, et al.
Published: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
by: Yang, Ruichao, et al.
Published: (2026)
by: Yang, Ruichao, et al.
Published: (2026)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
by: Beetham, James, et al.
Published: (2024)
by: Beetham, James, et al.
Published: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Unsupervised Concept Vector Extraction for Bias Control in LLMs
by: Cyberey, Hannah, et al.
Published: (2025)
by: Cyberey, Hannah, et al.
Published: (2025)
Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?
by: Zhang, Crystina, et al.
Published: (2024)
by: Zhang, Crystina, et al.
Published: (2024)
Speech LLMs are Contextual Reasoning Transcribers
by: Deng, Keqi, et al.
Published: (2026)
by: Deng, Keqi, et al.
Published: (2026)
Controlling Language Confusion in Multilingual LLMs
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
EssayCBM: Rubric-Aligned Concept Bottleneck Models for Transparent Essay Grading
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
by: Chaudhary, Kumar Satvik, et al.
Published: (2025)
Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging
by: Farn, Hua, et al.
Published: (2024)
by: Farn, Hua, et al.
Published: (2024)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2025)
by: Liang, Buyun, et al.
Published: (2025)
Length Controlled Generation for Black-box LLMs
by: Gu, Yuxuan, et al.
Published: (2024)
by: Gu, Yuxuan, et al.
Published: (2024)
Multi-Party Conversational Agents: A Survey
by: Sapkota, Sagar, et al.
Published: (2025)
by: Sapkota, Sagar, et al.
Published: (2025)
Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
by: Kriz, Anita, et al.
Published: (2025)
by: Kriz, Anita, et al.
Published: (2025)
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
by: Leong, Chak Tou, et al.
Published: (2025)
by: Leong, Chak Tou, et al.
Published: (2025)
Bringing Multimodality to Amazon Visual Search System
by: Zhu, Xinliang, et al.
Published: (2024)
by: Zhu, Xinliang, et al.
Published: (2024)
Bridging Dictionary: AI-Generated Dictionary of Partisan Language Use
by: Jiang, Hang, et al.
Published: (2024)
by: Jiang, Hang, et al.
Published: (2024)
Generating Concept Lexicalizations via Dictionary-Based Cross-Lingual Sense Projection
by: Basil, David, et al.
Published: (2026)
by: Basil, David, et al.
Published: (2026)
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
by: Cui, Shiyao, et al.
Published: (2025)
by: Cui, Shiyao, et al.
Published: (2025)
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
by: Jiang, Yilei, et al.
Published: (2024)
by: Jiang, Yilei, et al.
Published: (2024)
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
Multi-Physics: A Comprehensive Benchmark for Multimodal LLMs Reasoning on Chinese Multi-Subject Physics Problems
by: Luo, Zhongze, et al.
Published: (2025)
by: Luo, Zhongze, et al.
Published: (2025)
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs
by: Yin, Lake, et al.
Published: (2025)
by: Yin, Lake, et al.
Published: (2025)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
by: Wang, Zezhong, et al.
Published: (2023)
by: Wang, Zezhong, et al.
Published: (2023)
ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning
by: Li, Shuyue Stella, et al.
Published: (2025)
by: Li, Shuyue Stella, et al.
Published: (2025)
Similar Items
-
CompLLM: Compression for Long Context Q&A
by: Berton, Gabriele, et al.
Published: (2025) -
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs
by: Swetha, Sirnam, et al.
Published: (2024) -
Concept Lancet: Image Editing with Compositional Representation Transplant
by: Luo, Jinqi, et al.
Published: (2025) -
Contextual Knowledge Pursuit for Faithful Visual Synthesis
by: Luo, Jinqi, et al.
Published: (2023) -
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
by: Liang, Buyun, et al.
Published: (2025)