Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Hao, Zhang, Tianyi, Zhuang, Tianqu, Kong, Jiawei, Gao, Kuofeng, Chen, Bin, Zheng, Leqi, Xia, Shu-Tao, Xu, Ke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
by: Kong, Jiawei, et al.
Published: (2025)
by: Kong, Jiawei, et al.
Published: (2025)
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
by: Kong, Jiawei, et al.
Published: (2025)
by: Kong, Jiawei, et al.
Published: (2025)
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature Filtering
by: Yu, Hongyao, et al.
Published: (2024)
by: Yu, Hongyao, et al.
Published: (2024)
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
MIBench: A Comprehensive Framework for Benchmarking Model Inversion Attack and Defense
by: Qiu, Yixiang, et al.
Published: (2024)
by: Qiu, Yixiang, et al.
Published: (2024)
Denial-of-Service Poisoning Attacks against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing
by: Chen, Bin, et al.
Published: (2024)
by: Chen, Bin, et al.
Published: (2024)
ICAS: Detecting Training Data from Autoregressive Image Generative Models
by: Yu, Hongyao, et al.
Published: (2025)
by: Yu, Hongyao, et al.
Published: (2025)
Imperceptible Jailbreaking against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization
by: Kong, Jiawei, et al.
Published: (2026)
by: Kong, Jiawei, et al.
Published: (2026)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
by: Yang, Xiaochen, et al.
Published: (2026)
by: Yang, Xiaochen, et al.
Published: (2026)
Curriculum Learning-Guided Progressive Distillation in Large Language Models
by: Cao, Jincheng, et al.
Published: (2026)
by: Cao, Jincheng, et al.
Published: (2026)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
by: Sun, Shuoyang, et al.
Published: (2026)
by: Sun, Shuoyang, et al.
Published: (2026)
Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
Going Beyond Feature Similarity: Effective Dataset Distillation based on Class-Aware Conditional Mutual Information
by: Zhong, Xinhao, et al.
Published: (2024)
by: Zhong, Xinhao, et al.
Published: (2024)
A Survey on Knowledge Distillation of Large Language Models
by: Xu, Xiaohan, et al.
Published: (2024)
by: Xu, Xiaohan, et al.
Published: (2024)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
Privacy Leakage on DNNs: A Survey of Model Inversion Attacks and Defenses
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
by: Liu, Haitong, et al.
Published: (2025)
by: Liu, Haitong, et al.
Published: (2025)
Credence Calibration Game? Calibrating Large Language Models through Structured Play
by: Fang, Ke, et al.
Published: (2025)
by: Fang, Ke, et al.
Published: (2025)
Hierarchical Features Matter: A Deep Exploration of Progressive Parameterization Method for Dataset Distillation
by: Zhong, Xinhao, et al.
Published: (2024)
by: Zhong, Xinhao, et al.
Published: (2024)
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Revisiting the Privacy Risks of Split Inference: A GAN-Based Data Reconstruction Attack via Progressive Feature Optimization
by: Qiu, Yixiang, et al.
Published: (2025)
by: Qiu, Yixiang, et al.
Published: (2025)
Enhancing Gradient Inversion Attacks in Federated Learning via Hierarchical Feature Optimization
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
GI-NAS: Boosting Gradient Inversion Attacks Through Adaptive Neural Architecture Search
by: Yu, Wenbo, et al.
Published: (2024)
by: Yu, Wenbo, et al.
Published: (2024)
Similar Items
-
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026) -
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
by: Fang, Hao, et al.
Published: (2025) -
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025) -
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
by: Kong, Jiawei, et al.
Published: (2025) -
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
by: Kong, Jiawei, et al.
Published: (2025)