IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zewen, Yu, Zitong, Ye, Qilang, Xie, Weicheng, Zhuo, Wei, Shen, Linlin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Guided Multimodal Unified Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2026)
by: Li, Zewen, et al.
Published: (2026)
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
IM-IAD: Industrial Image Anomaly Detection Benchmark in Manufacturing
by: Xie, Guoyang, et al.
Published: (2023)
by: Xie, Guoyang, et al.
Published: (2023)
ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection
by: Chen, Qiuhui, et al.
Published: (2026)
by: Chen, Qiuhui, et al.
Published: (2026)
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing
by: Ma, Yingjie, et al.
Published: (2024)
by: Ma, Yingjie, et al.
Published: (2024)
HSS-IAD: A Heterogeneous Same-Sort Industrial Anomaly Detection Dataset
by: Wang, Qishan, et al.
Published: (2025)
by: Wang, Qishan, et al.
Published: (2025)
AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
by: Miao, Junwen, et al.
Published: (2025)
by: Miao, Junwen, et al.
Published: (2025)
LR-IAD:Mask-Free Industrial Anomaly Detection with Logical Reasoning
by: Zeng, Peijian, et al.
Published: (2025)
by: Zeng, Peijian, et al.
Published: (2025)
SFDA-rPPG: Source-Free Domain Adaptive Remote Physiological Measurement with Spatio-Temporal Consistency
by: Xie, Yiping, et al.
Published: (2024)
by: Xie, Yiping, et al.
Published: (2024)
IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly Detection
by: Li, Yanhui, et al.
Published: (2025)
by: Li, Yanhui, et al.
Published: (2025)
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
by: Zhang, Jiayu, et al.
Published: (2025)
by: Zhang, Jiayu, et al.
Published: (2025)
Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
by: Zhang, Jiayu, et al.
Published: (2026)
by: Zhang, Jiayu, et al.
Published: (2026)
Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era
by: Zhu, Wenbing, et al.
Published: (2025)
by: Zhu, Wenbing, et al.
Published: (2025)
AutoIAD: Manager-Driven Multi-Agent Collaboration for Automated Industrial Anomaly Detection
by: Ji, Dongwei, et al.
Published: (2025)
by: Ji, Dongwei, et al.
Published: (2025)
Denoising and Alignment: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing
by: Ma, Yingjie, et al.
Published: (2025)
by: Ma, Yingjie, et al.
Published: (2025)
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing
by: Xie, Yiping, et al.
Published: (2025)
by: Xie, Yiping, et al.
Published: (2025)
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection
by: Wang, Chengjie, et al.
Published: (2024)
by: Wang, Chengjie, et al.
Published: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
by: Xing, Bohao, et al.
Published: (2024)
by: Xing, Bohao, et al.
Published: (2024)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
by: Deng, Huilin, et al.
Published: (2024)
by: Deng, Huilin, et al.
Published: (2024)
Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection?
by: Chen, Zhiling, et al.
Published: (2025)
by: Chen, Zhiling, et al.
Published: (2025)
Real-IAD MVN: A Multi-View Normal Vector Dataset and Benchmark for High-Fidelity Industrial Anomaly Detection
by: Zhu, Wenbing, et al.
Published: (2026)
by: Zhu, Wenbing, et al.
Published: (2026)
Enhancing Adversarial Transferability by Balancing Exploration and Exploitation with Gradient-Guided Sampling
by: Niu, Zenghao, et al.
Published: (2025)
by: Niu, Zenghao, et al.
Published: (2025)
Mitigating Group-Level Fairness Disparities in Federated Visual Language Models
by: Chen, Chaomeng, et al.
Published: (2025)
by: Chen, Chaomeng, et al.
Published: (2025)
Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection
by: Zhang, Yaning, et al.
Published: (2024)
by: Zhang, Yaning, et al.
Published: (2024)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
by: Jiang, Xi, et al.
Published: (2024)
by: Jiang, Xi, et al.
Published: (2024)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
by: Zhang, Yaning, et al.
Published: (2026)
by: Zhang, Yaning, et al.
Published: (2026)
Question-Answer Cross Language Image Matching for Weakly Supervised Semantic Segmentation
by: Deng, Songhe, et al.
Published: (2024)
by: Deng, Songhe, et al.
Published: (2024)
Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly Detection
by: Zhu, Wenbing, et al.
Published: (2025)
by: Zhu, Wenbing, et al.
Published: (2025)
Myriad: Large Multimodal Model by Applying Vision Experts for Industrial Anomaly Detection
by: Li, Yuanze, et al.
Published: (2023)
by: Li, Yuanze, et al.
Published: (2023)
TRRG: Towards Truthful Radiology Report Generation With Cross-modal Disease Clue Enhanced Large Language Model
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
FAIR: Frequency-aware Image Restoration for Industrial Visual Anomaly Detection
by: Liu, Tongkun, et al.
Published: (2023)
by: Liu, Tongkun, et al.
Published: (2023)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
by: Hao, Jing, et al.
Published: (2025)
by: Hao, Jing, et al.
Published: (2025)
Multimodal Fake News Detection: MFND Dataset and Shallow-Deep Multitask Learning
by: Zhu, Ye, et al.
Published: (2025)
by: Zhu, Ye, et al.
Published: (2025)
EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models
by: Peng, Xiaomeng, et al.
Published: (2026)
by: Peng, Xiaomeng, et al.
Published: (2026)
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
by: Ma, Yingjie, et al.
Published: (2025)
by: Ma, Yingjie, et al.
Published: (2025)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
by: Zeng, Xingchen, et al.
Published: (2024)
by: Zeng, Xingchen, et al.
Published: (2024)
Similar Items
-
Text-Guided Multimodal Unified Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2026) -
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
by: Ye, Qilang, et al.
Published: (2024) -
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
by: Ye, Qilang, et al.
Published: (2024) -
IM-IAD: Industrial Image Anomaly Detection Benchmark in Manufacturing
by: Xie, Guoyang, et al.
Published: (2023) -
ZSG-IAD: A Multimodal Framework for Zero-Shot Grounded Industrial Anomaly Detection
by: Chen, Qiuhui, et al.
Published: (2026)