Large-Scale Universal Defect Generation: Foundation Models and Datasets
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Yuanting, Liu, Jun, Gao, Bin-Bin, Chen, Xiaochen, Lin, Yuhuan, Dai, Zhewei, Zhan, Jiawei, Wang, Chengjie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Multi-Class Anomaly Detection via Diffusion Refinement with Dual Conditioning
by: Zhan, Jiawei, et al.
Published: (2024)
by: Zhan, Jiawei, et al.
Published: (2024)
One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2026)
by: Gao, Bin-Bin, et al.
Published: (2026)
Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
by: Fan, Yuanting, et al.
Published: (2025)
by: Fan, Yuanting, et al.
Published: (2025)
ConsistentRFT: Reducing Visual Hallucinations in Flow-based Reinforcement Fine-Tuning
by: Tan, Xiaofeng, et al.
Published: (2026)
by: Tan, Xiaofeng, et al.
Published: (2026)
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy
by: Tan, Xiaofeng, et al.
Published: (2026)
by: Tan, Xiaofeng, et al.
Published: (2026)
DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
by: Zhan, Jiawei, et al.
Published: (2025)
by: Zhan, Jiawei, et al.
Published: (2025)
MatchDet: A Collaborative Framework for Image Matching and Object Detection
by: Lai, Jinxiang, et al.
Published: (2023)
by: Lai, Jinxiang, et al.
Published: (2023)
Clustered-patch Element Connection for Few-shot Learning
by: Lai, Jinxiang, et al.
Published: (2023)
by: Lai, Jinxiang, et al.
Published: (2023)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026)
by: Fu, Weifu, et al.
Published: (2026)
Decoupling Classifier for Boosting Few-shot Object Detection and Instance Segmentation
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
Few-Shot Anomaly-Driven Generation for Anomaly Classification and Segmentation
by: Gui, Guan, et al.
Published: (2025)
by: Gui, Guan, et al.
Published: (2025)
Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich Semantics
by: Yang, Shuai, et al.
Published: (2023)
by: Yang, Shuai, et al.
Published: (2023)
Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation
by: Luo, Jingnan, et al.
Published: (2026)
by: Luo, Jingnan, et al.
Published: (2026)
Real-IAD: A Real-World Multi-View Dataset for Benchmarking Versatile Industrial Anomaly Detection
by: Wang, Chengjie, et al.
Published: (2024)
by: Wang, Chengjie, et al.
Published: (2024)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Advancing Video Self-Supervised Learning via Image Foundation Models
by: Wu, Jingwei, et al.
Published: (2025)
by: Wu, Jingwei, et al.
Published: (2025)
BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
by: Lai, Jinxiang, et al.
Published: (2025)
by: Lai, Jinxiang, et al.
Published: (2025)
Iterative Volume Fusion for Asymmetric Stereo Matching
by: Gao, Yuanting, et al.
Published: (2025)
by: Gao, Yuanting, et al.
Published: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
by: Chen, Zhe, et al.
Published: (2023)
by: Chen, Zhe, et al.
Published: (2023)
MetaUAS: Universal Anomaly Segmentation with One-Prompt Meta-Learning
by: Gao, Bin-Bin
Published: (2025)
by: Gao, Bin-Bin
Published: (2025)
Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
by: Ni, TsaiChing, et al.
Published: (2025)
by: Ni, TsaiChing, et al.
Published: (2025)
Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
AdaDiffSR: Adaptive Region-aware Dynamic Acceleration Diffusion Model for Real-World Image Super-Resolution
by: Fan, Yuanting, et al.
Published: (2024)
by: Fan, Yuanting, et al.
Published: (2024)
Bootstrapping SparseFormers from Vision Foundation Models
by: Gao, Ziteng, et al.
Published: (2023)
by: Gao, Ziteng, et al.
Published: (2023)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
by: Jiang, Xi, et al.
Published: (2024)
by: Jiang, Xi, et al.
Published: (2024)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
Applications of Large Scale Foundation Models for Autonomous Driving
by: Huang, Yu, et al.
Published: (2023)
by: Huang, Yu, et al.
Published: (2023)
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks
by: Fang, Hao, et al.
Published: (2024)
by: Fang, Hao, et al.
Published: (2024)
LORS: Low-rank Residual Structure for Parameter-Efficient Network Stacking
by: Li, Jialin, et al.
Published: (2024)
by: Li, Jialin, et al.
Published: (2024)
SoftPatch+: Fully Unsupervised Anomaly Classification and Segmentation
by: Wang, Chengjie, et al.
Published: (2024)
by: Wang, Chengjie, et al.
Published: (2024)
AIR-VIEW: The Aviation Image Repository for Visibility Estimation of Weather, A Dataset and Benchmark
by: Mourning, Chad, et al.
Published: (2025)
by: Mourning, Chad, et al.
Published: (2025)
tSF: Transformer-based Semantic Filter for Few-Shot Learning
by: Lai, Jinxiang, et al.
Published: (2022)
by: Lai, Jinxiang, et al.
Published: (2022)
UniEM-3M: A Universal Electron Micrograph Dataset for Microstructural Segmentation and Generation
by: wang, Nan, et al.
Published: (2025)
by: wang, Nan, et al.
Published: (2025)
Transferable 3D Adversarial Shape Completion using Diffusion Models
by: Dai, Xuelong, et al.
Published: (2024)
by: Dai, Xuelong, et al.
Published: (2024)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
by: Li, Hui, et al.
Published: (2024)
by: Li, Hui, et al.
Published: (2024)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
by: Yang, Xiaochen, et al.
Published: (2026)
by: Yang, Xiaochen, et al.
Published: (2026)
Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Decision Boundary-aware Knowledge Consolidation Generates Better Instance-Incremental Learner
by: Nie, Qiang, et al.
Published: (2024)
by: Nie, Qiang, et al.
Published: (2024)
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
by: Yang, Yuhuan, et al.
Published: (2025)
by: Yang, Yuhuan, et al.
Published: (2025)
Similar Items
-
Enhancing Multi-Class Anomaly Detection via Diffusion Refinement with Dual Conditioning
by: Zhan, Jiawei, et al.
Published: (2024) -
One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2026) -
Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
by: Fan, Yuanting, et al.
Published: (2025) -
ConsistentRFT: Reducing Visual Hallucinations in Flow-based Reinforcement Fine-Tuning
by: Tan, Xiaofeng, et al.
Published: (2026) -
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy
by: Tan, Xiaofeng, et al.
Published: (2026)