Towards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Ni, TsaiChing, Chen, ZhenQi, Yang, YuanFu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
by: Chen, ZhenQi, et al.
Published: (2025)
by: Chen, ZhenQi, et al.
Published: (2025)
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
by: Hsu, YuChe, et al.
Published: (2025)
by: Hsu, YuChe, et al.
Published: (2025)
SceneFoundry: Generating Interactive Infinite 3D Worlds
by: Chen, ChunTeng, et al.
Published: (2026)
by: Chen, ChunTeng, et al.
Published: (2026)
Scaling Open-Vocabulary Action Detection
by: Sia, Zhen Hao, et al.
Published: (2025)
by: Sia, Zhen Hao, et al.
Published: (2025)
Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich Semantics
by: Yang, Shuai, et al.
Published: (2023)
by: Yang, Shuai, et al.
Published: (2023)
Open-Vocabulary SAM3D: Towards Training-free Open-Vocabulary 3D Scene Understanding
by: Tai, Hanchen, et al.
Published: (2024)
by: Tai, Hanchen, et al.
Published: (2024)
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
by: Xiang, Xinhao, et al.
Published: (2025)
by: Xiang, Xinhao, et al.
Published: (2025)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
by: Zhang, Zhenhao, et al.
Published: (2025)
by: Zhang, Zhenhao, et al.
Published: (2025)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
by: Jiang, Siyang, et al.
Published: (2025)
by: Jiang, Siyang, et al.
Published: (2025)
OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
by: Wu, Yanmin, et al.
Published: (2024)
by: Wu, Yanmin, et al.
Published: (2024)
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023)
by: Minderer, Matthias, et al.
Published: (2023)
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
by: Tan, Rongbin, et al.
Published: (2026)
by: Tan, Rongbin, et al.
Published: (2026)
Open Vocabulary Semantic Scene Sketch Understanding
by: Bourouis, Ahmed, et al.
Published: (2023)
by: Bourouis, Ahmed, et al.
Published: (2023)
ISP-AD: A Large-Scale Real-World Dataset for Advancing Industrial Anomaly Detection with Synthetic and Real Defects
by: Krassnig, Paul J., et al.
Published: (2025)
by: Krassnig, Paul J., et al.
Published: (2025)
Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images
by: Yang, Shuai, et al.
Published: (2026)
by: Yang, Shuai, et al.
Published: (2026)
Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation
by: Wang, Pengfei, et al.
Published: (2024)
by: Wang, Pengfei, et al.
Published: (2024)
Boosting Segment Anything Model Towards Open-Vocabulary Learning
by: Han, Xumeng, et al.
Published: (2023)
by: Han, Xumeng, et al.
Published: (2023)
OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
by: Nguyen, Hieu, et al.
Published: (2025)
by: Nguyen, Hieu, et al.
Published: (2025)
Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability
by: Jiang, Shuai, et al.
Published: (2025)
by: Jiang, Shuai, et al.
Published: (2025)
Towards Open Vocabulary Learning: A Survey
by: Wu, Jianzong, et al.
Published: (2023)
by: Wu, Jianzong, et al.
Published: (2023)
CAGS: Open-Vocabulary 3D Scene Understanding with Context-Aware Gaussian Splatting
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
by: Guo, Xianda, et al.
Published: (2025)
by: Guo, Xianda, et al.
Published: (2025)
Training-Free Class Purification for Open-Vocabulary Semantic Segmentation
by: Chen, Qi, et al.
Published: (2025)
by: Chen, Qi, et al.
Published: (2025)
Understanding Multi-Granularity for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2024)
by: Choi, Jiho, et al.
Published: (2024)
ProFuse: Efficient Cross-View Context Fusion for Open-Vocabulary 3D Gaussian Splatting
by: Chiou, Yen-Jen, et al.
Published: (2026)
by: Chiou, Yen-Jen, et al.
Published: (2026)
OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments
by: Bello, Hymalai, et al.
Published: (2026)
by: Bello, Hymalai, et al.
Published: (2026)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
by: Chen, Guo, et al.
Published: (2026)
by: Chen, Guo, et al.
Published: (2026)
Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation
by: Chen, Lin, et al.
Published: (2025)
by: Chen, Lin, et al.
Published: (2025)
OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments
by: Deng, Yinan, et al.
Published: (2024)
by: Deng, Yinan, et al.
Published: (2024)
Learning to Weigh Waste: A Physics-Informed Multimodal Fusion Framework and Large-Scale Dataset for Commercial and Industrial Applications
by: Islam, Md. Adnanul, et al.
Published: (2026)
by: Islam, Md. Adnanul, et al.
Published: (2026)
Open-Vocabulary Octree-Graph for 3D Scene Understanding
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
by: Xie, Jiangnan, et al.
Published: (2025)
by: Xie, Jiangnan, et al.
Published: (2025)
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
by: Hu, Chen, et al.
Published: (2025)
by: Hu, Chen, et al.
Published: (2025)
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection
by: Lei, Ting, et al.
Published: (2024)
by: Lei, Ting, et al.
Published: (2024)
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Large-Scale Universal Defect Generation: Foundation Models and Datasets
by: Fan, Yuanting, et al.
Published: (2026)
by: Fan, Yuanting, et al.
Published: (2026)
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
by: Yuan, Zhihao, et al.
Published: (2023)
by: Yuan, Zhihao, et al.
Published: (2023)
OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking
by: Liang, Haiji, et al.
Published: (2024)
by: Liang, Haiji, et al.
Published: (2024)
Similar Items
-
CritiFusion: Semantic Critique and Spectral Alignment for Faithful Text-to-Image Generation
by: Chen, ZhenQi, et al.
Published: (2025) -
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
by: Hsu, YuChe, et al.
Published: (2025) -
SceneFoundry: Generating Interactive Infinite 3D Worlds
by: Chen, ChunTeng, et al.
Published: (2026) -
Scaling Open-Vocabulary Action Detection
by: Sia, Zhen Hao, et al.
Published: (2025) -
Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich Semantics
by: Yang, Shuai, et al.
Published: (2023)