OCR is All you need: Importing Multi-Modality into Image-based Defect Detection System
Fuente:
arXiv
Guardado en:
| Autores principales: | Hsu, Chih-Chung, Lee, Chia-Ming, Sun, Chun-Hung, Wu, Kuang-Ming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
MISS: Memory-efficient Instance Segmentation Framework By Visual Inductive Priors Flow Propagation
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
Augment Before Copy-Paste: Data and Memory Efficiency-Oriented Instance Segmentation Framework for Sport-scenes
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
DenseSR: Image Shadow Removal as Dense Prediction
por: Lin, Yu-Fan, et al.
Publicado: (2025)
por: Lin, Yu-Fan, et al.
Publicado: (2025)
DRCT: Saving Image Super-resolution away from Information Bottleneck
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
WWE-UIE: A Wavelet & White Balance Efficient Network for Underwater Image Enhancement
por: Cheng, Ching-Heng, et al.
Publicado: (2025)
por: Cheng, Ching-Heng, et al.
Publicado: (2025)
CSAKD: Knowledge Distillation with Cross Self-Attention for Hyperspectral and Multispectral Image Fusion
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
Divide and Conquer: Grounding a Bleeding Areas in Gastrointestinal Image with Two-Stage Model
por: Lin, Yu-Fan, et al.
Publicado: (2024)
por: Lin, Yu-Fan, et al.
Publicado: (2024)
UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection
por: Lai, Ching-Yi, et al.
Publicado: (2025)
por: Lai, Ching-Yi, et al.
Publicado: (2025)
Robust Hyperspectral Image Panshapring via Sparse Spatial-Spectral Representation
por: Lee, Chia-Ming, et al.
Publicado: (2025)
por: Lee, Chia-Ming, et al.
Publicado: (2025)
VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results
por: Qiu, Bo-Cheng, et al.
Publicado: (2026)
por: Qiu, Bo-Cheng, et al.
Publicado: (2026)
Multi Source COVID-19 Detection via Kernel-Density-based Slice Sampling
por: Lee, Chia-Ming, et al.
Publicado: (2025)
por: Lee, Chia-Ming, et al.
Publicado: (2025)
Self-supervised Fusarium Head Blight Detection with Hyperspectral Image and Feature Mining
por: Lin, Yu-Fan, et al.
Publicado: (2024)
por: Lin, Yu-Fan, et al.
Publicado: (2024)
Towards Robust DeepFake Detection under Unstable Face Sequences: Adaptive Sparse Graph Embedding with Order-Free Representation and Explicit Laplacian Spectral Prior
por: Hsu, Chih-Chung, et al.
Publicado: (2025)
por: Hsu, Chih-Chung, et al.
Publicado: (2025)
Real-Time Compressed Sensing for Joint Hyperspectral Image Transmission and Restoration for CubeSat
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
HSSDCT: Factorized Spatial-Spectral Correlation for Hyperspectral Image Fusion
por: Lee, Chia-Ming, et al.
Publicado: (2026)
por: Lee, Chia-Ming, et al.
Publicado: (2026)
ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers
por: Hsu, Chih-Chung, et al.
Publicado: (2026)
por: Hsu, Chih-Chung, et al.
Publicado: (2026)
GRACE: Graph-Regularized Attentive Convolutional Entanglement with Laplacian Smoothing for Robust DeepFake Video Detection
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
HyFusion: Enhanced Reception Field Transformer for Hyperspectral Image Fusion
por: Lee, Chia-Ming, et al.
Publicado: (2025)
por: Lee, Chia-Ming, et al.
Publicado: (2025)
PhaSR: Generalized Image Shadow Removal with Physically Aligned Priors
por: Lee, Chia-Ming, et al.
Publicado: (2026)
por: Lee, Chia-Ming, et al.
Publicado: (2026)
Simple 2D Convolutional Neural Network-based Approach for COVID-19 Detection
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
GTATrack: Winner Solution to SoccerTrack 2025 with Deep-EIoU and Global Tracklet Association
por: Jian, Rong-Lin, et al.
Publicado: (2026)
por: Jian, Rong-Lin, et al.
Publicado: (2026)
ReflexSplit: Single Image Reflection Separation via Layer Fusion-Separation
por: Lee, Chia-Ming, et al.
Publicado: (2026)
por: Lee, Chia-Ming, et al.
Publicado: (2026)
CLEAR: Cross-Transformers with Pre-trained Language Model is All you need for Person Attribute Recognition and Retrieval
por: Bui, Doanh C., et al.
Publicado: (2024)
por: Bui, Doanh C., et al.
Publicado: (2024)
ChangeDINO: DINOv3-Driven Building Change Detection in Optical Remote Sensing Imagery
por: Cheng, Ching-Heng, et al.
Publicado: (2025)
por: Cheng, Ching-Heng, et al.
Publicado: (2025)
A Closer Look at Spatial-Slice Features Learning for COVID-19 Detection
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
GazeNLQ @ Ego4D Natural Language Queries Challenge 2025
por: Lin, Wei-Cheng, et al.
Publicado: (2025)
por: Lin, Wei-Cheng, et al.
Publicado: (2025)
VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection
por: Qiu, Bo-Cheng, et al.
Publicado: (2026)
por: Qiu, Bo-Cheng, et al.
Publicado: (2026)
Multi-Modal Face Anti-Spoofing via Cross-Modal Feature Transitions
por: Chong, Jun-Xiong, et al.
Publicado: (2025)
por: Chong, Jun-Xiong, et al.
Publicado: (2025)
Beyond Detection: Multi-Scale Hidden-Code for Natural Image Deepfake Recovery and Factual Retrieval
por: Chen, Yuan-Chih, et al.
Publicado: (2026)
por: Chen, Yuan-Chih, et al.
Publicado: (2026)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
por: Chen, Song, et al.
Publicado: (2025)
por: Chen, Song, et al.
Publicado: (2025)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
por: Li, Yuxuan, et al.
Publicado: (2026)
por: Li, Yuxuan, et al.
Publicado: (2026)
Image compositing is all you need for data augmentation
por: Shermaine, Ang Jia Ning, et al.
Publicado: (2025)
por: Shermaine, Ang Jia Ning, et al.
Publicado: (2025)
Self-Adaptive Gamma Context-Aware SSM-based Model for Metal Defect Detection
por: Sun, Sijin, et al.
Publicado: (2025)
por: Sun, Sijin, et al.
Publicado: (2025)
CANDLE: Illumination-Invariant Semantic Priors for Color Ambient Lighting Normalization
por: Jian, Rong-Lin, et al.
Publicado: (2026)
por: Jian, Rong-Lin, et al.
Publicado: (2026)
Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
por: Cioni, Dario, et al.
Publicado: (2024)
por: Cioni, Dario, et al.
Publicado: (2024)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
por: Huang, Kuan-Chih, et al.
Publicado: (2023)
por: Huang, Kuan-Chih, et al.
Publicado: (2023)
Robust Multi-Modal Face Anti-Spoofing with Domain Adaptation: Tackling Missing Modalities, Noisy Pseudo-Labels, and Model Degradation
por: Hsu, Ming-Tsung, et al.
Publicado: (2025)
por: Hsu, Ming-Tsung, et al.
Publicado: (2025)
PromptHSI: Universal Hyperspectral Image Restoration with Vision-Language Modulated Frequency Adaptation
por: Lee, Chia-Ming, et al.
Publicado: (2024)
por: Lee, Chia-Ming, et al.
Publicado: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
por: Tien, Dong Nguyen, et al.
Publicado: (2025)
por: Tien, Dong Nguyen, et al.
Publicado: (2025)
Ejemplares similares
-
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
por: Hsu, Chih-Chung, et al.
Publicado: (2024) -
MISS: Memory-efficient Instance Segmentation Framework By Visual Inductive Priors Flow Propagation
por: Hsu, Chih-Chung, et al.
Publicado: (2024) -
Augment Before Copy-Paste: Data and Memory Efficiency-Oriented Instance Segmentation Framework for Sport-scenes
por: Hsu, Chih-Chung, et al.
Publicado: (2024) -
DenseSR: Image Shadow Removal as Dense Prediction
por: Lin, Yu-Fan, et al.
Publicado: (2025) -
DRCT: Saving Image Super-resolution away from Information Bottleneck
por: Hsu, Chih-Chung, et al.
Publicado: (2024)