Evaluating Vision Transformer Models for Visual Quality Control in Industrial Manufacturing
Fuente:
arXiv
Saved in:
| Main Authors: | Alber, Miriam, Hönes, Christoph, Baier, Patrick |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
Improving the Precision of CNNs for Magnetic Resonance Spectral Modeling
by: LaMaster, John, et al.
Published: (2024)
by: LaMaster, John, et al.
Published: (2024)
Power Plant Detection for Energy Estimation using GIS with Remote Sensing, CNN & Vision Transformers
by: Austin-Gabriel, Blessing, et al.
Published: (2024)
by: Austin-Gabriel, Blessing, et al.
Published: (2024)
Explainable embeddings with Distance Explainer
by: Meijer, Christiaan, et al.
Published: (2025)
by: Meijer, Christiaan, et al.
Published: (2025)
Defining and Quantifying Creative Behavior in Popular Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2025)
by: Ramaswamy, Aditi, et al.
Published: (2025)
The Uncanny Valley: A Comprehensive Analysis of Diffusion Models
by: Ghanem, Karam, et al.
Published: (2024)
by: Ghanem, Karam, et al.
Published: (2024)
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
by: Liu, Liangjin, et al.
Published: (2025)
by: Liu, Liangjin, et al.
Published: (2025)
In Context Learning with Vision Transformers: Case Study
by: Zhao, Antony, et al.
Published: (2025)
by: Zhao, Antony, et al.
Published: (2025)
Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations
by: Penuela, Ricardo Gonzalez, et al.
Published: (2025)
by: Penuela, Ricardo Gonzalez, et al.
Published: (2025)
Visual Tuning
by: Yu, Bruce X. B., et al.
Published: (2023)
by: Yu, Bruce X. B., et al.
Published: (2023)
Visual Language Models show widespread visual deficits on neuropsychological tests
by: Tangtartharakul, Gene, et al.
Published: (2025)
by: Tangtartharakul, Gene, et al.
Published: (2025)
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
by: Wu, Diankun, et al.
Published: (2025)
by: Wu, Diankun, et al.
Published: (2025)
Towards virtual painting recolouring using Vision Transformer on X-Ray Fluorescence datacubes
by: Bombini, Alessandro, et al.
Published: (2024)
by: Bombini, Alessandro, et al.
Published: (2024)
NFR: Neural Feature-Guided Non-Rigid Shape Registration
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
Composite Data Augmentations for Synthetic Image Detection Against Real-World Perturbations
by: Amarantidou, Efthymia, et al.
Published: (2025)
by: Amarantidou, Efthymia, et al.
Published: (2025)
Curated endoscopic retrograde cholangiopancreatography images dataset
by: Andrade, Alda João, et al.
Published: (2026)
by: Andrade, Alda João, et al.
Published: (2026)
Guidelines for External Disturbance Factors in the Use of OCR in Real-World Environments
by: Iwata, Kenji, et al.
Published: (2025)
by: Iwata, Kenji, et al.
Published: (2025)
Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection
by: Ferrari, Niccolò, et al.
Published: (2026)
by: Ferrari, Niccolò, et al.
Published: (2026)
A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
by: Nazeri, Ali, et al.
Published: (2025)
by: Nazeri, Ali, et al.
Published: (2025)
SynthEnsemble: A Fusion of CNN, Vision Transformer, and Hybrid Models for Multi-Label Chest X-Ray Classification
by: Ashraf, S. M. Nabil, et al.
Published: (2023)
by: Ashraf, S. M. Nabil, et al.
Published: (2023)
A Spitting Image: Modular Superpixel Tokenization in Vision Transformers
by: Aasan, Marius, et al.
Published: (2024)
by: Aasan, Marius, et al.
Published: (2024)
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
by: Li, Tao, et al.
Published: (2024)
by: Li, Tao, et al.
Published: (2024)
When Fairness Metrics Disagree: Evaluating the Reliability of Demographic Fairness Assessment in Machine Learning
by: Alsayed, Khalid Adnan
Published: (2026)
by: Alsayed, Khalid Adnan
Published: (2026)
Eureka: Evaluating and Understanding Large Foundation Models
by: Balachandran, Vidhisha, et al.
Published: (2024)
by: Balachandran, Vidhisha, et al.
Published: (2024)
Digital Divides in Scene Recognition: Uncovering Socioeconomic Biases in Deep Learning Systems
by: Greene, Michelle R., et al.
Published: (2024)
by: Greene, Michelle R., et al.
Published: (2024)
Unsupervised Domain Adaptation via Style-Aware Self-intermediate Domain
by: Wang, Lianyu, et al.
Published: (2022)
by: Wang, Lianyu, et al.
Published: (2022)
ForAug: Recombining Foregrounds and Backgrounds to Improve Vision Transformer Training with Bias Mitigation
by: Nauen, Tobias Christian, et al.
Published: (2025)
by: Nauen, Tobias Christian, et al.
Published: (2025)
Which Transformer to Favor: A Comparative Analysis of Efficiency in Vision Transformers
by: Nauen, Tobias Christian, et al.
Published: (2023)
by: Nauen, Tobias Christian, et al.
Published: (2023)
Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models
by: Kedia, Akhil, et al.
Published: (2024)
by: Kedia, Akhil, et al.
Published: (2024)
Balanced Anomaly-guided Ego-graph Diffusion Model for Inductive Graph Anomaly Detection
by: Wei, Chunyu, et al.
Published: (2026)
by: Wei, Chunyu, et al.
Published: (2026)
SAG-ViT: A Scale-Aware, High-Fidelity Patching Approach with Graph Attention for Vision Transformers
by: Venkatraman, Shravan, et al.
Published: (2024)
by: Venkatraman, Shravan, et al.
Published: (2024)
Circuit Mechanisms for Spatial Relation Generation in Diffusion Transformers
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
Hybrid Knowledge Transfer through Attention and Logit Distillation for On-Device Vision Systems in Agricultural IoT
by: Mugisha, Stanley, et al.
Published: (2025)
by: Mugisha, Stanley, et al.
Published: (2025)
Deep Learning for automated multi-scale functional field boundaries extraction using multi-date Sentinel-2 and PlanetScope imagery: Case Study of Netherlands and Pakistan
by: Zahid, Saba, et al.
Published: (2024)
by: Zahid, Saba, et al.
Published: (2024)
Simple Self Organizing Map with Vision Transformers
by: Luo, Alan, et al.
Published: (2025)
by: Luo, Alan, et al.
Published: (2025)
CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization
by: Hong, Dasol, et al.
Published: (2025)
by: Hong, Dasol, et al.
Published: (2025)
From Rule-Based Models to Deep Learning Transformers Architectures for Natural Language Processing and Sign Language Translation Systems: Survey, Taxonomy and Performance Evaluation
by: Shahin, Nada, et al.
Published: (2024)
by: Shahin, Nada, et al.
Published: (2024)
Uncertainty Quantification in Continual Open-World Learning
by: Rios, Amanda S., et al.
Published: (2024)
by: Rios, Amanda S., et al.
Published: (2024)
CountPath: Automating Fragment Counting in Digital Pathology
by: Vieira, Ana Beatriz, et al.
Published: (2025)
by: Vieira, Ana Beatriz, et al.
Published: (2025)
SACA: Selective Attention-Based Clustering Algorithm
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2025)
by: Bilehsavar, Meysam Shirdel, et al.
Published: (2025)
Similar Items
-
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
by: Carnemolla, Simone, et al.
Published: (2025) -
Improving the Precision of CNNs for Magnetic Resonance Spectral Modeling
by: LaMaster, John, et al.
Published: (2024) -
Power Plant Detection for Energy Estimation using GIS with Remote Sensing, CNN & Vision Transformers
by: Austin-Gabriel, Blessing, et al.
Published: (2024) -
Explainable embeddings with Distance Explainer
by: Meijer, Christiaan, et al.
Published: (2025) -
Defining and Quantifying Creative Behavior in Popular Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2025)