JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Hyunju, Lee, Woohyun, Kim, Jaewon, Park, Hogun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly Detection
by: Kim, Soopil, et al.
Published: (2023)
by: Kim, Soopil, et al.
Published: (2023)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026)
by: Park, Jinho, et al.
Published: (2026)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
Parameter Efficient Multi-Class Intelligent Scheduling for Multimodal Online Distributed Industrial Anomaly Detection
by: Wang, Heqiang, et al.
Published: (2026)
by: Wang, Heqiang, et al.
Published: (2026)
Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement
by: Kee, Hogun, et al.
Published: (2025)
by: Kee, Hogun, et al.
Published: (2025)
Multimodal Real-Time Anomaly Detection and Industrial Applications
by: Verma, Aman, et al.
Published: (2025)
by: Verma, Aman, et al.
Published: (2025)
VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning?
by: Kim, Minkyu, et al.
Published: (2026)
by: Kim, Minkyu, et al.
Published: (2026)
Overcoming Data Inequality across Domains with Semi-Supervised Domain Generalization
by: Park, Jinha, et al.
Published: (2024)
by: Park, Jinha, et al.
Published: (2024)
Target-Oriented Single Domain Generalization
by: Heidari, Marzi, et al.
Published: (2025)
by: Heidari, Marzi, et al.
Published: (2025)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
by: Agarwal, Sakshi, et al.
Published: (2026)
by: Agarwal, Sakshi, et al.
Published: (2026)
Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models
by: Ahn, Donghoon, et al.
Published: (2025)
by: Ahn, Donghoon, et al.
Published: (2025)
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
by: Pande, Nilay, et al.
Published: (2025)
by: Pande, Nilay, et al.
Published: (2025)
UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models
by: Kang, Hyunju, et al.
Published: (2026)
by: Kang, Hyunju, et al.
Published: (2026)
Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
by: Kim, Shiwon, et al.
Published: (2026)
by: Kim, Shiwon, et al.
Published: (2026)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
by: Kim, Minung, et al.
Published: (2025)
by: Kim, Minung, et al.
Published: (2025)
MMToM-QA: Multimodal Theory of Mind Question Answering
by: Jin, Chuanyang, et al.
Published: (2024)
by: Jin, Chuanyang, et al.
Published: (2024)
MAMS: Model-Agnostic Module Selection Framework for Video Captioning
by: Lee, Sangho, et al.
Published: (2025)
by: Lee, Sangho, et al.
Published: (2025)
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
Recall-Oriented Continual Learning with Generative Adversarial Meta-Model
by: Kang, Haneol, et al.
Published: (2024)
by: Kang, Haneol, et al.
Published: (2024)
Self-Rectifying Diffusion Sampling with Perturbed-Attention Guidance
by: Ahn, Donghoon, et al.
Published: (2024)
by: Ahn, Donghoon, et al.
Published: (2024)
Anomaly Multi-classification in Industrial Scenarios: Transferring Few-shot Learning to a New Task
by: Liu, Jie, et al.
Published: (2024)
by: Liu, Jie, et al.
Published: (2024)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
A Noise is Worth Diffusion Guidance
by: Ahn, Donghoon, et al.
Published: (2024)
by: Ahn, Donghoon, et al.
Published: (2024)
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
by: Zhang, Gengwei, et al.
Published: (2026)
by: Zhang, Gengwei, et al.
Published: (2026)
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
by: Kim, Sungnyun, et al.
Published: (2023)
by: Kim, Sungnyun, et al.
Published: (2023)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
by: Park, Yeji, et al.
Published: (2024)
by: Park, Yeji, et al.
Published: (2024)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
DoMIX: An Efficient Framework for Exploiting Domain Knowledge in Fine-Tuning
by: Kim, Dohoon, et al.
Published: (2025)
by: Kim, Dohoon, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Reasoning-Augmented Representations for Multimodal Retrieval
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
by: Yerramilli, Sahiti, et al.
Published: (2025)
by: Yerramilli, Sahiti, et al.
Published: (2025)
Domain Generalizable Continual Learning
by: Yan, Hongwei, et al.
Published: (2025)
by: Yan, Hongwei, et al.
Published: (2025)
Motion-Oriented Compositional Neural Radiance Fields for Monocular Dynamic Human Modeling
by: Kim, Jaehyeok, et al.
Published: (2024)
by: Kim, Jaehyeok, et al.
Published: (2024)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
by: Kim, Yunho, et al.
Published: (2024)
by: Kim, Yunho, et al.
Published: (2024)
Constant Acceleration Flow
by: Park, Dogyun, et al.
Published: (2024)
by: Park, Dogyun, et al.
Published: (2024)
V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
The Role of Teacher Calibration in Knowledge Distillation
by: Kim, Suyoung, et al.
Published: (2025)
by: Kim, Suyoung, et al.
Published: (2025)
Similar Items
-
Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly Detection
by: Kim, Soopil, et al.
Published: (2023) -
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026) -
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025) -
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024) -
Parameter Efficient Multi-Class Intelligent Scheduling for Multimodal Online Distributed Industrial Anomaly Detection
by: Wang, Heqiang, et al.
Published: (2026)