A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
Fuente:
arXiv
Guardado en:
| Autores principales: | Malikussaid, Gohar, Imad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
por: Ashley, Dylan R., et al.
Publicado: (2026)
por: Ashley, Dylan R., et al.
Publicado: (2026)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025)
por: Tu, Songjun, et al.
Publicado: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
Balanced conic rectified flow
por: Kim, Shin Seong, et al.
Publicado: (2025)
por: Kim, Shin Seong, et al.
Publicado: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
por: Zhang, Junwen, et al.
Publicado: (2025)
por: Zhang, Junwen, et al.
Publicado: (2025)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
por: Gkountouras, John, et al.
Publicado: (2025)
por: Gkountouras, John, et al.
Publicado: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026)
por: Chen, Kewei, et al.
Publicado: (2026)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
por: Huo, Dongjie, et al.
Publicado: (2026)
por: Huo, Dongjie, et al.
Publicado: (2026)
Context-Dependent Affordance Computation in Vision-Language Models
por: Farzulla, Murad
Publicado: (2026)
por: Farzulla, Murad
Publicado: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
por: Lasbordes, Maxence, et al.
Publicado: (2026)
por: Lasbordes, Maxence, et al.
Publicado: (2026)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
por: Lysyi, Andrii, et al.
Publicado: (2025)
por: Lysyi, Andrii, et al.
Publicado: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025)
por: Chen, Kewei, et al.
Publicado: (2025)
Does CLIP perceive art the same way we do?
por: Asperti, Andrea, et al.
Publicado: (2025)
por: Asperti, Andrea, et al.
Publicado: (2025)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
por: Kalušev, Vladimir, et al.
Publicado: (2026)
por: Kalušev, Vladimir, et al.
Publicado: (2026)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
por: Ouyang, Rongxin, et al.
Publicado: (2024)
por: Ouyang, Rongxin, et al.
Publicado: (2024)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
por: Breneur, Oleksandr Marchenko, et al.
Publicado: (2026)
por: Breneur, Oleksandr Marchenko, et al.
Publicado: (2026)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
por: Qi, Dekang, et al.
Publicado: (2026)
por: Qi, Dekang, et al.
Publicado: (2026)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
por: Sharma, Aditya, et al.
Publicado: (2025)
por: Sharma, Aditya, et al.
Publicado: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
por: Yang, Yibo
Publicado: (2025)
por: Yang, Yibo
Publicado: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
por: Cui, Hejie, et al.
Publicado: (2024)
por: Cui, Hejie, et al.
Publicado: (2024)
Do Reasoning Models Enhance Embedding Models?
por: Chan, Wun Yu, et al.
Publicado: (2026)
por: Chan, Wun Yu, et al.
Publicado: (2026)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
por: Menon, Anjali R., et al.
Publicado: (2025)
por: Menon, Anjali R., et al.
Publicado: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
por: Koh, Hyunseo, et al.
Publicado: (2026)
por: Koh, Hyunseo, et al.
Publicado: (2026)
HieraEdgeNet: A Multi-Scale Edge-Enhanced Framework for Automated Pollen Recognition
por: Long, Yuchong, et al.
Publicado: (2025)
por: Long, Yuchong, et al.
Publicado: (2025)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
por: Skripkin, Matvey, et al.
Publicado: (2025)
por: Skripkin, Matvey, et al.
Publicado: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
por: Kaiser, Daniel, et al.
Publicado: (2025)
por: Kaiser, Daniel, et al.
Publicado: (2025)
Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
por: Khan, Md Ashik, et al.
Publicado: (2025)
por: Khan, Md Ashik, et al.
Publicado: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
Application of deep learning approaches for medieval historical documents transcription
por: Voloshchuk, Maksym, et al.
Publicado: (2025)
por: Voloshchuk, Maksym, et al.
Publicado: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
por: Zhang, Haichao, et al.
Publicado: (2026)
por: Zhang, Haichao, et al.
Publicado: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
por: Kuz, Mykola, et al.
Publicado: (2025)
por: Kuz, Mykola, et al.
Publicado: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
por: Wu, Robert, et al.
Publicado: (2024)
por: Wu, Robert, et al.
Publicado: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
por: Pather, Kaviraj, et al.
Publicado: (2025)
por: Pather, Kaviraj, et al.
Publicado: (2025)
Automated Cervical Cancer Detection through Visual Inspection with Acetic Acid in Resource-Poor Settings with Lightweight Deep Learning Models Deployed on an Android Device
por: Maben, Leander Melroy, et al.
Publicado: (2025)
por: Maben, Leander Melroy, et al.
Publicado: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
por: Patel, Urjitkumar, et al.
Publicado: (2025)
por: Patel, Urjitkumar, et al.
Publicado: (2025)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
por: Badshah, Sher, et al.
Publicado: (2025)
por: Badshah, Sher, et al.
Publicado: (2025)
Ejemplares similares
-
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024) -
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025) -
RACAS: Controlling Diverse Robots With a Single Agentic System
por: Ashley, Dylan R., et al.
Publicado: (2026) -
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
por: Tu, Songjun, et al.
Publicado: (2025) -
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)