Unlocking UML Class Diagram Understanding in Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Naboichenko, Artem, Peinl, René |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
by: Gautam, Sushant, et al.
Published: (2026)
by: Gautam, Sushant, et al.
Published: (2026)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
by: Lyu, Wenbo, et al.
Published: (2025)
by: Lyu, Wenbo, et al.
Published: (2025)
RailSafeNet: Visual Scene Understanding for Tram Safety
by: Valach, Ondřej, et al.
Published: (2025)
by: Valach, Ondřej, et al.
Published: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)
Adversarial Patch Attacks on Vision-Based Cargo Occupancy Estimation via Differentiable 3D Simulation
by: Hedna, Mohamed Rissal, et al.
Published: (2025)
by: Hedna, Mohamed Rissal, et al.
Published: (2025)
The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
by: Ömercikoğlu, Ahmet Can, et al.
Published: (2025)
by: Ömercikoğlu, Ahmet Can, et al.
Published: (2025)
Dual-sensing driving detection model
by: K, Leon C. C., et al.
Published: (2025)
by: K, Leon C. C., et al.
Published: (2025)
MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
by: Kwak, SeokJoo, et al.
Published: (2025)
by: Kwak, SeokJoo, et al.
Published: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
by: Yang, Baoyao, et al.
Published: (2025)
by: Yang, Baoyao, et al.
Published: (2025)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
by: Jung, Seoik, et al.
Published: (2025)
by: Jung, Seoik, et al.
Published: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
by: Patel, Urjitkumar, et al.
Published: (2025)
by: Patel, Urjitkumar, et al.
Published: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
by: Su, Yuetong, et al.
Published: (2025)
by: Su, Yuetong, et al.
Published: (2025)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Archival Faces: Detection of Faces in Digitized Historical Documents
by: Vaško, Marek, et al.
Published: (2025)
by: Vaško, Marek, et al.
Published: (2025)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
by: Ayanzadeh, Aydin, et al.
Published: (2026)
by: Ayanzadeh, Aydin, et al.
Published: (2026)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
by: Sajjadi, Arash, et al.
Published: (2025)
by: Sajjadi, Arash, et al.
Published: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
by: Ko, Hyun-kyu, et al.
Published: (2026)
by: Ko, Hyun-kyu, et al.
Published: (2026)
WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis
by: Lyu, Xinheng, et al.
Published: (2025)
by: Lyu, Xinheng, et al.
Published: (2025)
MVTamperBench: Evaluating Robustness of Vision-Language Models
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
by: Baydemir, Poyraz
Published: (2025)
by: Baydemir, Poyraz
Published: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
by: Rashid, Muhammad, et al.
Published: (2026)
by: Rashid, Muhammad, et al.
Published: (2026)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
by: Zinnen, Mathias, et al.
Published: (2025)
by: Zinnen, Mathias, et al.
Published: (2025)
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
by: Hu, Yangliu, et al.
Published: (2025)
by: Hu, Yangliu, et al.
Published: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
by: Chen, Jingkun, et al.
Published: (2025)
by: Chen, Jingkun, et al.
Published: (2025)
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Explaining What Machines See: XAI Strategies in Deep Object Detection Models
by: Seyedmomeni, FatemehSadat, et al.
Published: (2025)
by: Seyedmomeni, FatemehSadat, et al.
Published: (2025)
Objaverse++: Curated 3D Object Dataset with Quality Annotations
by: Lin, Chendi, et al.
Published: (2025)
by: Lin, Chendi, et al.
Published: (2025)
Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
by: Meziani, Yani
Published: (2026)
by: Meziani, Yani
Published: (2026)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
by: Svystun, Serhii, et al.
Published: (2025)
by: Svystun, Serhii, et al.
Published: (2025)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
by: Bruno, Pierangela, et al.
Published: (2024)
by: Bruno, Pierangela, et al.
Published: (2024)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
by: Lysyi, Andrii, et al.
Published: (2025)
by: Lysyi, Andrii, et al.
Published: (2025)
Feature Based Methods in Domain Adaptation for Object Detection: A Review Paper
by: Mohamadi, Helia, et al.
Published: (2024)
by: Mohamadi, Helia, et al.
Published: (2024)
Similar Items
-
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
by: Gautam, Sushant, et al.
Published: (2026) -
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025) -
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
by: Lyu, Wenbo, et al.
Published: (2025) -
RailSafeNet: Visual Scene Understanding for Tram Safety
by: Valach, Ondřej, et al.
Published: (2025) -
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
by: Yanambakkam, Hemanth Teja, et al.
Published: (2025)