ViRED: Prediction of Visual Relations in Engineering Drawings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Chao, Lin, Ke, Luo, Yiyang, Hou, Jiahui, Li, Xiang-Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026)
von: Li, Ming, et al.
Veröffentlicht: (2026)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
von: Han, Haonan, et al.
Veröffentlicht: (2026)
von: Han, Haonan, et al.
Veröffentlicht: (2026)
ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
von: Luo, Rundong, et al.
Veröffentlicht: (2025)
von: Luo, Rundong, et al.
Veröffentlicht: (2025)
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
von: Yang, Ziteng, et al.
Veröffentlicht: (2025)
von: Yang, Ziteng, et al.
Veröffentlicht: (2025)
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
von: Wu, Linquan, et al.
Veröffentlicht: (2026)
von: Wu, Linquan, et al.
Veröffentlicht: (2026)
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2024)
Natural Language Supervision for Low-light Image Enhancement
von: Tang, Jiahui, et al.
Veröffentlicht: (2025)
von: Tang, Jiahui, et al.
Veröffentlicht: (2025)
Navigating Efficiency in MobileViT through Gaussian Process on Global Architecture Factors
von: Meng, Ke, et al.
Veröffentlicht: (2024)
von: Meng, Ke, et al.
Veröffentlicht: (2024)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
von: Li, Kailing, et al.
Veröffentlicht: (2025)
von: Li, Kailing, et al.
Veröffentlicht: (2025)
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
von: Li, Quanjiang, et al.
Veröffentlicht: (2026)
von: Li, Quanjiang, et al.
Veröffentlicht: (2026)
Predictive Reasoning with Augmented Anomaly Contrastive Learning for Compositional Visual Relations
von: Li, Chengtai, et al.
Veröffentlicht: (2026)
von: Li, Chengtai, et al.
Veröffentlicht: (2026)
3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
von: Xiao, Hongcan, et al.
Veröffentlicht: (2026)
von: Xiao, Hongcan, et al.
Veröffentlicht: (2026)
SFMViT: SlowFast Meet ViT in Chaotic World
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
von: Tong, Haoyu, et al.
Veröffentlicht: (2026)
von: Tong, Haoyu, et al.
Veröffentlicht: (2026)
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
von: Zhang, Juntian, et al.
Veröffentlicht: (2025)
von: Zhang, Juntian, et al.
Veröffentlicht: (2025)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
von: Liao, Bencheng, et al.
Veröffentlicht: (2024)
von: Liao, Bencheng, et al.
Veröffentlicht: (2024)
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
von: Yan, Siming, et al.
Veröffentlicht: (2024)
von: Yan, Siming, et al.
Veröffentlicht: (2024)
Flexible ViG: Learning the Self-Saliency for Flexible Object Recognition
von: Zuo, Lin, et al.
Veröffentlicht: (2024)
von: Zuo, Lin, et al.
Veröffentlicht: (2024)
Automated Parsing of Engineering Drawings for Structured Information Extraction Using a Fine-tuned Document Understanding Transformer
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
Text-Enhanced Panoptic Symbol Spotting in CAD Drawings
von: Liu, Xianlin, et al.
Veröffentlicht: (2025)
von: Liu, Xianlin, et al.
Veröffentlicht: (2025)
From Drawings to Decisions: A Hybrid Vision-Language Framework for Parsing 2D Engineering Drawings into Structured Manufacturing Knowledge
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
von: Khan, Muhammad Tayyab, et al.
Veröffentlicht: (2025)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
von: Wu, Fengyi, et al.
Veröffentlicht: (2025)
von: Wu, Fengyi, et al.
Veröffentlicht: (2025)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
von: Zhang, Ben, et al.
Veröffentlicht: (2025)
von: Zhang, Ben, et al.
Veröffentlicht: (2025)
LaViC: Adapting Large Vision-Language Models to Visually-Aware Conversational Recommendation
von: Jeon, Hyunsik, et al.
Veröffentlicht: (2025)
von: Jeon, Hyunsik, et al.
Veröffentlicht: (2025)
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
von: Qin, Luozheng, et al.
Veröffentlicht: (2026)
von: Qin, Luozheng, et al.
Veröffentlicht: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
Frequency-Dynamic Attention Modulation for Dense Prediction
von: Chen, Linwei, et al.
Veröffentlicht: (2025)
von: Chen, Linwei, et al.
Veröffentlicht: (2025)
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
ViTree: Single-path Neural Tree for Step-wise Interpretable Fine-grained Visual Categorization
von: Lao, Danning, et al.
Veröffentlicht: (2024)
von: Lao, Danning, et al.
Veröffentlicht: (2024)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
von: Tuong, Nguyen Anh, et al.
Veröffentlicht: (2026)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
von: Li, Xueyang, et al.
Veröffentlicht: (2025)
von: Li, Xueyang, et al.
Veröffentlicht: (2025)
Frequency Dynamic Convolution for Dense Image Prediction
von: Chen, Linwei, et al.
Veröffentlicht: (2025)
von: Chen, Linwei, et al.
Veröffentlicht: (2025)
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
von: Wang, Xiyao, et al.
Veröffentlicht: (2026)
von: Wang, Xiyao, et al.
Veröffentlicht: (2026)
Two Causally Related Needles in a Video Haystack
von: Li, Miaoyu, et al.
Veröffentlicht: (2025)
von: Li, Miaoyu, et al.
Veröffentlicht: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
ViTaS: Visual Tactile Soft Fusion Contrastive Learning for Visuomotor Learning
von: Tian, Yufeng, et al.
Veröffentlicht: (2026)
von: Tian, Yufeng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026) -
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
von: Han, Haonan, et al.
Veröffentlicht: (2026) -
ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
von: Luo, Rundong, et al.
Veröffentlicht: (2025) -
ViP$^2$-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
von: Yang, Ziteng, et al.
Veröffentlicht: (2025) -
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
von: Wu, Meiqi, et al.
Veröffentlicht: (2025)