CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Jingwei, Wang, Chenyu, Zhao, Zibo, Liu, Wen, Ma, Yi, Gao, Shenghua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
von: Qi, Dacheng, et al.
Veröffentlicht: (2026)
von: Qi, Dacheng, et al.
Veröffentlicht: (2026)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
Text-to-CAD Retrieval: a Strong Baseline
von: Pan, Honghu, et al.
Veröffentlicht: (2026)
von: Pan, Honghu, et al.
Veröffentlicht: (2026)
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
von: Zou, Yueying, et al.
Veröffentlicht: (2025)
von: Zou, Yueying, et al.
Veröffentlicht: (2025)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
CAD-NeRF: Learning NeRFs from Uncalibrated Few-view Images by CAD Model Retrieval
von: Wen, Xin, et al.
Veröffentlicht: (2024)
von: Wen, Xin, et al.
Veröffentlicht: (2024)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
von: Xu, Binqian, et al.
Veröffentlicht: (2024)
von: Xu, Binqian, et al.
Veröffentlicht: (2024)
GeoCAD: Local Geometry-Controllable CAD Generation with Large Language Models
von: Zhang, Zhanwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhanwei, et al.
Veröffentlicht: (2025)
Img2CAD: Conditioned 3D CAD Model Generation from Single Image with Structured Visual Geometry
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
Drawing2CAD: Sequence-to-Sequence Learning for CAD Generation from Vector Drawings
von: Qin, Feiwei, et al.
Veröffentlicht: (2025)
von: Qin, Feiwei, et al.
Veröffentlicht: (2025)
Cantor: Inspiring Multimodal Chain-of-Thought of MLLM
von: Gao, Timin, et al.
Veröffentlicht: (2024)
von: Gao, Timin, et al.
Veröffentlicht: (2024)
CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
von: Huang, Binbin, et al.
Veröffentlicht: (2025)
von: Huang, Binbin, et al.
Veröffentlicht: (2025)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
von: Huang, Runhui, et al.
Veröffentlicht: (2025)
CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMs
von: Wang, Siyu, et al.
Veröffentlicht: (2024)
von: Wang, Siyu, et al.
Veröffentlicht: (2024)
ArtiCAD: Articulated CAD Assembly Design via Multi-Agent Code Generation
von: Shui, Yuan, et al.
Veröffentlicht: (2026)
von: Shui, Yuan, et al.
Veröffentlicht: (2026)
FlexCAD: Unified and Versatile Controllable CAD Generation with Fine-tuned Large Language Models
von: Zhang, Zhanwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhanwei, et al.
Veröffentlicht: (2024)
InstructX: Towards Unified Visual Editing with MLLM Guidance
von: Mou, Chong, et al.
Veröffentlicht: (2025)
von: Mou, Chong, et al.
Veröffentlicht: (2025)
ChatCAD+: Towards a Universal and Reliable Interactive CAD using LLMs
von: Zhao, Zihao, et al.
Veröffentlicht: (2023)
von: Zhao, Zihao, et al.
Veröffentlicht: (2023)
BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM
von: Wen, Haiquan, et al.
Veröffentlicht: (2025)
von: Wen, Haiquan, et al.
Veröffentlicht: (2025)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
von: Wang, Guanqun, et al.
Veröffentlicht: (2024)
von: Wang, Guanqun, et al.
Veröffentlicht: (2024)
CAD-Judge: Toward Efficient Morphological Grading and Verification for Text-to-CAD Generation
von: Zhou, Zheyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Zheyuan, et al.
Veröffentlicht: (2025)
GenCAD-Self-Repairing: Feasibility Enhancement for 3D CAD Generation
von: Tsuji, Chikaha, et al.
Veröffentlicht: (2025)
von: Tsuji, Chikaha, et al.
Veröffentlicht: (2025)
UNICBench: UNIfied Counting Benchmark for MLLM
von: Rong, Chenggang, et al.
Veröffentlicht: (2026)
von: Rong, Chenggang, et al.
Veröffentlicht: (2026)
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
Text2CAD: Text to 3D CAD Generation via Technical Drawings
von: Yavartanoo, Mohsen, et al.
Veröffentlicht: (2024)
von: Yavartanoo, Mohsen, et al.
Veröffentlicht: (2024)
CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation
von: Niu, Ke, et al.
Veröffentlicht: (2025)
von: Niu, Ke, et al.
Veröffentlicht: (2025)
DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image
von: Gao, Daoyi, et al.
Veröffentlicht: (2023)
von: Gao, Daoyi, et al.
Veröffentlicht: (2023)
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
von: Fan, Jiaqi, et al.
Veröffentlicht: (2024)
von: Fan, Jiaqi, et al.
Veröffentlicht: (2024)
ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
Img2CAD: Reverse Engineering 3D CAD Models from Images through VLM-Assisted Conditional Factorization
von: You, Yang, et al.
Veröffentlicht: (2024)
von: You, Yang, et al.
Veröffentlicht: (2024)
UI-UG: A Unified MLLM for UI Understanding and Generation
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
CAD-Recode: Reverse Engineering CAD Code from Point Clouds
von: Rukhovich, Danila, et al.
Veröffentlicht: (2024)
von: Rukhovich, Danila, et al.
Veröffentlicht: (2024)
AIC MLLM: Autonomous Interactive Correction MLLM for Robust Robotic Manipulation
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
von: Xiong, Chuyan, et al.
Veröffentlicht: (2024)
MLLM-CL: Continual Learning for Multimodal Large Language Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
CAD-Coder:Text-Guided CAD Files Code Generation
von: He, Changqi, et al.
Veröffentlicht: (2025)
von: He, Changqi, et al.
Veröffentlicht: (2025)
CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
von: Niu, Ke, et al.
Veröffentlicht: (2025)
von: Niu, Ke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
von: Qi, Dacheng, et al.
Veröffentlicht: (2026) -
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
von: Wang, Chenyu, et al.
Veröffentlicht: (2024) -
Text-to-CAD Retrieval: a Strong Baseline
von: Pan, Honghu, et al.
Veröffentlicht: (2026) -
CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
von: Wu, Jianyu, et al.
Veröffentlicht: (2025) -
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
von: Fang, Rongyao, et al.
Veröffentlicht: (2024)