CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jianyu, Wang, Yizhou, Yue, Xiangyu, Ma, Xinzhu, Guo, Jingyang, Zhou, Dongzhan, Ouyang, Wanli, Tang, Shixiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Point2Primitive: CAD Reconstruction from Point Cloud by Direct Primitive Prediction
by: Ma, Xinzhu, et al.
Published: (2025)
by: Ma, Xinzhu, et al.
Published: (2025)
3D Object Detection from Images for Autonomous Driving: A Survey
by: Ma, Xinzhu, et al.
Published: (2022)
by: Ma, Xinzhu, et al.
Published: (2022)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025)
by: Tang, Shixiang, et al.
Published: (2025)
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
by: Deng, Han, et al.
Published: (2025)
by: Deng, Han, et al.
Published: (2025)
DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM
by: Wu, Yixuan, et al.
Published: (2024)
by: Wu, Yixuan, et al.
Published: (2024)
A CLIP-Powered Framework for Robust and Generalizable Data Selection
by: Yang, Suorong, et al.
Published: (2024)
by: Yang, Suorong, et al.
Published: (2024)
UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines
by: Tang, Chen, et al.
Published: (2025)
by: Tang, Chen, et al.
Published: (2025)
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory
by: Cao, Jianbao, et al.
Published: (2026)
by: Cao, Jianbao, et al.
Published: (2026)
DiffFluid: Plain Diffusion Models are Effective Predictors of Flow Dynamics
by: Luo, Dongyu, et al.
Published: (2024)
by: Luo, Dongyu, et al.
Published: (2024)
Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
by: Yang, Suorong, et al.
Published: (2025)
by: Yang, Suorong, et al.
Published: (2025)
Hulk: A Universal Knowledge Translator for Human-Centric Tasks
by: Wang, Yizhou, et al.
Published: (2023)
by: Wang, Yizhou, et al.
Published: (2023)
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
by: Fan, Kaixuan, et al.
Published: (2025)
by: Fan, Kaixuan, et al.
Published: (2025)
Attention Reallocation: Towards Zero-cost and Controllable Hallucination Mitigation of MLLMs
by: Tu, Chongjun, et al.
Published: (2025)
by: Tu, Chongjun, et al.
Published: (2025)
Instruct-ReID++: Towards Universal Purpose Instruction-Guided Person Re-identification
by: He, Weizhen, et al.
Published: (2024)
by: He, Weizhen, et al.
Published: (2024)
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
Transition Models: Rethinking the Generative Learning Objective
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
Learning Geometry-Guided Depth via Projective Modeling for Monocular 3D Object Detection
by: Zhang, Yinmin, et al.
Published: (2021)
by: Zhang, Yinmin, et al.
Published: (2021)
LOCR: Location-Guided Transformer for Optical Character Recognition
by: Sun, Yu, et al.
Published: (2024)
by: Sun, Yu, et al.
Published: (2024)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
by: Zhang, Sha, et al.
Published: (2024)
by: Zhang, Sha, et al.
Published: (2024)
Native-Resolution Image Synthesis
by: Wang, Zidong, et al.
Published: (2025)
by: Wang, Zidong, et al.
Published: (2025)
Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time Shifts
by: Hu, Zixuan, et al.
Published: (2025)
by: Hu, Zixuan, et al.
Published: (2025)
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
by: Zhang, Ruiyang, et al.
Published: (2026)
by: Zhang, Ruiyang, et al.
Published: (2026)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023)
by: Zhang, Yaqi, et al.
Published: (2023)
MokA: Multimodal Low-Rank Adaptation for MLLMs
by: Wei, Yake, et al.
Published: (2025)
by: Wei, Yake, et al.
Published: (2025)
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions
by: He, Weizhen, et al.
Published: (2023)
by: He, Weizhen, et al.
Published: (2023)
EMR-Merging: Tuning-Free High-Performance Model Merging
by: Huang, Chenyu, et al.
Published: (2024)
by: Huang, Chenyu, et al.
Published: (2024)
COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation
by: Zhang, Xinlong, et al.
Published: (2026)
by: Zhang, Xinlong, et al.
Published: (2026)
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
CMT: Cross Modulation Transformer with Hybrid Loss for Pansharpening
by: Shu, Wen-Jie, et al.
Published: (2024)
by: Shu, Wen-Jie, et al.
Published: (2024)
Holistic-Motion2D: Scalable Whole-body Human Motion Generation in 2D Space
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
A Comprehensive Survey on 3D Content Generation
by: Liu, Jian, et al.
Published: (2024)
by: Liu, Jian, et al.
Published: (2024)
SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
by: Mai, Weijian, et al.
Published: (2025)
by: Mai, Weijian, et al.
Published: (2025)
GUPNet++: Geometry Uncertainty Propagation Network for Monocular 3D Object Detection
by: Lu, Yan, et al.
Published: (2023)
by: Lu, Yan, et al.
Published: (2023)
CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects
by: Tang, Zhenran, et al.
Published: (2026)
by: Tang, Zhenran, et al.
Published: (2026)
Solving all laminar flows around airfoils all-at-once using a parametric neural network solver
by: Cao, Wenbo, et al.
Published: (2025)
by: Cao, Wenbo, et al.
Published: (2025)
Charting Empirical Laws for LLM Fine-Tuning in Scientific Multi-Discipline Learning
by: Wang, Lintao, et al.
Published: (2026)
by: Wang, Lintao, et al.
Published: (2026)
Img2CAD: Conditioned 3D CAD Model Generation from Single Image with Structured Visual Geometry
by: Chen, Tianrun, et al.
Published: (2024)
by: Chen, Tianrun, et al.
Published: (2024)
Similar Items
-
Point2Primitive: CAD Reconstruction from Point Cloud by Direct Primitive Prediction
by: Ma, Xinzhu, et al.
Published: (2025) -
3D Object Detection from Images for Autonomous Driving: A Survey
by: Ma, Xinzhu, et al.
Published: (2022) -
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024) -
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
by: Tang, Shixiang, et al.
Published: (2025) -
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
by: Deng, Han, et al.
Published: (2025)