Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Haoyang, Ju, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches
by: Wu, Sifan, et al.
Published: (2024)
by: Wu, Sifan, et al.
Published: (2024)
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Text-to-CAD Evaluation with CADTests
by: Mallis, Dimitrios, et al.
Published: (2026)
by: Mallis, Dimitrios, et al.
Published: (2026)
CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers
by: Mallis, Dimitrios, et al.
Published: (2024)
by: Mallis, Dimitrios, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective
by: Ma, Xiaorui, et al.
Published: (2025)
by: Ma, Xiaorui, et al.
Published: (2025)
STIV: Scalable Text and Image Conditioned Video Generation
by: Lin, Zongyu, et al.
Published: (2024)
by: Lin, Zongyu, et al.
Published: (2024)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
by: Miao, Yanting, et al.
Published: (2026)
by: Miao, Yanting, et al.
Published: (2026)
Field Matching: an Electrostatic Paradigm to Generate and Transfer Data
by: Kolesov, Alexander, et al.
Published: (2025)
by: Kolesov, Alexander, et al.
Published: (2025)
STEP-Parts: Geometric Partitioning of Boundary Representations for Large-Scale CAD Processing
by: Fan, Shen, et al.
Published: (2026)
by: Fan, Shen, et al.
Published: (2026)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
by: Zhou, Shengli, et al.
Published: (2026)
by: Zhou, Shengli, et al.
Published: (2026)
Text-to-image Diffusion Models in Generative AI: A Survey
by: Zhang, Chenshuang, et al.
Published: (2023)
by: Zhang, Chenshuang, et al.
Published: (2023)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
by: Yu, Weihao, et al.
Published: (2023)
by: Yu, Weihao, et al.
Published: (2023)
FairCoT: Enhancing Fairness in Text-to-Image Generation via Chain of Thought Reasoning with Multimodal Large Language Models
by: Sahili, Zahraa Al, et al.
Published: (2024)
by: Sahili, Zahraa Al, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
BIKED++: A Multimodal Dataset of 1.4 Million Bicycle Image and Parametric CAD Designs
by: Regenwetter, Lyle, et al.
Published: (2024)
by: Regenwetter, Lyle, et al.
Published: (2024)
TextCraftor: Your Text Encoder Can be Image Quality Controller
by: Li, Yanyu, et al.
Published: (2024)
by: Li, Yanyu, et al.
Published: (2024)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
by: Tschannen, Michael, et al.
Published: (2024)
by: Tschannen, Michael, et al.
Published: (2024)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
How to Teach Large Multimodal Models New Skills
by: Zhu, Zhen, et al.
Published: (2025)
by: Zhu, Zhen, et al.
Published: (2025)
Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL
by: Wu, Junyi, et al.
Published: (2026)
by: Wu, Junyi, et al.
Published: (2026)
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
by: Huang, Wenxuan, et al.
Published: (2025)
by: Huang, Wenxuan, et al.
Published: (2025)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
by: Oertell, Owen, et al.
Published: (2024)
by: Oertell, Owen, et al.
Published: (2024)
Knowledge Translation: A New Pathway for Model Compression
by: Sun, Wujie, et al.
Published: (2024)
by: Sun, Wujie, et al.
Published: (2024)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
by: Wang, Ruiyu, et al.
Published: (2025)
by: Wang, Ruiyu, et al.
Published: (2025)
PoGDiff: Product-of-Gaussians Diffusion Models for Imbalanced Text-to-Image Generation
by: Wang, Ziyan, et al.
Published: (2025)
by: Wang, Ziyan, et al.
Published: (2025)
A Scalable and Generalized Deep Learning Framework for Anomaly Detection in Surveillance Videos
by: Jebur, Sabah Abdulazeez, et al.
Published: (2024)
by: Jebur, Sabah Abdulazeez, et al.
Published: (2024)
Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework
by: Kamal, Sadia, et al.
Published: (2025)
by: Kamal, Sadia, et al.
Published: (2025)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
by: Rao, Zhefan, et al.
Published: (2024)
by: Rao, Zhefan, et al.
Published: (2024)
Text-To-Image with Generative Adversarial Networks
by: Momen-Tayefeh, Mehrshad
Published: (2024)
by: Momen-Tayefeh, Mehrshad
Published: (2024)
ReText: Text Boosts Generalization in Image-Based Person Re-identification
by: Mamedov, Timur, et al.
Published: (2026)
by: Mamedov, Timur, et al.
Published: (2026)
Capabilities of Gemini Models in Medicine
by: Saab, Khaled, et al.
Published: (2024)
by: Saab, Khaled, et al.
Published: (2024)
Examining Common Paradigms in Multi-Task Learning
by: Elich, Cathrin, et al.
Published: (2023)
by: Elich, Cathrin, et al.
Published: (2023)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
by: Wang, Hengyi, et al.
Published: (2024)
by: Wang, Hengyi, et al.
Published: (2024)
Exploring Text-to-Motion Generation with Human Preference
by: Sheng, Jenny, et al.
Published: (2024)
by: Sheng, Jenny, et al.
Published: (2024)
EdgeFusion: On-Device Text-to-Image Generation
by: Castells, Thibault, et al.
Published: (2024)
by: Castells, Thibault, et al.
Published: (2024)
Similar Items
-
CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches
by: Wu, Sifan, et al.
Published: (2024) -
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024) -
Text-to-CAD Evaluation with CADTests
by: Mallis, Dimitrios, et al.
Published: (2026) -
CAD-Assistant: Tool-Augmented VLLMs as Generic CAD Task Solvers
by: Mallis, Dimitrios, et al.
Published: (2024) -
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)