Pushing Auto-regressive Models for 3D Shape Generation at Capacity and Scalability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Xuelin, Wang, Yu, Luo, Simian, Zhang, Yinda, Tai, Ying, Zhang, Zhenyu, Wang, Chengjie, Xue, Xiangyang, Zhao, Bo, Huang, Tiejun, Wu, Yunsheng, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
von: Zhang, Shijie, et al.
Veröffentlicht: (2024)
von: Zhang, Shijie, et al.
Veröffentlicht: (2024)
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
Exploring Fine-Grained Representation and Recomposition for Cloth-Changing Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2023)
von: Wang, Qizao, et al.
Veröffentlicht: (2023)
Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
Beyond 'Templates': Category-Agnostic Object Pose, Size, and Shape Estimation from a Single View
von: Zhang, Jinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jinyu, et al.
Veröffentlicht: (2025)
MinD-3D: Reconstruct High-quality 3D objects in Human Brain
von: Gao, Jianxiong, et al.
Veröffentlicht: (2023)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2023)
Distribution Aligned Semantics Adaption for Lifelong Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing
von: Cao, Chenjie, et al.
Veröffentlicht: (2024)
von: Cao, Chenjie, et al.
Veröffentlicht: (2024)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026)
CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
Synthesizing Efficient Data with Diffusion Models for Person Re-Identification Pre-Training
von: Niu, Ke, et al.
Veröffentlicht: (2024)
von: Niu, Ke, et al.
Veröffentlicht: (2024)
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
von: Jiang, Boyuan, et al.
Veröffentlicht: (2024)
von: Jiang, Boyuan, et al.
Veröffentlicht: (2024)
MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
von: Cao, Chenjie, et al.
Veröffentlicht: (2024)
von: Cao, Chenjie, et al.
Veröffentlicht: (2024)
Aligned Stable Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
von: Wang, Yikai, et al.
Veröffentlicht: (2026)
Improving Neural Surface Reconstruction with Feature Priors from Multi-View Image
von: Ren, Xinlin, et al.
Veröffentlicht: (2024)
von: Ren, Xinlin, et al.
Veröffentlicht: (2024)
RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base
von: Wang, Kuanning, et al.
Veröffentlicht: (2025)
von: Wang, Kuanning, et al.
Veröffentlicht: (2025)
SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
von: Tan, Weipeng, et al.
Veröffentlicht: (2024)
von: Tan, Weipeng, et al.
Veröffentlicht: (2024)
NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation
von: Xie, Yu, et al.
Veröffentlicht: (2025)
von: Xie, Yu, et al.
Veröffentlicht: (2025)
Pushing The Limit of LLM Capacity for Text Classification
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
von: Zhang, Yazhou, et al.
Veröffentlicht: (2024)
Hyper-Transformer for Amodal Completion
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
AutoOdom: Learning Auto-regressive Proprioceptive Odometry for Legged Locomotion
von: Luo, Changsheng, et al.
Veröffentlicht: (2025)
von: Luo, Changsheng, et al.
Veröffentlicht: (2025)
DiffFAE: Advancing High-fidelity One-shot Facial Appearance Editing with Space-sensitive Customization and Semantic Preservation
von: Wang, Qilin, et al.
Veröffentlicht: (2024)
von: Wang, Qilin, et al.
Veröffentlicht: (2024)
You Only Estimate Once: Unified, One-stage, Real-Time Category-level Articulated Object 6D Pose Estimation for Robotic Grasping
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
Multi-modal Auto-regressive Modeling via Visual Words
von: Peng, Tianshuo, et al.
Veröffentlicht: (2024)
von: Peng, Tianshuo, et al.
Veröffentlicht: (2024)
Chapter Flaws as features
von: Simian, Ricardo
Veröffentlicht: (2023)
von: Simian, Ricardo
Veröffentlicht: (2023)
EMOv2: Pushing 5M Vision Model Frontier
von: Zhang, Jiangning, et al.
Veröffentlicht: (2024)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2024)
NeuroPictor: Refining fMRI-to-Image Reconstruction via Multi-individual Pretraining and Multi-level Modulation
von: Huo, Jingyang, et al.
Veröffentlicht: (2024)
von: Huo, Jingyang, et al.
Veröffentlicht: (2024)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
von: Wang, Kuanning, et al.
Veröffentlicht: (2025)
von: Wang, Kuanning, et al.
Veröffentlicht: (2025)
Pre-training Auto-regressive Robotic Models with 4D Representations
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
von: Niu, Dantong, et al.
Veröffentlicht: (2025)
Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
von: Wang, Yikai, et al.
Veröffentlicht: (2023)
von: Wang, Yikai, et al.
Veröffentlicht: (2023)
Auto-selected Knowledge Adapters for Lifelong Person Re-identification
von: Qian, Xuelin, et al.
Veröffentlicht: (2024)
von: Qian, Xuelin, et al.
Veröffentlicht: (2024)
Enhancing Video Inpainting with Aligned Frame Interval Guidance
von: Xie, Ming, et al.
Veröffentlicht: (2025)
von: Xie, Ming, et al.
Veröffentlicht: (2025)
Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning
von: Yu, Qifan, et al.
Veröffentlicht: (2025)
von: Yu, Qifan, et al.
Veröffentlicht: (2025)
ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model
von: Xu, Chengming, et al.
Veröffentlicht: (2024)
von: Xu, Chengming, et al.
Veröffentlicht: (2024)
Deepfake Generation and Detection: A Benchmark and Survey
von: Pei, Gan, et al.
Veröffentlicht: (2024)
von: Pei, Gan, et al.
Veröffentlicht: (2024)
AutoSiMP: Autonomous Topology Optimization from Natural Language via LLM-Driven Problem Configuration and Adaptive Solver Control
von: Yang, Shaoliang, et al.
Veröffentlicht: (2026)
von: Yang, Shaoliang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
von: Zhang, Shijie, et al.
Veröffentlicht: (2024) -
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
von: Wang, Qizao, et al.
Veröffentlicht: (2024) -
Exploring Fine-Grained Representation and Recomposition for Cloth-Changing Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2023) -
Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2024) -
Beyond 'Templates': Category-Agnostic Object Pose, Size, and Shape Estimation from a Single View
von: Zhang, Jinyu, et al.
Veröffentlicht: (2025)