Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Yiwen, Guo, Zoey, Zhu, Kaixin, Zhang, Ray, Chen, Qizhi, Jiang, Dongzhi, Liu, Junli, Zeng, Bohan, Song, Haoming, Qu, Delin, Bai, Tianyi, Xu, Dan, Zhang, Wentao, Zhao, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Potential of Encoder-free Architectures in 3D LMMs
by: Tang, Yiwen, et al.
Published: (2025)
by: Tang, Yiwen, et al.
Published: (2025)
FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives
by: Chen, Qizhi, et al.
Published: (2024)
by: Chen, Qizhi, et al.
Published: (2024)
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
by: Zhu, Kaixin, et al.
Published: (2026)
by: Zhu, Kaixin, et al.
Published: (2026)
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
by: Gao, Xianqiang, et al.
Published: (2026)
by: Gao, Xianqiang, et al.
Published: (2026)
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)
by: Yao, Yuanqi, et al.
Published: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting Code
by: Lin, Haobo, et al.
Published: (2026)
by: Lin, Haobo, et al.
Published: (2026)
LiveScene: Language Embedding Interactive Radiance Fields for Physical Scene Rendering and Control
by: Qu, Delin, et al.
Published: (2024)
by: Qu, Delin, et al.
Published: (2024)
PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning
by: Lu, Keer, et al.
Published: (2025)
by: Lu, Keer, et al.
Published: (2025)
Point-PEFT: Parameter-Efficient Fine-Tuning for 3D Pre-trained Models
by: Tang, Yiwen, et al.
Published: (2023)
by: Tang, Yiwen, et al.
Published: (2023)
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
by: Liang, Hao, et al.
Published: (2025)
by: Liang, Hao, et al.
Published: (2025)
A Novel Method to Metigate Demographic and Expert Bias in ICD Coding with Causal Inference
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
A Novel ICD Coding Method Based on Associated and Hierarchical Code Description Distillation
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
by: Liu, Junli, et al.
Published: (2025)
by: Liu, Junli, et al.
Published: (2025)
Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
ClearDepth: Enhanced Stereo Perception of Transparent Objects for Robotic Manipulation
by: Bai, Kaixin, et al.
Published: (2024)
by: Bai, Kaixin, et al.
Published: (2024)
WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes
by: Yang, Ling, et al.
Published: (2025)
by: Yang, Ling, et al.
Published: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis
by: Zeng, Bohan, et al.
Published: (2024)
by: Zeng, Bohan, et al.
Published: (2024)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
by: Tang, Yiwen, et al.
Published: (2024)
by: Tang, Yiwen, et al.
Published: (2024)
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
by: Guo, Hailong, et al.
Published: (2025)
by: Guo, Hailong, et al.
Published: (2025)
Are Today's LLMs Ready to Explain Well-Being Concepts?
by: Jiang, Bohan, et al.
Published: (2025)
by: Jiang, Bohan, et al.
Published: (2025)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
by: Yan, Chi, et al.
Published: (2023)
by: Yan, Chi, et al.
Published: (2023)
Implicit Event-RGBD Neural SLAM
by: Qu, Delin, et al.
Published: (2023)
by: Qu, Delin, et al.
Published: (2023)
Language Guided Exploration for RL Agents in Text Environments
by: Golchha, Hitesh, et al.
Published: (2024)
by: Golchha, Hitesh, et al.
Published: (2024)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Are We Really Making Much Progress in Text Classification? A Comparative Review
by: Galke, Lukas, et al.
Published: (2022)
by: Galke, Lukas, et al.
Published: (2022)
"We Demand Justice!": Towards Social Context Grounding of Political Texts
by: Pujari, Rajkumar, et al.
Published: (2023)
by: Pujari, Rajkumar, et al.
Published: (2023)
Contextual Text Denoising with Masked Language Models
by: Sun, Yifu, et al.
Published: (2019)
by: Sun, Yifu, et al.
Published: (2019)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
by: Gao, Xianqiang, et al.
Published: (2024)
by: Gao, Xianqiang, et al.
Published: (2024)
Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
by: Cai, Qifeng, et al.
Published: (2025)
by: Cai, Qifeng, et al.
Published: (2025)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
TextShield-R1: Reinforced Reasoning for Tampered Text Detection
by: Qu, Chenfan, et al.
Published: (2026)
by: Qu, Chenfan, et al.
Published: (2026)
Information Extraction from Clinical Notes: Are We Ready to Switch to Large Language Models?
by: Hu, Yan, et al.
Published: (2024)
by: Hu, Yan, et al.
Published: (2024)
Preference Learning Unlocks LLMs' Psycho-Counseling Skills
by: Zhang, Mian, et al.
Published: (2025)
by: Zhang, Mian, et al.
Published: (2025)
An Interactive Visual Enhancement for Prompted Programmatic Weak Supervision in Text Classification
by: Y. Lin, et al.
Published: (2025)
by: Y. Lin, et al.
Published: (2025)
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
by: Zeng, Bohan, et al.
Published: (2026)
by: Zeng, Bohan, et al.
Published: (2026)
Similar Items
-
Exploring the Potential of Encoder-free Architectures in 3D LMMs
by: Tang, Yiwen, et al.
Published: (2025) -
FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives
by: Chen, Qizhi, et al.
Published: (2024) -
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
by: Zhu, Kaixin, et al.
Published: (2026) -
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
by: Gao, Xianqiang, et al.
Published: (2026) -
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)