From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hsu, Joy, Jin, Emily, Wu, Jiajun, Mitra, Niloy J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Predicate Hierarchies Improve Few-Shot State Classification
von: Jin, Emily, et al.
Veröffentlicht: (2025)
von: Jin, Emily, et al.
Veröffentlicht: (2025)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025)
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025)
LIM: Large Interpolator Model for Dynamic Reconstruction
von: Sabathier, Remy, et al.
Veröffentlicht: (2025)
von: Sabathier, Remy, et al.
Veröffentlicht: (2025)
Neural Geometry Processing via Spherical Neural Surfaces
von: Williamson, Romy, et al.
Veröffentlicht: (2024)
von: Williamson, Romy, et al.
Veröffentlicht: (2024)
WonderZoom: Multi-Scale 3D World Generation
von: Cao, Jin, et al.
Veröffentlicht: (2025)
von: Cao, Jin, et al.
Veröffentlicht: (2025)
JOG3R: Towards 3D-Consistent Video Generators
von: Huang, Chun-Hao Paul, et al.
Veröffentlicht: (2025)
von: Huang, Chun-Hao Paul, et al.
Veröffentlicht: (2025)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
von: Feng, Chun, et al.
Veröffentlicht: (2024)
von: Feng, Chun, et al.
Veröffentlicht: (2024)
SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
von: Shi, Yukai, et al.
Veröffentlicht: (2025)
von: Shi, Yukai, et al.
Veröffentlicht: (2025)
FocusDD: Real-World Scene Infusion for Robust Dataset Distillation
von: Hu, Youbing, et al.
Veröffentlicht: (2025)
von: Hu, Youbing, et al.
Veröffentlicht: (2025)
General Scene Adaptation for Vision-and-Language Navigation
von: Hong, Haodong, et al.
Veröffentlicht: (2025)
von: Hong, Haodong, et al.
Veröffentlicht: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
von: Cao, Yue, et al.
Veröffentlicht: (2024)
von: Cao, Yue, et al.
Veröffentlicht: (2024)
DynamicPAE: Generating Scene-Aware Physical Adversarial Examples in Real-Time
von: Hu, Jin, et al.
Veröffentlicht: (2024)
von: Hu, Jin, et al.
Veröffentlicht: (2024)
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
von: Lin, Juekai, et al.
Veröffentlicht: (2026)
von: Lin, Juekai, et al.
Veröffentlicht: (2026)
Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory
von: Wu, Ruiqi, et al.
Veröffentlicht: (2026)
von: Wu, Ruiqi, et al.
Veröffentlicht: (2026)
WorldScore: A Unified Evaluation Benchmark for World Generation
von: Duan, Haoyi, et al.
Veröffentlicht: (2025)
von: Duan, Haoyi, et al.
Veröffentlicht: (2025)
Exploring Pose-Based Anomaly Detection for Retail Security: A Real-World Shoplifting Dataset and Benchmark
von: Rashvand, Narges, et al.
Veröffentlicht: (2025)
von: Rashvand, Narges, et al.
Veröffentlicht: (2025)
Bridging Synthetic and Real-World Domains: A Human-in-the-Loop Weakly-Supervised Framework for Industrial Toxic Emission Segmentation
von: Tao, Yida, et al.
Veröffentlicht: (2025)
von: Tao, Yida, et al.
Veröffentlicht: (2025)
RealWonder: Real-Time Physical Action-Conditioned Video Generation
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
von: Shen, Yuan, et al.
Veröffentlicht: (2024)
von: Shen, Yuan, et al.
Veröffentlicht: (2024)
Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction
von: Bejerano, Emily, et al.
Veröffentlicht: (2026)
von: Bejerano, Emily, et al.
Veröffentlicht: (2026)
Planning with the Views via Scene Self-Exploration
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
von: Wang, Kangrui, et al.
Veröffentlicht: (2026)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
Composable Part-Based Manipulation
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
von: Liu, Weiyu, et al.
Veröffentlicht: (2024)
VRAG: Learning World Models for Interactive Video Generation
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
von: Chen, Taiye, et al.
Veröffentlicht: (2025)
BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
von: Zhang, Yumeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yumeng, et al.
Veröffentlicht: (2024)
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
von: Shen, Yedong, et al.
Veröffentlicht: (2026)
von: Shen, Yedong, et al.
Veröffentlicht: (2026)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
SceneFoundry: Generating Interactive Infinite 3D Worlds
von: Chen, ChunTeng, et al.
Veröffentlicht: (2026)
von: Chen, ChunTeng, et al.
Veröffentlicht: (2026)
WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions
von: Li, Zizhang, et al.
Veröffentlicht: (2025)
von: Li, Zizhang, et al.
Veröffentlicht: (2025)
SPLite Hand: Sparsity-Aware Lightweight 3D Hand Pose Estimation
von: Hao, Yeh Keng, et al.
Veröffentlicht: (2025)
von: Hao, Yeh Keng, et al.
Veröffentlicht: (2025)
Visually Descriptive Language Model for Vector Graphics Reasoning
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2024)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2024)
Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
Pose as Clinical Prior: Learning Dual Representations for Scoliosis Screening
von: Zhou, Zirui, et al.
Veröffentlicht: (2025)
von: Zhou, Zirui, et al.
Veröffentlicht: (2025)
DynaPose4D: High-Quality 4D Dynamic Content Generation via Pose Alignment Loss
von: Yang, Jing, et al.
Veröffentlicht: (2025)
von: Yang, Jing, et al.
Veröffentlicht: (2025)
From Easy to Hard: Learning Curricular Shape-aware Features for Robust Panoptic Scene Graph Generation
von: Shi, Hanrong, et al.
Veröffentlicht: (2024)
von: Shi, Hanrong, et al.
Veröffentlicht: (2024)
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
von: Meng, Yanxu, et al.
Veröffentlicht: (2025)
von: Meng, Yanxu, et al.
Veröffentlicht: (2025)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
von: Yang, Yu, et al.
Veröffentlicht: (2025)
von: Yang, Yu, et al.
Veröffentlicht: (2025)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
von: Kabra, Rishabh, et al.
Veröffentlicht: (2026)
von: Kabra, Rishabh, et al.
Veröffentlicht: (2026)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Predicate Hierarchies Improve Few-Shot State Classification
von: Jin, Emily, et al.
Veröffentlicht: (2025) -
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024) -
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
von: Fang, Shaoheng, et al.
Veröffentlicht: (2025) -
LIM: Large Interpolator Model for Dynamic Reconstruction
von: Sabathier, Remy, et al.
Veröffentlicht: (2025) -
Neural Geometry Processing via Spherical Neural Surfaces
von: Williamson, Romy, et al.
Veröffentlicht: (2024)