Synthesizing Moving People with 3D Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Boyi, Chen, Junming, Rajasegaran, Jathushan, Gandelsman, Yossi, Efros, Alexei A., Malik, Jitendra |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Empirical Study of Autoregressive Pre-training from Videos
di: Rajasegaran, Jathushan, et al.
Pubblicazione: (2025)
di: Rajasegaran, Jathushan, et al.
Pubblicazione: (2025)
Interpreting CLIP's Image Representation via Text-Based Decomposition
di: Gandelsman, Yossi, et al.
Pubblicazione: (2023)
di: Gandelsman, Yossi, et al.
Pubblicazione: (2023)
Scaling Properties of Diffusion Models for Perceptual Tasks
di: Ravishankar, Rahul, et al.
Pubblicazione: (2024)
di: Ravishankar, Rahul, et al.
Pubblicazione: (2024)
Vision Transformers Don't Need Trained Registers
di: Jiang, Nick, et al.
Pubblicazione: (2025)
di: Jiang, Nick, et al.
Pubblicazione: (2025)
Gaussian Masked Autoencoders
di: Rajasegaran, Jathushan, et al.
Pubblicazione: (2025)
di: Rajasegaran, Jathushan, et al.
Pubblicazione: (2025)
Tracking by Predicting 3-D Gaussians Over Time
di: Baranwal, Tanish, et al.
Pubblicazione: (2025)
di: Baranwal, Tanish, et al.
Pubblicazione: (2025)
Interpreting the Second-Order Effects of Neurons in CLIP
di: Gandelsman, Yossi, et al.
Pubblicazione: (2024)
di: Gandelsman, Yossi, et al.
Pubblicazione: (2024)
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
di: Bu, Edmund, et al.
Pubblicazione: (2025)
di: Bu, Edmund, et al.
Pubblicazione: (2025)
Poly-Autoregressive Prediction for Modeling Interactions
di: Thakkar, Neerja, et al.
Pubblicazione: (2025)
di: Thakkar, Neerja, et al.
Pubblicazione: (2025)
Quantifying and Enabling the Interpretability of CLIP-like Models
di: Madasu, Avinash, et al.
Pubblicazione: (2024)
di: Madasu, Avinash, et al.
Pubblicazione: (2024)
Synergy and Synchrony in Couple Dances
di: Maluleke, Vongani, et al.
Pubblicazione: (2024)
di: Maluleke, Vongani, et al.
Pubblicazione: (2024)
Jailbreaking Vision-Language Models Through the Visual Modality
di: Azulay, Aharon, et al.
Pubblicazione: (2026)
di: Azulay, Aharon, et al.
Pubblicazione: (2026)
LLMs can see and hear without any training
di: Ashutosh, Kumar, et al.
Pubblicazione: (2025)
di: Ashutosh, Kumar, et al.
Pubblicazione: (2025)
Diffusion Models as Data Mining Tools
di: Siglidis, Ioannis, et al.
Pubblicazione: (2024)
di: Siglidis, Ioannis, et al.
Pubblicazione: (2024)
Test-Time Training on Video Streams
di: Wang, Renhao, et al.
Pubblicazione: (2023)
di: Wang, Renhao, et al.
Pubblicazione: (2023)
FewShotNeRF: Meta-Learning-based Novel View Synthesis for Rapid Scene-Specific Adaptation
di: Sivakumar, Piraveen, et al.
Pubblicazione: (2024)
di: Sivakumar, Piraveen, et al.
Pubblicazione: (2024)
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
di: Ekin, Yigit, et al.
Pubblicazione: (2026)
di: Ekin, Yigit, et al.
Pubblicazione: (2026)
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
di: Koepke, A. Sophia, et al.
Pubblicazione: (2026)
di: Koepke, A. Sophia, et al.
Pubblicazione: (2026)
Humanoid Locomotion as Next Token Prediction
di: Radosavovic, Ilija, et al.
Pubblicazione: (2024)
di: Radosavovic, Ilija, et al.
Pubblicazione: (2024)
Learning Video Representations without Natural Videos
di: Yu, Xueyang, et al.
Pubblicazione: (2024)
di: Yu, Xueyang, et al.
Pubblicazione: (2024)
Interpreting the Weight Space of Customized Diffusion Models
di: Dravid, Amil, et al.
Pubblicazione: (2024)
di: Dravid, Amil, et al.
Pubblicazione: (2024)
MV-MOS: Multi-View Feature Fusion for 3D Moving Object Segmentation
di: Cheng, Jintao, et al.
Pubblicazione: (2024)
di: Cheng, Jintao, et al.
Pubblicazione: (2024)
xT: Nested Tokenization for Larger Context in Large Images
di: Gupta, Ritwik, et al.
Pubblicazione: (2024)
di: Gupta, Ritwik, et al.
Pubblicazione: (2024)
The More You See in 2D, the More You Perceive in 3D
di: Han, Xinyang, et al.
Pubblicazione: (2024)
di: Han, Xinyang, et al.
Pubblicazione: (2024)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
di: Wang, Renhao, et al.
Pubblicazione: (2025)
di: Wang, Renhao, et al.
Pubblicazione: (2025)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
di: Cai, Xinhao, et al.
Pubblicazione: (2025)
Steering CLIP's vision transformer with sparse autoencoders
di: Joseph, Sonia, et al.
Pubblicazione: (2025)
di: Joseph, Sonia, et al.
Pubblicazione: (2025)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
di: Tian, Ran, et al.
Pubblicazione: (2023)
di: Tian, Ran, et al.
Pubblicazione: (2023)
HalluGen: Synthesizing Realistic and Controllable Hallucinations for Evaluating Image Restoration
di: Kim, Seunghoi, et al.
Pubblicazione: (2025)
di: Kim, Seunghoi, et al.
Pubblicazione: (2025)
Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
di: Zhao, Chen, et al.
Pubblicazione: (2024)
di: Zhao, Chen, et al.
Pubblicazione: (2024)
PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D Data
di: Yang, ChangHee, et al.
Pubblicazione: (2025)
di: Yang, ChangHee, et al.
Pubblicazione: (2025)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
di: Li, Ying, et al.
Pubblicazione: (2025)
di: Li, Ying, et al.
Pubblicazione: (2025)
SAM 3D: 3Dfy Anything in Images
di: SAM 3D Team, et al.
Pubblicazione: (2025)
di: SAM 3D Team, et al.
Pubblicazione: (2025)
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024)
di: Gao, Ruiyuan, et al.
Pubblicazione: (2024)
Synthesizing Physically Plausible Human Motions in 3D Scenes
di: Pan, Liang, et al.
Pubblicazione: (2023)
di: Pan, Liang, et al.
Pubblicazione: (2023)
3DiFACE: Synthesizing and Editing Holistic 3D Facial Animation
di: Thambiraja, Balamurugan, et al.
Pubblicazione: (2025)
di: Thambiraja, Balamurugan, et al.
Pubblicazione: (2025)
Estimating Body and Hand Motion in an Ego-sensed World
di: Yi, Brent, et al.
Pubblicazione: (2024)
di: Yi, Brent, et al.
Pubblicazione: (2024)
Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
di: Chen, Yuxin, et al.
Pubblicazione: (2025)
di: Chen, Yuxin, et al.
Pubblicazione: (2025)
Chain of Time: In-Context Physical Simulation with Image Generation Models
di: Wang, YingQiao, et al.
Pubblicazione: (2025)
di: Wang, YingQiao, et al.
Pubblicazione: (2025)
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
di: Liu, Hongyuan, et al.
Pubblicazione: (2026)
di: Liu, Hongyuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
An Empirical Study of Autoregressive Pre-training from Videos
di: Rajasegaran, Jathushan, et al.
Pubblicazione: (2025) -
Interpreting CLIP's Image Representation via Text-Based Decomposition
di: Gandelsman, Yossi, et al.
Pubblicazione: (2023) -
Scaling Properties of Diffusion Models for Perceptual Tasks
di: Ravishankar, Rahul, et al.
Pubblicazione: (2024) -
Vision Transformers Don't Need Trained Registers
di: Jiang, Nick, et al.
Pubblicazione: (2025) -
Gaussian Masked Autoencoders
di: Rajasegaran, Jathushan, et al.
Pubblicazione: (2025)