Learning Action and Reasoning-Centric Image Editing from Videos and Simulations
Fuente:
arXiv
Saved in:
| Main Authors: | Krojer, Benno, Vattikonda, Dheeraj, Lara, Luis, Jampani, Varun, Portelance, Eva, Pal, Christopher, Reddy, Siva |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Promise of RL for Autoregressive Image Editing
by: Ahmadi, Saba, et al.
Published: (2025)
by: Ahmadi, Saba, et al.
Published: (2025)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
by: Krojer, Benno, et al.
Published: (2026)
by: Krojer, Benno, et al.
Published: (2026)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
Human Video Generation from a Single Image with 3D Pose and View Control
by: Wang, Tiantian, et al.
Published: (2026)
by: Wang, Tiantian, et al.
Published: (2026)
SLACK: Attacking LiDAR-based SLAM with Adversarial Point Injections
by: Kumar, Prashant, et al.
Published: (2025)
by: Kumar, Prashant, et al.
Published: (2025)
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing
by: Li, Hengjia, et al.
Published: (2026)
by: Li, Hengjia, et al.
Published: (2026)
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
by: Krojer, Benno, et al.
Published: (2025)
by: Krojer, Benno, et al.
Published: (2025)
SyncNoise: Geometrically Consistent Noise Prediction for Text-based 3D Scene Editing
by: Li, Ruihuang, et al.
Published: (2024)
by: Li, Ruihuang, et al.
Published: (2024)
ICE-G: Image Conditional Editing of 3D Gaussian Splats
by: Jaganathan, Vishnu, et al.
Published: (2024)
by: Jaganathan, Vishnu, et al.
Published: (2024)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
by: Yao, Chun-Han, et al.
Published: (2025)
by: Yao, Chun-Han, et al.
Published: (2025)
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
SV3.3B: A Sports Video Understanding Model for Action Recognition
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
Stable Video-Driven Portraits
by: R., Mallikarjun B., et al.
Published: (2025)
by: R., Mallikarjun B., et al.
Published: (2025)
Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
by: Rowles, Ciara, et al.
Published: (2025)
by: Rowles, Ciara, et al.
Published: (2025)
FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image
by: Yin, Fei, et al.
Published: (2025)
by: Yin, Fei, et al.
Published: (2025)
MARBLE: Material Recomposition and Blending in CLIP-Space
by: Cheng, Ta-Ying, et al.
Published: (2025)
by: Cheng, Ta-Ying, et al.
Published: (2025)
FROMAT: Multiview Material Appearance Transfer via Few-Shot Self-Attention Adaptation
by: Kompanowski, Hubert, et al.
Published: (2025)
by: Kompanowski, Hubert, et al.
Published: (2025)
SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement
by: Boss, Mark, et al.
Published: (2024)
by: Boss, Mark, et al.
Published: (2024)
SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
Detecting Localized Deepfake Manipulations Using Action Unit-Guided Video Representations
by: Anand, Tharun, et al.
Published: (2025)
by: Anand, Tharun, et al.
Published: (2025)
PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
by: Nguyen, Hieu T., et al.
Published: (2024)
by: Nguyen, Hieu T., et al.
Published: (2024)
SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion
by: Voleti, Vikram, et al.
Published: (2024)
by: Voleti, Vikram, et al.
Published: (2024)
SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
by: Engelhardt, Andreas, et al.
Published: (2025)
by: Engelhardt, Andreas, et al.
Published: (2025)
Block Cascading: Training Free Acceleration of Block-Causal Video Models
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
by: Bandyopadhyay, Hmrishav, et al.
Published: (2025)
Action Reimagined: Text-to-Pose Video Editing for Dynamic Human Actions
by: Wang, Lan, et al.
Published: (2024)
by: Wang, Lan, et al.
Published: (2024)
Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition
by: Li, Xunsong, et al.
Published: (2024)
by: Li, Xunsong, et al.
Published: (2024)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
by: Liu, Yuanxin, et al.
Published: (2025)
by: Liu, Yuanxin, et al.
Published: (2025)
Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
by: Chatterjee, Agneet, et al.
Published: (2025)
by: Chatterjee, Agneet, et al.
Published: (2025)
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
by: Xie, Yiming, et al.
Published: (2024)
by: Xie, Yiming, et al.
Published: (2024)
ZeroShape: Regression-based Zero-shot Shape Reconstruction
by: Huang, Zixuan, et al.
Published: (2023)
by: Huang, Zixuan, et al.
Published: (2023)
3D Congealing: 3D-Aware Image Alignment in the Wild
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
ReSWD: ReSTIR'd, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction
by: Boss, Mark, et al.
Published: (2025)
by: Boss, Mark, et al.
Published: (2025)
Similar Items
-
The Promise of RL for Autoregressive Image Editing
by: Ahmadi, Saba, et al.
Published: (2025) -
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
by: Hu, Tao, et al.
Published: (2026) -
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
by: Krojer, Benno, et al.
Published: (2026) -
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024) -
Human Video Generation from a Single Image with 3D Pose and View Control
by: Wang, Tiantian, et al.
Published: (2026)