Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jiayi, Cai, Haoming, Fermuller, Cornelia, Metzler, Christopher, Aloimonos, Yiannis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
by: Wu, Jiayi, et al.
Published: (2024)
by: Wu, Jiayi, et al.
Published: (2024)
Diving Deep into the Motion Representation of Video-Text Models
by: Devaraj, Chinmaya, et al.
Published: (2024)
by: Devaraj, Chinmaya, et al.
Published: (2024)
Event3DGS: Event-Based 3D Gaussian Splatting for High-Speed Robot Egomotion
by: Xiong, Tianyi, et al.
Published: (2024)
by: Xiong, Tianyi, et al.
Published: (2024)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
Single-Step Latent Diffusion for Underwater Image Restoration
by: Wu, Jiayi, et al.
Published: (2025)
by: Wu, Jiayi, et al.
Published: (2025)
Decodable and Sample Invariant Continuous Object Encoder
by: Yuan, Dehao, et al.
Published: (2023)
by: Yuan, Dehao, et al.
Published: (2023)
Learning Normal Flow Directly From Event Neighborhoods
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
Active Human Pose Estimation via an Autonomous UAV Agent
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
A Real-Time Event-Based Normal Flow Estimator
by: Yuan, Dehao, et al.
Published: (2025)
by: Yuan, Dehao, et al.
Published: (2025)
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
by: Chen, Jingxi, et al.
Published: (2025)
by: Chen, Jingxi, et al.
Published: (2025)
FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
First Frame Is the Place to Go for Video Content Customization
by: Chen, Jingxi, et al.
Published: (2025)
by: Chen, Jingxi, et al.
Published: (2025)
CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras
by: Shah, Sachin, et al.
Published: (2024)
by: Shah, Sachin, et al.
Published: (2024)
ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image
by: Luo, Dongyu, et al.
Published: (2025)
by: Luo, Dongyu, et al.
Published: (2025)
AcTExplore: Active Tactile Exploration of Unknown Objects
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2023)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2023)
Object Pose Estimation through Dexterous Touch
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2025)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2025)
CodedVO: Coded Visual Odometry
by: Shah, Sachin, et al.
Published: (2024)
by: Shah, Sachin, et al.
Published: (2024)
Flash-Split: 2D Reflection Removal with Flash Cues and Latent Diffusion Separation
by: Wang, Tianfu, et al.
Published: (2024)
by: Wang, Tianfu, et al.
Published: (2024)
Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models
by: Cai, Zijie, et al.
Published: (2025)
by: Cai, Zijie, et al.
Published: (2025)
Maelstrom Networks
by: Evanusa, Matthew, et al.
Published: (2024)
by: Evanusa, Matthew, et al.
Published: (2024)
Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
by: Cai, Haoming, et al.
Published: (2025)
by: Cai, Haoming, et al.
Published: (2025)
EmbodiSwap for Zero-Shot Robot Imitation Learning
by: Dessalene, Eadom, et al.
Published: (2025)
by: Dessalene, Eadom, et al.
Published: (2025)
Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning
by: Liu, Haofeng, et al.
Published: (2024)
by: Liu, Haofeng, et al.
Published: (2024)
Monocular Reconstruction of Neural Tactile Fields
by: Mantripragada, Pavan, et al.
Published: (2026)
by: Mantripragada, Pavan, et al.
Published: (2026)
Odometry Without Correspondence from Inertially Constrained Ruled Surfaces
by: Zhu, Chenqi, et al.
Published: (2025)
by: Zhu, Chenqi, et al.
Published: (2025)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
by: Lin, Yijing, et al.
Published: (2025)
by: Lin, Yijing, et al.
Published: (2025)
GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
by: Deng, Ken, et al.
Published: (2025)
by: Deng, Ken, et al.
Published: (2025)
Real-time Motion Segmentation with Event-based Normal Flow
by: Zhong, Sheng, et al.
Published: (2026)
by: Zhong, Sheng, et al.
Published: (2026)
EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
by: Zeng, Chengxi, et al.
Published: (2025)
by: Zeng, Chengxi, et al.
Published: (2025)
RealWonder: Real-Time Physical Action-Conditioned Video Generation
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
Adapting VACE for Real-Time Autoregressive Video Diffusion
by: Fosdick, Ryan
Published: (2026)
by: Fosdick, Ryan
Published: (2026)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Artificial Microsaccade Compensation: Stable Vision for an Ornithopter
by: Burner, Levi, et al.
Published: (2025)
by: Burner, Levi, et al.
Published: (2025)
H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
SAM2 for Image and Video Segmentation: A Comprehensive Survey
by: Jiaxing, Zhang, et al.
Published: (2025)
by: Jiaxing, Zhang, et al.
Published: (2025)
CamSAM2: Segment Anything Accurately in Camouflaged Videos
by: Zhou, Yuli, et al.
Published: (2025)
by: Zhou, Yuli, et al.
Published: (2025)
Similar Items
-
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
by: Chen, Jingxi, et al.
Published: (2024) -
MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
by: Wu, Jiayi, et al.
Published: (2024) -
Diving Deep into the Motion Representation of Video-Text Models
by: Devaraj, Chinmaya, et al.
Published: (2024) -
Event3DGS: Event-Based 3D Gaussian Splatting for High-Speed Robot Egomotion
by: Xiong, Tianyi, et al.
Published: (2024) -
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
by: Chen, Jingxi, et al.
Published: (2024)