Diving Deep into the Motion Representation of Video-Text Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Devaraj, Chinmaya, Fermuller, Cornelia, Aloimonos, Yiannis |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
di: Wu, Jiayi, et al.
Pubblicazione: (2024)
di: Wu, Jiayi, et al.
Pubblicazione: (2024)
Decodable and Sample Invariant Continuous Object Encoder
di: Yuan, Dehao, et al.
Pubblicazione: (2023)
di: Yuan, Dehao, et al.
Pubblicazione: (2023)
Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
di: Wu, Jiayi, et al.
Pubblicazione: (2026)
di: Wu, Jiayi, et al.
Pubblicazione: (2026)
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
di: Yuan, Dehao, et al.
Pubblicazione: (2024)
di: Yuan, Dehao, et al.
Pubblicazione: (2024)
Active Human Pose Estimation via an Autonomous UAV Agent
di: Chen, Jingxi, et al.
Pubblicazione: (2024)
di: Chen, Jingxi, et al.
Pubblicazione: (2024)
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
di: Chen, Jingxi, et al.
Pubblicazione: (2024)
di: Chen, Jingxi, et al.
Pubblicazione: (2024)
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
di: Chen, Jingxi, et al.
Pubblicazione: (2025)
di: Chen, Jingxi, et al.
Pubblicazione: (2025)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
di: Chen, Jingxi, et al.
Pubblicazione: (2024)
di: Chen, Jingxi, et al.
Pubblicazione: (2024)
FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
di: Shahidzadeh, Amir-Hossein, et al.
Pubblicazione: (2024)
di: Shahidzadeh, Amir-Hossein, et al.
Pubblicazione: (2024)
Learning Normal Flow Directly From Event Neighborhoods
di: Yuan, Dehao, et al.
Pubblicazione: (2024)
di: Yuan, Dehao, et al.
Pubblicazione: (2024)
First Frame Is the Place to Go for Video Content Customization
di: Chen, Jingxi, et al.
Pubblicazione: (2025)
di: Chen, Jingxi, et al.
Pubblicazione: (2025)
Event3DGS: Event-Based 3D Gaussian Splatting for High-Speed Robot Egomotion
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
di: Xiong, Tianyi, et al.
Pubblicazione: (2024)
ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image
di: Luo, Dongyu, et al.
Pubblicazione: (2025)
di: Luo, Dongyu, et al.
Pubblicazione: (2025)
AcTExplore: Active Tactile Exploration of Unknown Objects
di: Shahidzadeh, Amir-Hossein, et al.
Pubblicazione: (2023)
di: Shahidzadeh, Amir-Hossein, et al.
Pubblicazione: (2023)
Object Pose Estimation through Dexterous Touch
di: Shahidzadeh, Amir-Hossein, et al.
Pubblicazione: (2025)
di: Shahidzadeh, Amir-Hossein, et al.
Pubblicazione: (2025)
Single-Step Latent Diffusion for Underwater Image Restoration
di: Wu, Jiayi, et al.
Pubblicazione: (2025)
di: Wu, Jiayi, et al.
Pubblicazione: (2025)
Motion Segmentation and Egomotion Estimation from Event-Based Normal Flow
di: Hua, Zhiyuan, et al.
Pubblicazione: (2025)
di: Hua, Zhiyuan, et al.
Pubblicazione: (2025)
A Real-Time Event-Based Normal Flow Estimator
di: Yuan, Dehao, et al.
Pubblicazione: (2025)
di: Yuan, Dehao, et al.
Pubblicazione: (2025)
Odometry Without Correspondence from Inertially Constrained Ruled Surfaces
di: Zhu, Chenqi, et al.
Pubblicazione: (2025)
di: Zhu, Chenqi, et al.
Pubblicazione: (2025)
Artificial Microsaccade Compensation: Stable Vision for an Ornithopter
di: Burner, Levi, et al.
Pubblicazione: (2025)
di: Burner, Levi, et al.
Pubblicazione: (2025)
OysterNet: Enhanced Oyster Detection Using Simulation
di: Lin, Xiaomin, et al.
Pubblicazione: (2022)
di: Lin, Xiaomin, et al.
Pubblicazione: (2022)
EREBUS: End-to-end Robust Event Based Underwater Simulation
di: Kyatham, Hitesh, et al.
Pubblicazione: (2025)
di: Kyatham, Hitesh, et al.
Pubblicazione: (2025)
Real-time Motion Segmentation with Event-based Normal Flow
di: Zhong, Sheng, et al.
Pubblicazione: (2026)
di: Zhong, Sheng, et al.
Pubblicazione: (2026)
Embodiment: Self-Supervised Depth Estimation Based on Camera Models
di: Zhang, Jinchang, et al.
Pubblicazione: (2024)
di: Zhang, Jinchang, et al.
Pubblicazione: (2024)
VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference
di: Yoo, Seong Jong, et al.
Pubblicazione: (2024)
di: Yoo, Seong Jong, et al.
Pubblicazione: (2024)
Embodied Visuomotor Representation
di: Burner, Levi, et al.
Pubblicazione: (2024)
di: Burner, Levi, et al.
Pubblicazione: (2024)
Maelstrom Networks
di: Evanusa, Matthew, et al.
Pubblicazione: (2024)
di: Evanusa, Matthew, et al.
Pubblicazione: (2024)
Whale Detection Enhancement through Synthetic Satellite Images
di: Gaur, Akshaj, et al.
Pubblicazione: (2023)
di: Gaur, Akshaj, et al.
Pubblicazione: (2023)
Microsaccade-inspired Event Camera for Robotics
di: He, Botao, et al.
Pubblicazione: (2024)
di: He, Botao, et al.
Pubblicazione: (2024)
Exploring Emerging Trends and Research Opportunities in Visual Place Recognition
di: Gasteratos, Antonios, et al.
Pubblicazione: (2024)
di: Gasteratos, Antonios, et al.
Pubblicazione: (2024)
CodedVO: Coded Visual Odometry
di: Shah, Sachin, et al.
Pubblicazione: (2024)
di: Shah, Sachin, et al.
Pubblicazione: (2024)
Recent Event Camera Innovations: A Survey
di: Chakravarthi, Bharatesh, et al.
Pubblicazione: (2024)
di: Chakravarthi, Bharatesh, et al.
Pubblicazione: (2024)
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
di: Chefer, Hila, et al.
Pubblicazione: (2025)
di: Chefer, Hila, et al.
Pubblicazione: (2025)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
di: Tang, Yolo Y., et al.
Pubblicazione: (2025)
di: Tang, Yolo Y., et al.
Pubblicazione: (2025)
EmbodiSwap for Zero-Shot Robot Imitation Learning
di: Dessalene, Eadom, et al.
Pubblicazione: (2025)
di: Dessalene, Eadom, et al.
Pubblicazione: (2025)
Future Aspects in Human Action Recognition: Exploring Emerging Techniques and Ethical Influences
di: Gasteratos, Antonios, et al.
Pubblicazione: (2024)
di: Gasteratos, Antonios, et al.
Pubblicazione: (2024)
Monocular Reconstruction of Neural Tactile Fields
di: Mantripragada, Pavan, et al.
Pubblicazione: (2026)
di: Mantripragada, Pavan, et al.
Pubblicazione: (2026)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
MotionBooth: Motion-Aware Customized Text-to-Video Generation
di: Wu, Jianzong, et al.
Pubblicazione: (2024)
di: Wu, Jianzong, et al.
Pubblicazione: (2024)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
di: Ren, Yixuan, et al.
Pubblicazione: (2024)
di: Ren, Yixuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
di: Wu, Jiayi, et al.
Pubblicazione: (2024) -
Decodable and Sample Invariant Continuous Object Encoder
di: Yuan, Dehao, et al.
Pubblicazione: (2023) -
Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
di: Wu, Jiayi, et al.
Pubblicazione: (2026) -
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
di: Yuan, Dehao, et al.
Pubblicazione: (2024) -
Active Human Pose Estimation via an Autonomous UAV Agent
di: Chen, Jingxi, et al.
Pubblicazione: (2024)