First Frame Is the Place to Go for Video Content Customization
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Jingxi, Li, Zongxia, Liu, Zhichao, Shi, Guangyao, Wu, Xiyang, Liu, Fuxiao, Fermuller, Cornelia, Feng, Brandon Y., Aloimonos, Yiannis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diving Deep into the Motion Representation of Video-Text Models
by: Devaraj, Chinmaya, et al.
Published: (2024)
by: Devaraj, Chinmaya, et al.
Published: (2024)
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
by: Chen, Jingxi, et al.
Published: (2025)
by: Chen, Jingxi, et al.
Published: (2025)
Active Human Pose Estimation via an Autonomous UAV Agent
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
by: Chen, Jingxi, et al.
Published: (2024)
by: Chen, Jingxi, et al.
Published: (2024)
Learning Normal Flow Directly From Event Neighborhoods
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
Decodable and Sample Invariant Continuous Object Encoder
by: Yuan, Dehao, et al.
Published: (2023)
by: Yuan, Dehao, et al.
Published: (2023)
Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
by: Wu, Jiayi, et al.
Published: (2026)
by: Wu, Jiayi, et al.
Published: (2026)
MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
by: Wu, Jiayi, et al.
Published: (2024)
by: Wu, Jiayi, et al.
Published: (2024)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
by: Yuan, Dehao, et al.
Published: (2024)
by: Yuan, Dehao, et al.
Published: (2024)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2024)
Event3DGS: Event-Based 3D Gaussian Splatting for High-Speed Robot Egomotion
by: Xiong, Tianyi, et al.
Published: (2024)
by: Xiong, Tianyi, et al.
Published: (2024)
ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image
by: Luo, Dongyu, et al.
Published: (2025)
by: Luo, Dongyu, et al.
Published: (2025)
AcTExplore: Active Tactile Exploration of Unknown Objects
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2023)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2023)
Object Pose Estimation through Dexterous Touch
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2025)
by: Shahidzadeh, Amir-Hossein, et al.
Published: (2025)
Single-Step Latent Diffusion for Underwater Image Restoration
by: Wu, Jiayi, et al.
Published: (2025)
by: Wu, Jiayi, et al.
Published: (2025)
Maelstrom Networks
by: Evanusa, Matthew, et al.
Published: (2024)
by: Evanusa, Matthew, et al.
Published: (2024)
Embodied Visuomotor Representation
by: Burner, Levi, et al.
Published: (2024)
by: Burner, Levi, et al.
Published: (2024)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
by: Wu, Xiyang, et al.
Published: (2025)
by: Wu, Xiyang, et al.
Published: (2025)
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
by: Li, Zongxia, et al.
Published: (2026)
by: Li, Zongxia, et al.
Published: (2026)
Exploring Emerging Trends and Research Opportunities in Visual Place Recognition
by: Gasteratos, Antonios, et al.
Published: (2024)
by: Gasteratos, Antonios, et al.
Published: (2024)
A Real-Time Event-Based Normal Flow Estimator
by: Yuan, Dehao, et al.
Published: (2025)
by: Yuan, Dehao, et al.
Published: (2025)
Odometry Without Correspondence from Inertially Constrained Ruled Surfaces
by: Zhu, Chenqi, et al.
Published: (2025)
by: Zhu, Chenqi, et al.
Published: (2025)
Microsaccade-inspired Event Camera for Robotics
by: He, Botao, et al.
Published: (2024)
by: He, Botao, et al.
Published: (2024)
Artificial Microsaccade Compensation: Stable Vision for an Ornithopter
by: Burner, Levi, et al.
Published: (2025)
by: Burner, Levi, et al.
Published: (2025)
Self-Rewarding Vision-Language Model via Reasoning Decomposition
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Motion Segmentation and Egomotion Estimation from Event-Based Normal Flow
by: Hua, Zhiyuan, et al.
Published: (2025)
by: Hua, Zhiyuan, et al.
Published: (2025)
Whale Detection Enhancement through Synthetic Satellite Images
by: Gaur, Akshaj, et al.
Published: (2023)
by: Gaur, Akshaj, et al.
Published: (2023)
EREBUS: End-to-end Robust Event Based Underwater Simulation
by: Kyatham, Hitesh, et al.
Published: (2025)
by: Kyatham, Hitesh, et al.
Published: (2025)
ViewActive: Active viewpoint optimization from a single image
by: Wu, Jiayi, et al.
Published: (2024)
by: Wu, Jiayi, et al.
Published: (2024)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
by: Ren, Yixuan, et al.
Published: (2024)
by: Ren, Yixuan, et al.
Published: (2024)
OysterNet: Enhanced Oyster Detection Using Simulation
by: Lin, Xiaomin, et al.
Published: (2022)
by: Lin, Xiaomin, et al.
Published: (2022)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
by: Guan, Tianrui, et al.
Published: (2023)
by: Guan, Tianrui, et al.
Published: (2023)
Place Anything into Any Video
by: Liu, Ziling, et al.
Published: (2024)
by: Liu, Ziling, et al.
Published: (2024)
TRACE: Object Motion Editing in Videos with First-Frame Trajectory Guidance
by: Phung, Quynh, et al.
Published: (2026)
by: Phung, Quynh, et al.
Published: (2026)
CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras
by: Shah, Sachin, et al.
Published: (2024)
by: Shah, Sachin, et al.
Published: (2024)
Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-scale Complex Unknown Environments
by: He, Botao, et al.
Published: (2024)
by: He, Botao, et al.
Published: (2024)
EmbodiSwap for Zero-Shot Robot Imitation Learning
by: Dessalene, Eadom, et al.
Published: (2025)
by: Dessalene, Eadom, et al.
Published: (2025)
Similar Items
-
Diving Deep into the Motion Representation of Video-Text Models
by: Devaraj, Chinmaya, et al.
Published: (2024) -
TimeRewind: Rewinding Time with Image-and-Events Video Diffusion
by: Chen, Jingxi, et al.
Published: (2024) -
From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
by: Chen, Jingxi, et al.
Published: (2025) -
Active Human Pose Estimation via an Autonomous UAV Agent
by: Chen, Jingxi, et al.
Published: (2024) -
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
by: Chen, Jingxi, et al.
Published: (2024)