LAMP: Language-Assisted Motion Planning for Controllable Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kizil, Muhammed Burak, Sanli, Enes, Mitra, Niloy J., Erdem, Erkut, Erdem, Aykut, Ceylan, Duygu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025)
by: Çapuk, Hakan, et al.
Published: (2025)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
MonetGPT: Solving Puzzles Enhances MLLMs' Image Retouching Skills
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
by: Dutt, Niladri Shekhar, et al.
Published: (2025)
JOG3R: Towards 3D-Consistent Video Generators
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
BLiSS: Bootstrapped Linear Shape Space
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
by: Muralikrishnan, Sanjeev, et al.
Published: (2023)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
by: Jeong, Hyeonho, et al.
Published: (2024)
by: Jeong, Hyeonho, et al.
Published: (2024)
Infrared Domain Adaptation with Zero-Shot Quantization
by: Sevsay, Burak, et al.
Published: (2024)
by: Sevsay, Burak, et al.
Published: (2024)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
SuperGaussian: Repurposing Video Models for 3D Super Resolution
by: Shen, Yuan, et al.
Published: (2024)
by: Shen, Yuan, et al.
Published: (2024)
Boosting Camera Motion Control for Video Diffusion Transformers
by: Cheong, Soon Yau, et al.
Published: (2024)
by: Cheong, Soon Yau, et al.
Published: (2024)
SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
by: Kan, Mia, et al.
Published: (2025)
by: Kan, Mia, et al.
Published: (2025)
GeoFusionLRM: Geometry-Aware Self-Correction for Consistent 3D Reconstruction
by: Yildirim, Ahmet Burak, et al.
Published: (2026)
by: Yildirim, Ahmet Burak, et al.
Published: (2026)
WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
MD-ProjTex: Texturing 3D Shapes with Multi-Diffusion Projection
by: Yildirim, Ahmet Burak, et al.
Published: (2025)
by: Yildirim, Ahmet Burak, et al.
Published: (2025)
VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
by: Koo, Juil, et al.
Published: (2025)
by: Koo, Juil, et al.
Published: (2025)
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
by: Sabathier, Remy, et al.
Published: (2024)
by: Sabathier, Remy, et al.
Published: (2024)
How to Augment for Atmospheric Turbulence Effects on Thermal Adapted Object Detection Models?
by: Uzun, Engin, et al.
Published: (2024)
by: Uzun, Engin, et al.
Published: (2024)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
3D Stylization via Large Reconstruction Model
by: Oztas, Ipek, et al.
Published: (2025)
by: Oztas, Ipek, et al.
Published: (2025)
FUSION: Full-Body Unified Motion Prior for Body and Hands via Diffusion
by: Duran, Enes, et al.
Published: (2026)
by: Duran, Enes, et al.
Published: (2026)
ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
LoST: Level of Semantics Tokenization for 3D Shapes
by: Dutt, Niladri Shekhar, et al.
Published: (2026)
by: Dutt, Niladri Shekhar, et al.
Published: (2026)
Motion Modes: What Could Happen Next?
by: Pandey, Karran, et al.
Published: (2024)
by: Pandey, Karran, et al.
Published: (2024)
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
by: Sabathier, Remy, et al.
Published: (2026)
by: Sabathier, Remy, et al.
Published: (2026)
FloAt: Flow Warping of Self-Attention for Clothing Animation Generation
by: Mishra, Swasti Shreya, et al.
Published: (2024)
by: Mishra, Swasti Shreya, et al.
Published: (2024)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
Similar Items
-
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026) -
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025) -
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024) -
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
by: Çapuk, Hakan, et al.
Published: (2025) -
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)