What Are You Doing? A Closer Look at Controllable Human Video Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Bugliarello, Emanuele, Arnab, Anurag, Paiss, Roni, Kindermans, Pieter-Jan, Schmid, Cordelia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dense Video Object Captioning from Disjoint Supervision
di: Zhou, Xingyi, et al.
Pubblicazione: (2023)
di: Zhou, Xingyi, et al.
Pubblicazione: (2023)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025)
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
Streaming Dense Video Captioning
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
di: Uijlings, Jasper, et al.
Pubblicazione: (2025)
di: Uijlings, Jasper, et al.
Pubblicazione: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2025)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2025)
Versatile Editing of Video Content, Actions, and Dynamics without Training
di: Kulikov, Vladimir, et al.
Pubblicazione: (2026)
di: Kulikov, Vladimir, et al.
Pubblicazione: (2026)
BrickNet: Graph-Backed Generative Brick Assembly
di: Kulits, Peter, et al.
Pubblicazione: (2026)
di: Kulits, Peter, et al.
Pubblicazione: (2026)
Audiovisual Masked Autoencoders
di: Georgescu, Mariana-Iuliana, et al.
Pubblicazione: (2022)
di: Georgescu, Mariana-Iuliana, et al.
Pubblicazione: (2022)
Still-Moving: Customized Video Generation without Customized Video Data
di: Chefer, Hila, et al.
Pubblicazione: (2024)
di: Chefer, Hila, et al.
Pubblicazione: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
di: Khan, Zeeshan, et al.
Pubblicazione: (2025)
di: Khan, Zeeshan, et al.
Pubblicazione: (2025)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
di: Ghosh, Partha, et al.
Pubblicazione: (2024)
di: Ghosh, Partha, et al.
Pubblicazione: (2024)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
di: Sokar, Ghada, et al.
Pubblicazione: (2025)
di: Sokar, Ghada, et al.
Pubblicazione: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
Learning text-to-video retrieval from image captioning
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
di: Chen, Zerui, et al.
Pubblicazione: (2024)
di: Chen, Zerui, et al.
Pubblicazione: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
di: Garcia, Ricardo, et al.
Pubblicazione: (2024)
di: Garcia, Ricardo, et al.
Pubblicazione: (2024)
FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
di: Nkegoum, Manuel, et al.
Pubblicazione: (2025)
di: Nkegoum, Manuel, et al.
Pubblicazione: (2025)
A Closer Look at the Explainability of Contrastive Language-Image Pre-training
di: Li, Yi, et al.
Pubblicazione: (2023)
di: Li, Yi, et al.
Pubblicazione: (2023)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
Feature Augmentation for Self-supervised Contrastive Learning: A Closer Look
di: Zhang, Yong, et al.
Pubblicazione: (2024)
di: Zhang, Yong, et al.
Pubblicazione: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2023)
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2023)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation
di: Li, Lin, et al.
Pubblicazione: (2025)
di: Li, Lin, et al.
Pubblicazione: (2025)
Dense Optical Tracking: Connecting the Dots
di: Moing, Guillaume Le, et al.
Pubblicazione: (2023)
di: Moing, Guillaume Le, et al.
Pubblicazione: (2023)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
di: Zhang, David Junhao, et al.
Pubblicazione: (2024)
di: Zhang, David Junhao, et al.
Pubblicazione: (2024)
A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models
di: Silva-Rodríguez, Julio, et al.
Pubblicazione: (2023)
di: Silva-Rodríguez, Julio, et al.
Pubblicazione: (2023)
A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification
di: Marks, Markus, et al.
Pubblicazione: (2024)
di: Marks, Markus, et al.
Pubblicazione: (2024)
Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
di: Zhong, Xinhao, et al.
Pubblicazione: (2025)
di: Zhong, Xinhao, et al.
Pubblicazione: (2025)
LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoyu, et al.
Pubblicazione: (2025)
Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation
di: Chen, Yuanhong, et al.
Pubblicazione: (2023)
di: Chen, Yuanhong, et al.
Pubblicazione: (2023)
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
di: Kamenetsky, Ronen, et al.
Pubblicazione: (2025)
di: Kamenetsky, Ronen, et al.
Pubblicazione: (2025)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024)
di: Taghipour, Ashkan, et al.
Pubblicazione: (2024)
VidPanos: Generative Panoramic Videos from Casual Panning Videos
di: Ma, Jingwei, et al.
Pubblicazione: (2024)
di: Ma, Jingwei, et al.
Pubblicazione: (2024)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dense Video Object Captioning from Disjoint Supervision
di: Zhou, Xingyi, et al.
Pubblicazione: (2023) -
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025) -
Time-, Memory- and Parameter-Efficient Visual Adaptation
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024) -
Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024) -
Streaming Dense Video Captioning
di: Zhou, Xingyi, et al.
Pubblicazione: (2024)