What Are You Doing? A Closer Look at Controllable Human Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Bugliarello, Emanuele, Arnab, Anurag, Paiss, Roni, Kindermans, Pieter-Jan, Schmid, Cordelia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dense Video Object Captioning from Disjoint Supervision
por: Zhou, Xingyi, et al.
Publicado: (2023)
por: Zhou, Xingyi, et al.
Publicado: (2023)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
por: Wysoczańska, Monika, et al.
Publicado: (2025)
por: Wysoczańska, Monika, et al.
Publicado: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
por: Mercea, Otniel-Bogdan, et al.
Publicado: (2024)
por: Mercea, Otniel-Bogdan, et al.
Publicado: (2024)
Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2024)
por: Kazakos, Evangelos, et al.
Publicado: (2024)
Streaming Dense Video Captioning
por: Zhou, Xingyi, et al.
Publicado: (2024)
por: Zhou, Xingyi, et al.
Publicado: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
por: Uijlings, Jasper, et al.
Publicado: (2025)
por: Uijlings, Jasper, et al.
Publicado: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2025)
por: Kazakos, Evangelos, et al.
Publicado: (2025)
Versatile Editing of Video Content, Actions, and Dynamics without Training
por: Kulikov, Vladimir, et al.
Publicado: (2026)
por: Kulikov, Vladimir, et al.
Publicado: (2026)
BrickNet: Graph-Backed Generative Brick Assembly
por: Kulits, Peter, et al.
Publicado: (2026)
por: Kulits, Peter, et al.
Publicado: (2026)
Audiovisual Masked Autoencoders
por: Georgescu, Mariana-Iuliana, et al.
Publicado: (2022)
por: Georgescu, Mariana-Iuliana, et al.
Publicado: (2022)
Still-Moving: Customized Video Generation without Customized Video Data
por: Chefer, Hila, et al.
Publicado: (2024)
por: Chefer, Hila, et al.
Publicado: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
por: Khan, Zeeshan, et al.
Publicado: (2025)
por: Khan, Zeeshan, et al.
Publicado: (2025)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
por: Ventura, Lucas, et al.
Publicado: (2025)
por: Ventura, Lucas, et al.
Publicado: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
por: Ghosh, Partha, et al.
Publicado: (2024)
por: Ghosh, Partha, et al.
Publicado: (2024)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
por: Ventura, Lucas, et al.
Publicado: (2023)
por: Ventura, Lucas, et al.
Publicado: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
por: Caron, Mathilde, et al.
Publicado: (2024)
por: Caron, Mathilde, et al.
Publicado: (2024)
Continual Learning in Vision-Language Models via Aligned Model Merging
por: Sokar, Ghada, et al.
Publicado: (2025)
por: Sokar, Ghada, et al.
Publicado: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
por: Bousselham, Walid, et al.
Publicado: (2025)
por: Bousselham, Walid, et al.
Publicado: (2025)
Learning text-to-video retrieval from image captioning
por: Ventura, Lucas, et al.
Publicado: (2024)
por: Ventura, Lucas, et al.
Publicado: (2024)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
por: Chen, Zerui, et al.
Publicado: (2024)
por: Chen, Zerui, et al.
Publicado: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
por: Garcia, Ricardo, et al.
Publicado: (2024)
por: Garcia, Ricardo, et al.
Publicado: (2024)
FSMODNet: A Closer Look at Few-Shot Detection in Multispectral Data
por: Nkegoum, Manuel, et al.
Publicado: (2025)
por: Nkegoum, Manuel, et al.
Publicado: (2025)
A Closer Look at the Explainability of Contrastive Language-Image Pre-training
por: Li, Yi, et al.
Publicado: (2023)
por: Li, Yi, et al.
Publicado: (2023)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
por: Zhang, Ji, et al.
Publicado: (2025)
por: Zhang, Ji, et al.
Publicado: (2025)
Feature Augmentation for Self-supervised Contrastive Learning: A Closer Look
por: Zhang, Yong, et al.
Publicado: (2024)
por: Zhang, Yong, et al.
Publicado: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
por: Kahatapitiya, Kumara, et al.
Publicado: (2023)
por: Kahatapitiya, Kumara, et al.
Publicado: (2023)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
por: Kim, Jae Myung, et al.
Publicado: (2025)
por: Kim, Jae Myung, et al.
Publicado: (2025)
Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation
por: Li, Lin, et al.
Publicado: (2025)
por: Li, Lin, et al.
Publicado: (2025)
Dense Optical Tracking: Connecting the Dots
por: Moing, Guillaume Le, et al.
Publicado: (2023)
por: Moing, Guillaume Le, et al.
Publicado: (2023)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
por: Chen, Shizhe, et al.
Publicado: (2026)
por: Chen, Shizhe, et al.
Publicado: (2026)
ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning
por: Zhang, David Junhao, et al.
Publicado: (2024)
por: Zhang, David Junhao, et al.
Publicado: (2024)
A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models
por: Silva-Rodríguez, Julio, et al.
Publicado: (2023)
por: Silva-Rodríguez, Julio, et al.
Publicado: (2023)
A Closer Look at Benchmarking Self-Supervised Pre-training with Image Classification
por: Marks, Markus, et al.
Publicado: (2024)
por: Marks, Markus, et al.
Publicado: (2024)
Rectified Decoupled Dataset Distillation: A Closer Look for Fair and Comprehensive Evaluation
por: Zhong, Xinhao, et al.
Publicado: (2025)
por: Zhong, Xinhao, et al.
Publicado: (2025)
LookCloser: Frequency-aware Radiance Field for Tiny-Detail Scene
por: Zhang, Xiaoyu, et al.
Publicado: (2025)
por: Zhang, Xiaoyu, et al.
Publicado: (2025)
Unraveling Instance Associations: A Closer Look for Audio-Visual Segmentation
por: Chen, Yuanhong, et al.
Publicado: (2023)
por: Chen, Yuanhong, et al.
Publicado: (2023)
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
por: Kamenetsky, Ronen, et al.
Publicado: (2025)
por: Kamenetsky, Ronen, et al.
Publicado: (2025)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
por: Taghipour, Ashkan, et al.
Publicado: (2024)
por: Taghipour, Ashkan, et al.
Publicado: (2024)
VidPanos: Generative Panoramic Videos from Casual Panning Videos
por: Ma, Jingwei, et al.
Publicado: (2024)
por: Ma, Jingwei, et al.
Publicado: (2024)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
por: Caron, Mathilde, et al.
Publicado: (2024)
por: Caron, Mathilde, et al.
Publicado: (2024)
Ejemplares similares
-
Dense Video Object Captioning from Disjoint Supervision
por: Zhou, Xingyi, et al.
Publicado: (2023) -
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
por: Wysoczańska, Monika, et al.
Publicado: (2025) -
Time-, Memory- and Parameter-Efficient Visual Adaptation
por: Mercea, Otniel-Bogdan, et al.
Publicado: (2024) -
Grounded Video Caption Generation
por: Kazakos, Evangelos, et al.
Publicado: (2024) -
Streaming Dense Video Captioning
por: Zhou, Xingyi, et al.
Publicado: (2024)