Extending Video Masked Autoencoders to 128 frames
Fuente:
arXiv
Guardado en:
| Autores principales: | Gundavarapu, Nitesh Bharadwaj, Friedman, Luke, Goyal, Raghav, Hegde, Chaitra, Agustsson, Eirikur, Waghmare, Sagar M., Sirotenko, Mikhail, Yang, Ming-Hsuan, Weyand, Tobias, Gong, Boqing, Sigal, Leonid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoGLUE: Video General Understanding Evaluation of Foundation Models
por: Yuan, Liangzhe, et al.
Publicado: (2023)
por: Yuan, Liangzhe, et al.
Publicado: (2023)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
por: Nagrani, Arsha, et al.
Publicado: (2024)
por: Nagrani, Arsha, et al.
Publicado: (2024)
VideoPrism: A Foundational Visual Encoder for Video Understanding
por: Zhao, Long, et al.
Publicado: (2024)
por: Zhao, Long, et al.
Publicado: (2024)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
por: Goyal, Raghav, et al.
Publicado: (2023)
por: Goyal, Raghav, et al.
Publicado: (2023)
High-Fidelity Image Compression with Score-based Generative Models
por: Hoogeboom, Emiel, et al.
Publicado: (2023)
por: Hoogeboom, Emiel, et al.
Publicado: (2023)
Factorized Video Autoencoders for Efficient Generative Modelling
por: Suhail, Mohammed, et al.
Publicado: (2024)
por: Suhail, Mohammed, et al.
Publicado: (2024)
Fus-MAE: A cross-attention-based data fusion approach for Masked Autoencoders in remote sensing
por: Chan-To-Hing, Hugo, et al.
Publicado: (2024)
por: Chan-To-Hing, Hugo, et al.
Publicado: (2024)
Higgs Decay Signal Strengths in an Extended 2HDM
por: Bharadwaj, Hrishabh, et al.
Publicado: (2025)
por: Bharadwaj, Hrishabh, et al.
Publicado: (2025)
UGMAE: A Unified Framework for Graph Masked Autoencoders
por: Tian, Yijun, et al.
Publicado: (2024)
por: Tian, Yijun, et al.
Publicado: (2024)
MAX: Masked Autoencoder for X-ray Fluorescence in Geological Investigation
por: Lee, An-Sheng, et al.
Publicado: (2024)
por: Lee, An-Sheng, et al.
Publicado: (2024)
$W-$mass and Muon $g-2$ in Inert 2HDM Extended by Singlet Complex Scalar
por: Bharadwaj, Hrishabh, et al.
Publicado: (2024)
por: Bharadwaj, Hrishabh, et al.
Publicado: (2024)
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset
por: Waghmare, Sagar M., et al.
Publicado: (2023)
por: Waghmare, Sagar M., et al.
Publicado: (2023)
Sobre la realidad de la vida cotidiana de los jóvenes en poblaciones en el nuevo orden democrático: «ni tan protagonista ni tan víctima»
por: Michaela Weyand
Publicado: (1993)
por: Michaela Weyand
Publicado: (1993)
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
por: Li, Lingxiao, et al.
Publicado: (2024)
por: Li, Lingxiao, et al.
Publicado: (2024)
Coherent Spectral Feature Extraction Using Symmetric Autoencoders
por: Bhattacharjee, Archisman, et al.
Publicado: (2023)
por: Bhattacharjee, Archisman, et al.
Publicado: (2023)
Cryptocurrency forensics: Forensic analysis of the Electrum wallet to uncover artifacts
por: Saloni Jain, et al.
Publicado: (2026)
por: Saloni Jain, et al.
Publicado: (2026)
MINERVA: Evaluating Complex Video Reasoning
por: Nagrani, Arsha, et al.
Publicado: (2025)
por: Nagrani, Arsha, et al.
Publicado: (2025)
Covariant Measures of Non-Markovianity in Curved Spacetime
por: Waghmare, Tushar
Publicado: (2025)
por: Waghmare, Tushar
Publicado: (2025)
Designing for Human-Agent Alignment: Understanding what humans want from their agents
por: Goyal, Nitesh, et al.
Publicado: (2024)
por: Goyal, Nitesh, et al.
Publicado: (2024)
A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluations
por: Berman, Glen, et al.
Publicado: (2024)
por: Berman, Glen, et al.
Publicado: (2024)
"You have to prove the threat is real": Understanding the needs of Female Journalists and Activists to Document and Report Online Harassment
por: Goyal, Nitesh, et al.
Publicado: (2022)
por: Goyal, Nitesh, et al.
Publicado: (2022)
Deranged Perfect Matchings on complete graph and balanced complete r-partite graph
por: Deng, Boqing
Publicado: (2025)
por: Deng, Boqing
Publicado: (2025)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
por: Rahman, Tanzila, et al.
Publicado: (2026)
por: Rahman, Tanzila, et al.
Publicado: (2026)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
por: Bhatt, Gaurav, et al.
Publicado: (2024)
por: Bhatt, Gaurav, et al.
Publicado: (2024)
Uniting the World by Dividing it: Federated Maps to Enable Spatial Applications
por: Bharadwaj, Sagar, et al.
Publicado: (2025)
por: Bharadwaj, Sagar, et al.
Publicado: (2025)
Audiovisual Masked Autoencoders
por: Georgescu, Mariana-Iuliana, et al.
Publicado: (2022)
por: Georgescu, Mariana-Iuliana, et al.
Publicado: (2022)
Gaussian Masked Autoencoders
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
Masked Capsule Autoencoders
por: Everett, Miles, et al.
Publicado: (2024)
por: Everett, Miles, et al.
Publicado: (2024)
The interplay of geopolitics and agricultural commodity prices
por: Raghav Goyal, et al.
Publicado: (2024)
por: Raghav Goyal, et al.
Publicado: (2024)
Spectral Determinants of Almost Equilateral Quantum Graphs
por: Harrison, Jonathan, et al.
Publicado: (2024)
por: Harrison, Jonathan, et al.
Publicado: (2024)
Improving Masked Autoencoders by Learning Where to Mask
por: Chen, Haijian, et al.
Publicado: (2023)
por: Chen, Haijian, et al.
Publicado: (2023)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
por: Yu, Lijun, et al.
Publicado: (2023)
por: Yu, Lijun, et al.
Publicado: (2023)
Where to Mask: Structure-Guided Masking for Graph Masked Autoencoders
por: Liu, Chuang, et al.
Publicado: (2024)
por: Liu, Chuang, et al.
Publicado: (2024)
Zero Shot Context-Based Object Segmentation using SLIP (SAM+CLIP)
por: Gundavarapu, Saaketh Koundinya, et al.
Publicado: (2024)
por: Gundavarapu, Saaketh Koundinya, et al.
Publicado: (2024)
Evidence of an inertialess Kapitza instability due to viscosity stratification
por: Gundavarapu, Shravya, et al.
Publicado: (2026)
por: Gundavarapu, Shravya, et al.
Publicado: (2026)
Skip the Hessian, Keep the Rates: Globalized Semismooth Newton with Lazy Hessian Updates
por: Alphonse, Amal, et al.
Publicado: (2026)
por: Alphonse, Amal, et al.
Publicado: (2026)
Dictionary Learning Based Regularization in Quantitative MRI: A Nested Alternating Optimization Framework
por: Dong, Guozhi, et al.
Publicado: (2025)
por: Dong, Guozhi, et al.
Publicado: (2025)
Emergent self-inhibition governs the landscape of stable states in complex ecosystems
por: Patro, Nitesh Kumar, et al.
Publicado: (2025)
por: Patro, Nitesh Kumar, et al.
Publicado: (2025)
Recurrent Video Masked Autoencoders
por: Zoran, Daniel, et al.
Publicado: (2025)
por: Zoran, Daniel, et al.
Publicado: (2025)
Self-Guided Masked Autoencoder
por: Shin, Jeongwoo, et al.
Publicado: (2025)
por: Shin, Jeongwoo, et al.
Publicado: (2025)
Ejemplares similares
-
VideoGLUE: Video General Understanding Evaluation of Foundation Models
por: Yuan, Liangzhe, et al.
Publicado: (2023) -
Neptune: The Long Orbit to Benchmarking Long Video Understanding
por: Nagrani, Arsha, et al.
Publicado: (2024) -
VideoPrism: A Foundational Visual Encoder for Video Understanding
por: Zhao, Long, et al.
Publicado: (2024) -
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
por: Goyal, Raghav, et al.
Publicado: (2023) -
High-Fidelity Image Compression with Score-based Generative Models
por: Hoogeboom, Emiel, et al.
Publicado: (2023)