Audiovisual Masked Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Georgescu, Mariana-Iuliana, Fonseca, Eduardo, Ionescu, Radu Tudor, Lucic, Mario, Schmid, Cordelia, Arnab, Anurag |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
by: Grigore, Diana-Nicoleta, et al.
Published: (2024)
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
by: Dalmonte, Francesco, et al.
Published: (2025)
by: Dalmonte, Francesco, et al.
Published: (2025)
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
by: Zheng, Yuqian, et al.
Published: (2026)
by: Zheng, Yuqian, et al.
Published: (2026)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
CL-MAE: Curriculum-Learned Masked Autoencoders
by: Madan, Neelu, et al.
Published: (2023)
by: Madan, Neelu, et al.
Published: (2023)
CBM: Curriculum by Masking
by: Jarca, Andrei, et al.
Published: (2024)
by: Jarca, Andrei, et al.
Published: (2024)
Towards Few-Call Model Stealing via Active Self-Paced Knowledge Distillation and Diffusion-Based Image Generation
by: Hondru, Vlad, et al.
Published: (2023)
by: Hondru, Vlad, et al.
Published: (2023)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
by: Hong, Joanna, et al.
Published: (2025)
by: Hong, Joanna, et al.
Published: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
by: Mercea, Otniel-Bogdan, et al.
Published: (2024)
Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study
by: Liu, Wuao, et al.
Published: (2026)
by: Liu, Wuao, et al.
Published: (2026)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
Semantic Segmentation in Satellite Hyperspectral Imagery by Deep Learning
by: Justo, Jon Alvarez, et al.
Published: (2023)
by: Justo, Jon Alvarez, et al.
Published: (2023)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
by: Zeng, Donghuo, et al.
Published: (2026)
by: Zeng, Donghuo, et al.
Published: (2026)
Machine Unlearning in the Era of Quantum Machine Learning: An Empirical Study
by: Crivoi, Carla, et al.
Published: (2025)
by: Crivoi, Carla, et al.
Published: (2025)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
by: Vyas, Apoorv, et al.
Published: (2025)
by: Vyas, Apoorv, et al.
Published: (2025)
Speech Audio Generation from dynamic MRI via a Knowledge Enhanced Conditional Variational Autoencoder
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
LaDe: Unified Multi-Layered Graphic Media Generation and Decomposition
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
by: Lungu-Stan, Vlad-Constantin, et al.
Published: (2026)
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
by: Zhang, Xiangyue, et al.
Published: (2025)
by: Zhang, Xiangyue, et al.
Published: (2025)
MTL-MAD: Multi-Task Learners are Effective Medical Anomaly Detectors
by: Bercean, Bogdan Alexandru, et al.
Published: (2026)
by: Bercean, Bogdan Alexandru, et al.
Published: (2026)
Masked Image Modeling: A Survey
by: Hondru, Vlad, et al.
Published: (2024)
by: Hondru, Vlad, et al.
Published: (2024)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
by: Araujo, Edson, et al.
Published: (2025)
by: Araujo, Edson, et al.
Published: (2025)
Subspace-Boosted Model Merging
by: Skorobogat, Ronald, et al.
Published: (2025)
by: Skorobogat, Ronald, et al.
Published: (2025)
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
by: Lee, Seung-jae, et al.
Published: (2025)
by: Lee, Seung-jae, et al.
Published: (2025)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
by: Wu, Yihan, et al.
Published: (2024)
by: Wu, Yihan, et al.
Published: (2024)
Multi-Level Feature Distillation of Joint Teachers Trained on Distinct Image Datasets
by: Iordache, Adrian, et al.
Published: (2024)
by: Iordache, Adrian, et al.
Published: (2024)
Curriculum Multi-Task Self-Supervision Improves Lightweight Architectures for Onboard Satellite Hyperspectral Image Segmentation
by: Carlesso, Hugo, et al.
Published: (2025)
by: Carlesso, Hugo, et al.
Published: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
by: Sun, Licai, et al.
Published: (2024)
by: Sun, Licai, et al.
Published: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
SegMate: Asymmetric Attention-Based Lightweight Architecture for Efficient Multi-Organ Segmentation
by: Bunea, Andrei-Alexandru, et al.
Published: (2026)
by: Bunea, Andrei-Alexandru, et al.
Published: (2026)
Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly Detectors
by: Ristea, Nicolae-Catalin, et al.
Published: (2023)
by: Ristea, Nicolae-Catalin, et al.
Published: (2023)
M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment
by: Nguyen-Phuoc, Long, et al.
Published: (2024)
by: Nguyen-Phuoc, Long, et al.
Published: (2024)
Learning Using Generated Privileged Information by Text-to-Image Diffusion Models
by: Menadil, Rafael-Edy, et al.
Published: (2023)
by: Menadil, Rafael-Edy, et al.
Published: (2023)
VQPP: Video Query Performance Prediction Benchmark
by: Lutu, Adrian Catalin, et al.
Published: (2026)
by: Lutu, Adrian Catalin, et al.
Published: (2026)
PRNU-Bench: A Novel Benchmark and Model for PRNU-Based Camera Identification
by: Croitoru, Florinel Alin, et al.
Published: (2025)
by: Croitoru, Florinel Alin, et al.
Published: (2025)
Similar Items
-
Weight Copy and Low-Rank Adaptation for Few-Shot Distillation of Vision Transformers
by: Grigore, Diana-Nicoleta, et al.
Published: (2024) -
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
by: Dalmonte, Francesco, et al.
Published: (2025) -
X-Aligner: Composed Visual Retrieval without the Bells and Whistles
by: Zheng, Yuqian, et al.
Published: (2026) -
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025) -
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)