MU-MAE: Multimodal Masked Autoencoders-Based One-Shot Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Rex, Liu, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024)
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
von: Sun, Hao, et al.
Veröffentlicht: (2024)
von: Sun, Hao, et al.
Veröffentlicht: (2024)
Detached and Interactive Multimodal Learning
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
von: Fan, Yunfeng, et al.
Veröffentlicht: (2024)
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion
von: Chen, Sen, et al.
Veröffentlicht: (2022)
von: Chen, Sen, et al.
Veröffentlicht: (2022)
SMC++: Masked Learning of Unsupervised Video Semantic Compression
von: Tian, Yuan, et al.
Veröffentlicht: (2024)
von: Tian, Yuan, et al.
Veröffentlicht: (2024)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
von: Cai, Haonan, et al.
Veröffentlicht: (2026)
von: Cai, Haonan, et al.
Veröffentlicht: (2026)
XY-Cut++: Advanced Layout Ordering via Hierarchical Mask Mechanism on a Novel Benchmark
von: Liu, Shuai, et al.
Veröffentlicht: (2025)
von: Liu, Shuai, et al.
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
von: Le, Anh-Duy, et al.
Veröffentlicht: (2026)
von: Le, Anh-Duy, et al.
Veröffentlicht: (2026)
Scaling and Masking: A New Paradigm of Data Sampling for Image and Video Quality Assessment
von: Liu, Yongxu, et al.
Veröffentlicht: (2024)
von: Liu, Yongxu, et al.
Veröffentlicht: (2024)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
von: He, Liu, et al.
Veröffentlicht: (2024)
von: He, Liu, et al.
Veröffentlicht: (2024)
Generalizable Deepfake Detection Based on Forgery-aware Layer Masking and Multi-artifact Subspace Decomposition
von: Zhang, Xiang, et al.
Veröffentlicht: (2026)
von: Zhang, Xiang, et al.
Veröffentlicht: (2026)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
A Dual-Module Denoising Approach with Curriculum Learning for Enhancing Multimodal Aspect-Based Sentiment Analysis
von: Van Doan, Nguyen, et al.
Veröffentlicht: (2024)
von: Van Doan, Nguyen, et al.
Veröffentlicht: (2024)
OneDiff: A Generalist Model for Image Difference Captioning
von: Hu, Erdong, et al.
Veröffentlicht: (2024)
von: Hu, Erdong, et al.
Veröffentlicht: (2024)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
VCoME: Verbal Video Composition with Multimodal Editing Effects
von: Gong, Weibo, et al.
Veröffentlicht: (2024)
von: Gong, Weibo, et al.
Veröffentlicht: (2024)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
Advancing Unsupervised Low-light Image Enhancement: Noise Estimation, Illumination Interpolation, and Self-Regulation
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Tao, Ziyuan, et al.
Veröffentlicht: (2025)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
von: Wang, Yilin, et al.
Veröffentlicht: (2025)
von: Wang, Yilin, et al.
Veröffentlicht: (2025)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
von: Xu, Zitong, et al.
Veröffentlicht: (2025)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
von: Dong, Ziyi, et al.
Veröffentlicht: (2022)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
von: Huang, Zhijian, et al.
Veröffentlicht: (2024)
von: Huang, Zhijian, et al.
Veröffentlicht: (2024)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
QPT V2: Masked Image Modeling Advances Visual Scoring
von: Xie, Qizhi, et al.
Veröffentlicht: (2024)
von: Xie, Qizhi, et al.
Veröffentlicht: (2024)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
von: Xie, Liping, et al.
Veröffentlicht: (2025)
von: Xie, Liping, et al.
Veröffentlicht: (2025)
M2ORT: Many-To-One Regression Transformer for Spatial Transcriptomics Prediction from Histopathology Images
von: Wang, Hongyi, et al.
Veröffentlicht: (2024)
von: Wang, Hongyi, et al.
Veröffentlicht: (2024)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
von: Taniguchi, Takara, et al.
Veröffentlicht: (2024)
von: Taniguchi, Takara, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024) -
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024) -
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
von: Araujo, Edson, et al.
Veröffentlicht: (2025) -
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
von: Fernandez-Lopez, Adriana, et al.
Veröffentlicht: (2024) -
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
von: Tang, Hao, et al.
Veröffentlicht: (2025)