Data Collection-free Masked Video Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Ishikawa, Yuchi, Kondo, Masayoshi, Aoki, Yoshimitsu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-training with Synthetic Patterns for Audio
by: Ishikawa, Yuchi, et al.
Published: (2024)
by: Ishikawa, Yuchi, et al.
Published: (2024)
BoundMatch: Boundary detection applied to semi-supervised segmentation
by: Ishikawa, Haruya, et al.
Published: (2025)
by: Ishikawa, Haruya, et al.
Published: (2025)
MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024)
by: Kawano, Yasufumi, et al.
Published: (2024)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
Listening without Looking: Modality Bias in Audio-Visual Captioning
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
TAG: Guidance-free Open-Vocabulary Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024)
by: Kawano, Yasufumi, et al.
Published: (2024)
PCT: Perspective Cue Training Framework for Multi-Camera BEV Segmentation
by: Ishikawa, Haruya, et al.
Published: (2024)
by: Ishikawa, Haruya, et al.
Published: (2024)
Vision-Language Models Learn Super Images for Efficient Partially Relevant Video Retrieval
by: Nishimura, Taichi, et al.
Published: (2023)
by: Nishimura, Taichi, et al.
Published: (2023)
Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering
by: Mori, Erika, et al.
Published: (2025)
by: Mori, Erika, et al.
Published: (2025)
3D Human Scan With A Moving Event Camera
by: Kohyama, Kai, et al.
Published: (2024)
by: Kohyama, Kai, et al.
Published: (2024)
Secrets of Event-Based Optical Flow
by: Shiba, Shintaro, et al.
Published: (2022)
by: Shiba, Shintaro, et al.
Published: (2022)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Event Collapse in Contrast Maximization Frameworks
by: Shiba, Shintaro, et al.
Published: (2022)
by: Shiba, Shintaro, et al.
Published: (2022)
Fast Event-based Optical Flow Estimation by Triplet Matching
by: Shiba, Shintaro, et al.
Published: (2022)
by: Shiba, Shintaro, et al.
Published: (2022)
A Fast Geometric Regularizer to Mitigate Event Collapse in the Contrast Maximization Framework
by: Shiba, Shintaro, et al.
Published: (2022)
by: Shiba, Shintaro, et al.
Published: (2022)
UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
by: Zhang, Zhenghao, et al.
Published: (2023)
by: Zhang, Zhenghao, et al.
Published: (2023)
Simultaneous Motion And Noise Estimation with Event Cameras
by: Shiba, Shintaro, et al.
Published: (2025)
by: Shiba, Shintaro, et al.
Published: (2025)
Geometric-Photometric Event-based 3D Gaussian Ray Tracing
by: Kohyama, Kai, et al.
Published: (2025)
by: Kohyama, Kai, et al.
Published: (2025)
Event-based Background-Oriented Schlieren
by: Shiba, Shintaro, et al.
Published: (2023)
by: Shiba, Shintaro, et al.
Published: (2023)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
by: Ni, Jingcheng, et al.
Published: (2025)
by: Ni, Jingcheng, et al.
Published: (2025)
SIGMA: Sinkhorn-Guided Masked Video Modeling
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Image-based Joint-level Detection for Inflammation in Rheumatoid Arthritis from Small and Imbalanced Data
by: Kato, Shun, et al.
Published: (2026)
by: Kato, Shun, et al.
Published: (2026)
Rethinking Image Super-Resolution from Training Data Perspectives
by: Ohtani, Go, et al.
Published: (2024)
by: Ohtani, Go, et al.
Published: (2024)
Rethinking Video Segmentation with Masked Video Consistency: Did the Model Learn as Intended?
by: Liang, Chen, et al.
Published: (2024)
by: Liang, Chen, et al.
Published: (2024)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding
by: Torimi, Kohei, et al.
Published: (2025)
by: Torimi, Kohei, et al.
Published: (2025)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
by: Zheng, Haoyu, et al.
Published: (2024)
by: Zheng, Haoyu, et al.
Published: (2024)
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
by: Wan, Zhifan, et al.
Published: (2024)
by: Wan, Zhifan, et al.
Published: (2024)
Hierarchical Masked 3D Diffusion Model for Video Outpainting
by: Fan, Fanda, et al.
Published: (2023)
by: Fan, Fanda, et al.
Published: (2023)
MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
by: Tang, Zhicong, et al.
Published: (2026)
by: Tang, Zhicong, et al.
Published: (2026)
Iterative Event-based Motion Segmentation by Variational Contrast Maximization
by: Yamaki, Ryo, et al.
Published: (2025)
by: Yamaki, Ryo, et al.
Published: (2025)
Text-Guided Video Masked Autoencoder
by: Fan, David, et al.
Published: (2024)
by: Fan, David, et al.
Published: (2024)
LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
by: Xiao, Zhihan, et al.
Published: (2025)
by: Xiao, Zhihan, et al.
Published: (2025)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
HR Human: Modeling Human Avatars with Triangular Mesh and High-Resolution Textures from Videos
by: Chen, Qifeng, et al.
Published: (2024)
by: Chen, Qifeng, et al.
Published: (2024)
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
by: Shibata, Yuto, et al.
Published: (2026)
by: Shibata, Yuto, et al.
Published: (2026)
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
by: Ma, Haoyu, et al.
Published: (2023)
by: Ma, Haoyu, et al.
Published: (2023)
VideoMAC: Video Masked Autoencoders Meet ConvNets
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
by: Wang, Zanyi, et al.
Published: (2025)
by: Wang, Zanyi, et al.
Published: (2025)
DVD-Quant: Data-free Video Diffusion Transformers Quantization
by: Li, Zhiteng, et al.
Published: (2025)
by: Li, Zhiteng, et al.
Published: (2025)
Similar Items
-
Pre-training with Synthetic Patterns for Audio
by: Ishikawa, Yuchi, et al.
Published: (2024) -
BoundMatch: Boundary detection applied to semi-supervised segmentation
by: Ishikawa, Haruya, et al.
Published: (2025) -
MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024) -
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
by: Ishikawa, Yuchi, et al.
Published: (2025) -
Listening without Looking: Modality Bias in Audio-Visual Captioning
by: Ishikawa, Yuchi, et al.
Published: (2025)