Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Monsefi, Amin Karimi, Zhou, Mengxi, Monsefi, Nastaran Karimi, Lim, Ser-Nam, Chao, Wei-Lun, Ramnath, Rajiv |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
by: Meyarian, Abolfazl, et al.
Published: (2026)
by: Meyarian, Abolfazl, et al.
Published: (2026)
Controlla: Learning Controllability via Graph-Constrained Latent Geometry
by: Murthy, Jamuna S., et al.
Published: (2026)
by: Murthy, Jamuna S., et al.
Published: (2026)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models
by: Navard, Pouyan, et al.
Published: (2024)
by: Navard, Pouyan, et al.
Published: (2024)
Masked LoGoNet: Fast and Accurate 3D Image Analysis for Medical Domain
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
TaxaDiffusion: Progressively Trained Diffusion Model for Fine-Grained Species Generation
by: Monsefi, Amin Karimi, et al.
Published: (2025)
by: Monsefi, Amin Karimi, et al.
Published: (2025)
SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability
by: Monsefi, Amin Karimi, et al.
Published: (2026)
by: Monsefi, Amin Karimi, et al.
Published: (2026)
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life
by: Khurana, Mridul, et al.
Published: (2026)
by: Khurana, Mridul, et al.
Published: (2026)
CrashFormer: A Multimodal Architecture to Predict the Risk of Crash
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
T-MAE: Temporal Masked Autoencoders for Point Cloud Representation Learning
by: Wei, Weijie, et al.
Published: (2023)
by: Wei, Weijie, et al.
Published: (2023)
Towards Chunk-Wise Generation for Long Videos
by: Zhang, Siyang, et al.
Published: (2024)
by: Zhang, Siyang, et al.
Published: (2024)
DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot Segmentation
by: Karimi, Amin, et al.
Published: (2025)
by: Karimi, Amin, et al.
Published: (2025)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision
by: Huang, Jiani, et al.
Published: (2023)
by: Huang, Jiani, et al.
Published: (2023)
Self-Balanced R-CNN for Instance Segmentation
by: Rossi, Leonardo, et al.
Published: (2024)
by: Rossi, Leonardo, et al.
Published: (2024)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Evolved Hierarchical Masking for Self-Supervised Learning
by: Feng, Zhanzhou, et al.
Published: (2025)
by: Feng, Zhanzhou, et al.
Published: (2025)
Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
by: Wang, Ruilang, et al.
Published: (2025)
by: Wang, Ruilang, et al.
Published: (2025)
MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Instruct-ICL: Instruction-Guided In-Context Learning for Post-Disaster Damage Assessment
by: Zarbaft, Armin, et al.
Published: (2026)
by: Zarbaft, Armin, et al.
Published: (2026)
Entropy Reveals Block Importance in Masked Self-Supervised Vision Transformers
by: Xiang, Peihao, et al.
Published: (2026)
by: Xiang, Peihao, et al.
Published: (2026)
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
by: Gao, Yunhe, et al.
Published: (2026)
by: Gao, Yunhe, et al.
Published: (2026)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
by: Huang, Jiyang, et al.
Published: (2026)
by: Huang, Jiyang, et al.
Published: (2026)
Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis
by: Yang, Xinquan, et al.
Published: (2026)
by: Yang, Xinquan, et al.
Published: (2026)
Enhancing Diffusion-based Restoration Models via Difficulty-Adaptive Reinforcement Learning with IQA Reward
by: Xu, Xiaogang, et al.
Published: (2025)
by: Xu, Xiaogang, et al.
Published: (2025)
MARINE: A Computer Vision Model for Detecting Rare Predator-Prey Interactions in Animal Videos
by: Katona, Zsófia, et al.
Published: (2024)
by: Katona, Zsófia, et al.
Published: (2024)
Quantifying Deep Learning Model Uncertainty in Conformal Prediction
by: Karimi, Hamed, et al.
Published: (2023)
by: Karimi, Hamed, et al.
Published: (2023)
Continual Self-Supervised Learning with Masked Autoencoders in Remote Sensing
by: Möllenbrok, Lars, et al.
Published: (2025)
by: Möllenbrok, Lars, et al.
Published: (2025)
FSViewFusion: Few-Shots View Generation of Novel Objects
by: Hussain, Rukhshanda, et al.
Published: (2024)
by: Hussain, Rukhshanda, et al.
Published: (2024)
Similar Items
-
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
by: Meyarian, Abolfazl, et al.
Published: (2026) -
Controlla: Learning Controllability via Graph-Constrained Latent Geometry
by: Murthy, Jamuna S., et al.
Published: (2026) -
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024) -
KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models
by: Navard, Pouyan, et al.
Published: (2024) -
Masked LoGoNet: Fast and Accurate 3D Image Analysis for Medical Domain
by: Monsefi, Amin Karimi, et al.
Published: (2024)