PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Junmyeong, Hwang, Eui Jun, Cho, Sukmin, Park, Jong C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Efficient Sign Language Translation Using Spatial Configuration and Motion Dynamics with LLMs
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2024)
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2024)
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2024)
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2024)
Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2023)
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2023)
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
von: Yuan, Yike, et al.
Veröffentlicht: (2024)
von: Yuan, Yike, et al.
Veröffentlicht: (2024)
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
von: Zhu, Lei, et al.
Veröffentlicht: (2025)
Towards Latent Masked Image Modeling for Self-Supervised Visual Representation Learning
von: Wei, Yibing, et al.
Veröffentlicht: (2024)
von: Wei, Yibing, et al.
Veröffentlicht: (2024)
Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping
von: Lee, Junmyeong, et al.
Veröffentlicht: (2026)
von: Lee, Junmyeong, et al.
Veröffentlicht: (2026)
PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Yu, et al.
Veröffentlicht: (2024)
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding
von: Zhang, Mingming, et al.
Veröffentlicht: (2023)
von: Zhang, Mingming, et al.
Veröffentlicht: (2023)
VasoMIM: Vascular Anatomy-Aware Masked Image Modeling for Vessel Segmentation
von: Huang, De-Xing, et al.
Veröffentlicht: (2025)
von: Huang, De-Xing, et al.
Veröffentlicht: (2025)
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
von: Lee, Junmyeong, et al.
Veröffentlicht: (2024)
von: Lee, Junmyeong, et al.
Veröffentlicht: (2024)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
von: Zou, Jialv, et al.
Veröffentlicht: (2024)
von: Zou, Jialv, et al.
Veröffentlicht: (2024)
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
von: Lu, Yifan, et al.
Veröffentlicht: (2026)
von: Lu, Yifan, et al.
Veröffentlicht: (2026)
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
HybridMIM: A Hybrid Masked Image Modeling Framework for 3D Medical Image Segmentation
von: Xing, Zhaohu, et al.
Veröffentlicht: (2023)
von: Xing, Zhaohu, et al.
Veröffentlicht: (2023)
Towards Test-time Efficient Visual Place Recognition via Asymmetric Query Processing
von: Kim, Jaeyoon, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyoon, et al.
Veröffentlicht: (2025)
Disentangling Disentangled Representations: Towards Improved Latent Units via Diffusion Models
von: Jun, Youngjun, et al.
Veröffentlicht: (2024)
von: Jun, Youngjun, et al.
Veröffentlicht: (2024)
Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion Models
von: Bradbury, Rowan, et al.
Veröffentlicht: (2025)
von: Bradbury, Rowan, et al.
Veröffentlicht: (2025)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
von: Lee, Youngwan, et al.
Veröffentlicht: (2023)
von: Lee, Youngwan, et al.
Veröffentlicht: (2023)
From Pixels to Components: Eigenvector Masking for Visual Representation Learning
von: Bizeul, Alice, et al.
Veröffentlicht: (2025)
von: Bizeul, Alice, et al.
Veröffentlicht: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration
von: Park, Juhan, et al.
Veröffentlicht: (2025)
von: Park, Juhan, et al.
Veröffentlicht: (2025)
Latent Diffusion Models with Masked AutoEncoders
von: Lee, Junho, et al.
Veröffentlicht: (2025)
von: Lee, Junho, et al.
Veröffentlicht: (2025)
Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network
von: Huang, Pin-Jui, et al.
Veröffentlicht: (2025)
von: Huang, Pin-Jui, et al.
Veröffentlicht: (2025)
Accurate and Fast Pixel Retrieval with Spatial and Uncertainty Aware Hypergraph Diffusion
von: An, Guoyuan, et al.
Veröffentlicht: (2024)
von: An, Guoyuan, et al.
Veröffentlicht: (2024)
Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2025)
von: Xiang, Wenzhao, et al.
Veröffentlicht: (2025)
DMT-JEPA: Discriminative Masked Targets for Joint-Embedding Predictive Architecture
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
RFL-CDNet: Towards Accurate Change Detection via Richer Feature Learning
von: Gan, Yuhang, et al.
Veröffentlicht: (2024)
von: Gan, Yuhang, et al.
Veröffentlicht: (2024)
CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple Images
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
von: Shin, Jisu, et al.
Veröffentlicht: (2024)
LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
von: Huang, Weiquan, et al.
Veröffentlicht: (2024)
von: Huang, Weiquan, et al.
Veröffentlicht: (2024)
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
von: Xie, Chang, et al.
Veröffentlicht: (2025)
von: Xie, Chang, et al.
Veröffentlicht: (2025)
Anomaly Score: Evaluating Generative Models and Individual Generated Images based on Complexity and Vulnerability
von: Hwang, Jaehui, et al.
Veröffentlicht: (2023)
von: Hwang, Jaehui, et al.
Veröffentlicht: (2023)
SC-Pro: Training-Free Framework for Defending Unsafe Image Synthesis Attack
von: Park, Junha, et al.
Veröffentlicht: (2025)
von: Park, Junha, et al.
Veröffentlicht: (2025)
Enhancing Visual Re-ranking through Denoising Nearest Neighbor Graph via Continuous CRF
von: Kim, Jaeyoon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyoon, et al.
Veröffentlicht: (2024)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
von: Sun, Peng, et al.
Veröffentlicht: (2026)
von: Sun, Peng, et al.
Veröffentlicht: (2026)
Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
von: Seon, Juhyeong, et al.
Veröffentlicht: (2024)
von: Seon, Juhyeong, et al.
Veröffentlicht: (2024)
MINR: Implicit Neural Representations with Masked Image Modelling
von: Lee, Sua, et al.
Veröffentlicht: (2025)
von: Lee, Sua, et al.
Veröffentlicht: (2025)
DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
von: Wu, Weijia, et al.
Veröffentlicht: (2023)
von: Wu, Weijia, et al.
Veröffentlicht: (2023)
DroneKey: Drone 3D Pose Estimation in Image Sequences using Gated Key-representation and Pose-adaptive Learning
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2025)
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2025)
DroneKey++: A Size Prior-free Method and New Benchmark for Drone 3D Pose Estimation from Sequential Images
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2026)
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
An Efficient Sign Language Translation Using Spatial Configuration and Motion Dynamics with LLMs
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2024) -
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2024) -
Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2023) -
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
von: Yuan, Yike, et al.
Veröffentlicht: (2024) -
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
von: Zhu, Lei, et al.
Veröffentlicht: (2025)