Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Dong-Hee, Cho, Sungduk, Cho, Hyeonwoo, Park, Chanmin, Kim, Jinyoung, Kim, Won Hwa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CNG-SFDA:Clean-and-Noisy Region Guided Online-Offline Source-Free Domain Adaptation
by: Cho, Hyeonwoo, et al.
Published: (2024)
by: Cho, Hyeonwoo, et al.
Published: (2024)
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
by: Kim, Hyeonwoo, et al.
Published: (2026)
by: Kim, Hyeonwoo, et al.
Published: (2026)
Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
by: Zhu, Haoran, et al.
Published: (2025)
by: Zhu, Haoran, et al.
Published: (2025)
Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud
by: Saito, Ayumu, et al.
Published: (2024)
by: Saito, Ayumu, et al.
Published: (2024)
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)
by: Cho, Hyeonwoo, et al.
Published: (2026)
Mine-JEPA: In-Domain Self-Supervised Learning for Mine-Like Object Classification in Side-Scan Sonar
by: Kwon, Taeyoun, et al.
Published: (2026)
by: Kwon, Taeyoun, et al.
Published: (2026)
3D-JEPA: A Joint Embedding Predictive Architecture for 3D Self-Supervised Representation Learning
by: Hu, Naiwen, et al.
Published: (2024)
by: Hu, Naiwen, et al.
Published: (2024)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
by: Kim, Inho, et al.
Published: (2025)
by: Kim, Inho, et al.
Published: (2025)
PBVS 2024 Solution: Self-Supervised Learning and Sampling Strategies for SAR Classification in Extreme Long-Tail Distribution
by: Kim, Yuhyun, et al.
Published: (2024)
by: Kim, Yuhyun, et al.
Published: (2024)
CNN-JEPA: Self-Supervised Pretraining Convolutional Neural Networks Using Joint Embedding Predictive Architecture
by: Kalapos, András, et al.
Published: (2024)
by: Kalapos, András, et al.
Published: (2024)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
S-JEA: Stacked Joint Embedding Architectures for Self-Supervised Visual Representation Learning
by: Manová, Alžběta, et al.
Published: (2023)
by: Manová, Alžběta, et al.
Published: (2023)
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection
by: Kim, Jongha, et al.
Published: (2024)
by: Kim, Jongha, et al.
Published: (2024)
Predicting Gradient is Better: Exploring Self-Supervised Learning for SAR ATR with a Joint-Embedding Predictive Architecture
by: Li, Weijie, et al.
Published: (2023)
by: Li, Weijie, et al.
Published: (2023)
Weakly Supervised Video Scene Graph Generation via Natural Language Supervision
by: Kim, Kibum, et al.
Published: (2025)
by: Kim, Kibum, et al.
Published: (2025)
JEPA-T: Joint-Embedding Predictive Architecture with Text Fusion for Image Generation
by: Wan, Siheng, et al.
Published: (2025)
by: Wan, Siheng, et al.
Published: (2025)
JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning
by: Kenneweg, Tristan, et al.
Published: (2025)
by: Kenneweg, Tristan, et al.
Published: (2025)
Improving Joint Embedding Predictive Architecture with Diffusion Noise
by: Qiu, Yuping, et al.
Published: (2025)
by: Qiu, Yuping, et al.
Published: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
by: Kim, Jisoo, et al.
Published: (2024)
by: Kim, Jisoo, et al.
Published: (2024)
Prompt Learning via Meta-Regularization
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
by: He, Xiangteng, et al.
Published: (2025)
by: He, Xiangteng, et al.
Published: (2025)
Self-Rectifying Diffusion Sampling with Perturbed-Attention Guidance
by: Ahn, Donghoon, et al.
Published: (2024)
by: Ahn, Donghoon, et al.
Published: (2024)
MF-LPR$^2$: Multi-Frame License Plate Image Restoration and Recognition using Optical Flow
by: Na, Kihyun, et al.
Published: (2025)
by: Na, Kihyun, et al.
Published: (2025)
Denoising with a Joint-Embedding Predictive Architecture
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models
by: Baik, Sangwon, et al.
Published: (2025)
by: Baik, Sangwon, et al.
Published: (2025)
DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Zero-Shot Scene Change Detection
by: Cho, Kyusik, et al.
Published: (2024)
by: Cho, Kyusik, et al.
Published: (2024)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024)
by: Park, Chunghyun, et al.
Published: (2024)
Overcoming Data Inequality across Domains with Semi-Supervised Domain Generalization
by: Park, Jinha, et al.
Published: (2024)
by: Park, Jinha, et al.
Published: (2024)
Masked Spatial Propagation Network for Sparsity-Adaptive Depth Refinement
by: Jun, Jinyoung, et al.
Published: (2024)
by: Jun, Jinyoung, et al.
Published: (2024)
RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
by: Park, Jihwan, et al.
Published: (2026)
by: Park, Jihwan, et al.
Published: (2026)
Diffusion Model Compression for Image-to-Image Translation
by: Kim, Geonung, et al.
Published: (2024)
by: Kim, Geonung, et al.
Published: (2024)
TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoising
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
Long-tailed Adversarial Training with Self-Distillation
by: Cho, Seungju, et al.
Published: (2025)
by: Cho, Seungju, et al.
Published: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
Towards Efficient Vision State Space Models via Token Merging
by: Park, Jinyoung, et al.
Published: (2025)
by: Park, Jinyoung, et al.
Published: (2025)
Similar Items
-
CNG-SFDA:Clean-and-Noisy Region Guided Online-Offline Source-Free Domain Adaptation
by: Cho, Hyeonwoo, et al.
Published: (2024) -
DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
by: Kim, Hyeonwoo, et al.
Published: (2026) -
Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object Detection
by: Zhu, Haoran, et al.
Published: (2025) -
Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud
by: Saito, Ayumu, et al.
Published: (2024) -
Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference
by: Cho, Hyeonwoo, et al.
Published: (2026)