JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Sousa, Joao, Darabi, Roya, Sousa, Armando, Brueckner, Frank, Reis, Luís Paulo, Reis, Ana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026)
by: Darabi, Nastaran, et al.
Published: (2026)
The Good, the Better, and the Best: Improving the Discriminability of Face Embeddings through Attribute-aware Learning
by: Dias, Ana, et al.
Published: (2026)
by: Dias, Ana, et al.
Published: (2026)
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
by: Jiang, Haonan, et al.
Published: (2026)
by: Jiang, Haonan, et al.
Published: (2026)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Optimizing CLIP Models for Image Retrieval with Maintained Joint-Embedding Alignment
by: Schall, Konstantin, et al.
Published: (2024)
by: Schall, Konstantin, et al.
Published: (2024)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
A Framework for a Capability-driven Evaluation of Scenario Understanding for Multimodal Large Language Models in Autonomous Driving
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
Joint Superpixel and Self-Representation Learning for Scalable Hyperspectral Image Clustering
by: Li, Xianlu, et al.
Published: (2025)
by: Li, Xianlu, et al.
Published: (2025)
Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
JEPA for RL: Investigating Joint-Embedding Predictive Architectures for Reinforcement Learning
by: Kenneweg, Tristan, et al.
Published: (2025)
by: Kenneweg, Tristan, et al.
Published: (2025)
Topo-GS: Continuous Volumetric Embedding of High-Dimensional Data via Topological Gaussian Splatting
by: Gois, João Paulo, et al.
Published: (2026)
by: Gois, João Paulo, et al.
Published: (2026)
Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding Approach
by: Zou, Mian, et al.
Published: (2024)
by: Zou, Mian, et al.
Published: (2024)
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
by: Yang, Wenjie, et al.
Published: (2026)
by: Yang, Wenjie, et al.
Published: (2026)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Multimodal Alignment and Fusion: A Survey
by: Li, Songtao, et al.
Published: (2024)
by: Li, Songtao, et al.
Published: (2024)
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
by: Vo, Khang H. N., et al.
Published: (2025)
by: Vo, Khang H. N., et al.
Published: (2025)
Joint-Embedding Predictive Architecture for Self-Supervised Learning of Mask Classification Architecture
by: Kim, Dong-Hee, et al.
Published: (2024)
by: Kim, Dong-Hee, et al.
Published: (2024)
Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud
by: Saito, Ayumu, et al.
Published: (2024)
by: Saito, Ayumu, et al.
Published: (2024)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
by: Wang, Hongyi, et al.
Published: (2025)
by: Wang, Hongyi, et al.
Published: (2025)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
by: Yu, Shiyao, et al.
Published: (2025)
by: Yu, Shiyao, et al.
Published: (2025)
CoRe: Joint Optimization with Contrastive Learning for Medical Image Registration
by: Kats, Eytan, et al.
Published: (2026)
by: Kats, Eytan, et al.
Published: (2026)
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
by: Du, Yuxuan, et al.
Published: (2025)
by: Du, Yuxuan, et al.
Published: (2025)
Sleep Stage Classification using Multimodal Embedding Fusion from EOG and PSM
by: Papillon, Olivier, et al.
Published: (2025)
by: Papillon, Olivier, et al.
Published: (2025)
Joint Self-Supervised Video Alignment and Action Segmentation
by: Ali, Ali Shah, et al.
Published: (2025)
by: Ali, Ali Shah, et al.
Published: (2025)
Joint Identity Verification and Pose Alignment for Partial Fingerprints
by: Guan, Xiongjun, et al.
Published: (2024)
by: Guan, Xiongjun, et al.
Published: (2024)
TiCoSS: Tightening the Coupling between Semantic Segmentation and Stereo Matching within A Joint Learning Framework
by: Tang, Guanfeng, et al.
Published: (2024)
by: Tang, Guanfeng, et al.
Published: (2024)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
Improving Joint Embedding Predictive Architecture with Diffusion Noise
by: Qiu, Yuping, et al.
Published: (2025)
by: Qiu, Yuping, et al.
Published: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026)
by: Khosla, Savya, et al.
Published: (2026)
Embedding Space Allocation with Angle-Norm Joint Classifiers for Few-Shot Class-Incremental Learning
by: Tu, Dunwei, et al.
Published: (2024)
by: Tu, Dunwei, et al.
Published: (2024)
CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
by: Hossain, Md Aminur, et al.
Published: (2026)
by: Hossain, Md Aminur, et al.
Published: (2026)
The Indra Representation Hypothesis for Multimodal Alignment
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
3D-JEPA: A Joint Embedding Predictive Architecture for 3D Self-Supervised Representation Learning
by: Hu, Naiwen, et al.
Published: (2024)
by: Hu, Naiwen, et al.
Published: (2024)
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
by: Tang, Jianting, et al.
Published: (2025)
by: Tang, Jianting, et al.
Published: (2025)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
by: Qian, Chengxuan, et al.
Published: (2025)
by: Qian, Chengxuan, et al.
Published: (2025)
Similar Items
-
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
by: Darabi, Nastaran, et al.
Published: (2026) -
The Good, the Better, and the Best: Improving the Discriminability of Face Embeddings through Attribute-aware Learning
by: Dias, Ana, et al.
Published: (2026) -
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
by: Jiang, Haonan, et al.
Published: (2026) -
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
by: Liu, Yang, et al.
Published: (2025) -
Optimizing CLIP Models for Image Retrieval with Maintained Joint-Embedding Alignment
by: Schall, Konstantin, et al.
Published: (2024)