Learning Relative Representations for Fine-Grained Multimodal Alignment with Limited Data
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Shiwon, Park, Yu Rang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gramian Multimodal Representation Learning and Alignment
by: Cicchetti, Giordano, et al.
Published: (2024)
by: Cicchetti, Giordano, et al.
Published: (2024)
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
by: Gröger, Fabian, et al.
Published: (2025)
by: Gröger, Fabian, et al.
Published: (2025)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
by: Seo, Ara, et al.
Published: (2025)
by: Seo, Ara, et al.
Published: (2025)
Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
by: Hong, Yunqi, et al.
Published: (2025)
by: Hong, Yunqi, et al.
Published: (2025)
Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
by: Pawlonka, Szymon, et al.
Published: (2025)
by: Pawlonka, Szymon, et al.
Published: (2025)
Using Self-Supervised Auxiliary Tasks to Improve Fine-Grained Facial Representation
by: Pourmirzaei, Mahdi, et al.
Published: (2021)
by: Pourmirzaei, Mahdi, et al.
Published: (2021)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
Boosting Fine-Grained Visual Anomaly Detection with Coarse-Knowledge-Aware Adversarial Learning
by: Fang, Qingqing, et al.
Published: (2024)
by: Fang, Qingqing, et al.
Published: (2024)
Federated Learning with Feedback Alignment
by: Baek, Incheol, et al.
Published: (2025)
by: Baek, Incheol, et al.
Published: (2025)
Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios
by: Fan, Qi, et al.
Published: (2023)
by: Fan, Qi, et al.
Published: (2023)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
Learning from Limited and Imperfect Data
by: Rangwani, Harsh
Published: (2025)
by: Rangwani, Harsh
Published: (2025)
No Alignment Needed for Generation: Learning Linearly Separable Representations in Diffusion Models
by: Yun, Junno, et al.
Published: (2025)
by: Yun, Junno, et al.
Published: (2025)
Biological Plausibility and Representational Alignment of Feedback Alignment in Convolutional Networks
by: Lance, Jake, et al.
Published: (2026)
by: Lance, Jake, et al.
Published: (2026)
Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
by: Chatterjee, Abhiroop, et al.
Published: (2025)
by: Chatterjee, Abhiroop, et al.
Published: (2025)
Enhancing Fine-Grained Visual Recognition in the Low-Data Regime Through Feature Magnitude Regularization
by: Chapman, Avraham, et al.
Published: (2024)
by: Chapman, Avraham, et al.
Published: (2024)
Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision
by: Cao, Yi, et al.
Published: (2023)
by: Cao, Yi, et al.
Published: (2023)
Towards Fine-Grained Video Question Answering
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
Learning Part Knowledge to Facilitate Category Understanding for Fine-Grained Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2025)
by: Wang, Enguang, et al.
Published: (2025)
Self-supervised Transformation Learning for Equivariant Representations
by: Yu, Jaemyung, et al.
Published: (2025)
by: Yu, Jaemyung, et al.
Published: (2025)
FORGE: Fine-grained Multimodal Evaluation for Manufacturing Scenarios
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
SimO Loss: Anchor-Free Contrastive Loss for Fine-Grained Supervised Contrastive Learning
by: Bouhsine, Taha, et al.
Published: (2024)
by: Bouhsine, Taha, et al.
Published: (2024)
On Computational Limits and Provably Efficient Criteria of Visual Autoregressive Models: A Fine-Grained Complexity Analysis
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
HueManity: Probing Fine-Grained Visual Perception in MLLMs
by: Grover, Rynaa, et al.
Published: (2025)
by: Grover, Rynaa, et al.
Published: (2025)
Recent Deep Semi-supervised Learning Approaches and Related Works
by: Kim, Gyeongho
Published: (2021)
by: Kim, Gyeongho
Published: (2021)
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
by: Kang, Hyunju, et al.
Published: (2026)
by: Kang, Hyunju, et al.
Published: (2026)
Learning Equi-angular Representations for Online Continual Learning
by: Seo, Minhyuk, et al.
Published: (2024)
by: Seo, Minhyuk, et al.
Published: (2024)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
by: Brouwer, Eric, et al.
Published: (2024)
by: Brouwer, Eric, et al.
Published: (2024)
Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated Perturbation
by: Xu, Jie, et al.
Published: (2025)
by: Xu, Jie, et al.
Published: (2025)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
Prototype-Based Knowledge Guidance for Fine-Grained Structured Radiology Reporting
by: Pellegrini, Chantal, et al.
Published: (2026)
by: Pellegrini, Chantal, et al.
Published: (2026)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
by: Rastegar, Sarah, et al.
Published: (2024)
by: Rastegar, Sarah, et al.
Published: (2024)
Exploring Perceptual Limitation of Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity
by: Cicchetti, Giordano, et al.
Published: (2025)
by: Cicchetti, Giordano, et al.
Published: (2025)
A Closer Look at Multimodal Representation Collapse
by: Chaudhuri, Abhra, et al.
Published: (2025)
by: Chaudhuri, Abhra, et al.
Published: (2025)
Overcoming Data Inequality across Domains with Semi-Supervised Domain Generalization
by: Park, Jinha, et al.
Published: (2024)
by: Park, Jinha, et al.
Published: (2024)
LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models
by: Yu, Shu, et al.
Published: (2025)
by: Yu, Shu, et al.
Published: (2025)
Test-time Alignment of Diffusion Models without Reward Over-optimization
by: Kim, Sunwoo, et al.
Published: (2025)
by: Kim, Sunwoo, et al.
Published: (2025)
Similar Items
-
Gramian Multimodal Representation Learning and Alignment
by: Cicchetti, Giordano, et al.
Published: (2024) -
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
by: Gröger, Fabian, et al.
Published: (2025) -
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
by: Jiang, Jiachen, et al.
Published: (2025) -
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
by: Seo, Ara, et al.
Published: (2025) -
Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
by: Hong, Yunqi, et al.
Published: (2025)