Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Saporta, Adriel, Puli, Aahlad, Goldstein, Mark, Ranganath, Rajesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation
by: Puli, Aahlad, et al.
Published: (2022)
by: Puli, Aahlad, et al.
Published: (2022)
Explanations that reveal all through the definition of encoding
by: Puli, Aahlad, et al.
Published: (2024)
by: Puli, Aahlad, et al.
Published: (2024)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
by: Srivastava, Archita, et al.
Published: (2025)
by: Srivastava, Archita, et al.
Published: (2025)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis
by: Patro, Badri N., et al.
Published: (2024)
by: Patro, Badri N., et al.
Published: (2024)
VLCD: Vision-Language Contrastive Distillation for Accurate and Efficient Automatic Placenta Analysis
by: Mehta, Manas, et al.
Published: (2025)
by: Mehta, Manas, et al.
Published: (2025)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
by: Huang, Chengyue, et al.
Published: (2025)
by: Huang, Chengyue, et al.
Published: (2025)
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
by: Lan, Zhibin, et al.
Published: (2025)
by: Lan, Zhibin, et al.
Published: (2025)
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
Prompt-Driven Contrastive Learning for Transferable Adversarial Attacks
by: Yang, Hunmin, et al.
Published: (2024)
by: Yang, Hunmin, et al.
Published: (2024)
$\nabla τ$: Gradient-based and Task-Agnostic machine Unlearning
by: Trippa, Daniel, et al.
Published: (2024)
by: Trippa, Daniel, et al.
Published: (2024)
VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
A Cross-Modal Rumor Detection Scheme via Contrastive Learning by Exploring Text and Image internal Correlations
by: Ma, Bin, et al.
Published: (2025)
by: Ma, Bin, et al.
Published: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models
by: Han, Zongbo, et al.
Published: (2024)
by: Han, Zongbo, et al.
Published: (2024)
Contrastive-to-Self-Supervised: A Two-Stage Framework for Script Similarity Learning
by: Roman, Claire, et al.
Published: (2026)
by: Roman, Claire, et al.
Published: (2026)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2026)
by: Villegas, Danae Sánchez, et al.
Published: (2026)
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
by: Zhang, Zilun, et al.
Published: (2024)
by: Zhang, Zilun, et al.
Published: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
by: Liu, Chenxi, et al.
Published: (2025)
by: Liu, Chenxi, et al.
Published: (2025)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
by: Hua, Tianze, et al.
Published: (2025)
by: Hua, Tianze, et al.
Published: (2025)
CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
by: Baumann, Alexander, et al.
Published: (2025)
by: Baumann, Alexander, et al.
Published: (2025)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
by: Lin, Dongyan, et al.
Published: (2026)
by: Lin, Dongyan, et al.
Published: (2026)
Examining Modality Incongruity in Multimodal Federated Learning for Medical Vision and Language-based Disease Detection
by: Saha, Pramit, et al.
Published: (2024)
by: Saha, Pramit, et al.
Published: (2024)
DOCCI: Descriptions of Connected and Contrasting Images
by: Onoe, Yasumasa, et al.
Published: (2024)
by: Onoe, Yasumasa, et al.
Published: (2024)
Efficient Contrastive Decoding with Probabilistic Hallucination Detection - Mitigating Hallucinations in Large Vision Language Models -
by: Fieback, Laura, et al.
Published: (2025)
by: Fieback, Laura, et al.
Published: (2025)
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
by: Wu, Chang-Hsun, et al.
Published: (2025)
by: Wu, Chang-Hsun, et al.
Published: (2025)
Robust Pre-Training of Medical Vision-and-Language Models with Domain-Invariant Multi-Modal Masked Reconstruction
by: Filvantorkaman, Melika, et al.
Published: (2026)
by: Filvantorkaman, Melika, et al.
Published: (2026)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
by: Chakrabarty, Sayak, et al.
Published: (2024)
by: Chakrabarty, Sayak, et al.
Published: (2024)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
by: Zhao, Haozhe, et al.
Published: (2023)
by: Zhao, Haozhe, et al.
Published: (2023)
CURLing the Dream: Contrastive Representations for World Modeling in Reinforcement Learning
by: Kich, Victor Augusto, et al.
Published: (2024)
by: Kich, Victor Augusto, et al.
Published: (2024)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
by: Ye, Jiabo, et al.
Published: (2024)
by: Ye, Jiabo, et al.
Published: (2024)
Adaptive Sampling of k-Space in Magnetic Resonance for Rapid Pathology Prediction
by: Yen, Chen-Yu, et al.
Published: (2024)
by: Yen, Chen-Yu, et al.
Published: (2024)
Similar Items
-
Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation
by: Puli, Aahlad, et al.
Published: (2022) -
Explanations that reveal all through the definition of encoding
by: Puli, Aahlad, et al.
Published: (2024) -
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
by: Srivastava, Archita, et al.
Published: (2025) -
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
by: Wang, Xiyao, et al.
Published: (2024) -
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)