Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Jeong, Minoh, Kim, Zae Myung, Namgung, Min, Kang, Dongyeop, Chiang, Yao-Yi, Hero, Alfred |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WalkCLIP: Multimodal Learning for Urban Walkability Prediction
by: Xiang, Shilong, et al.
Published: (2025)
by: Xiang, Shilong, et al.
Published: (2025)
Generalizing Supervised Contrastive learning: A Projection Perspective
by: Jeong, Minoh, et al.
Published: (2025)
by: Jeong, Minoh, et al.
Published: (2025)
Probabilistic Variational Contrastive Learning
by: Jeong, Minoh, et al.
Published: (2025)
by: Jeong, Minoh, et al.
Published: (2025)
LIGHT: Multi-Modal Text Linking on Historical Maps
by: Lin, Yijun, et al.
Published: (2025)
by: Lin, Yijun, et al.
Published: (2025)
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
Event-based Facial Keypoint Alignment via Cross-Modal Fusion Attention and Self-Supervised Multi-Event Representation Learning
by: Kang, Donghwa, et al.
Published: (2025)
by: Kang, Donghwa, et al.
Published: (2025)
Dual-Modality Anchor-Guided Filtering for Test-time Prompt Tuning
by: Choi, Jungwon, et al.
Published: (2026)
by: Choi, Jungwon, et al.
Published: (2026)
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
by: Shao, YiKang, et al.
Published: (2025)
by: Shao, YiKang, et al.
Published: (2025)
UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings
by: Qin, Jiajun, et al.
Published: (2025)
by: Qin, Jiajun, et al.
Published: (2025)
Feature Attenuation of Defective Representation Can Resolve Incomplete Masking on Anomaly Detection
by: Park, YeongHyeon, et al.
Published: (2024)
by: Park, YeongHyeon, et al.
Published: (2024)
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
by: Kim, Hyeonyu, et al.
Published: (2025)
by: Kim, Hyeonyu, et al.
Published: (2025)
Challenging Forgets: Unveiling the Worst-Case Forget Sets in Machine Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification
by: Wan, Zishuo, et al.
Published: (2025)
by: Wan, Zishuo, et al.
Published: (2025)
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023)
by: Chen, Xin, et al.
Published: (2023)
Modality Unified Attack for Omni-Modality Person Re-Identification
by: Bian, Yuan, et al.
Published: (2025)
by: Bian, Yuan, et al.
Published: (2025)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
by: Lee, Yongjin, et al.
Published: (2024)
by: Lee, Yongjin, et al.
Published: (2024)
Hyper-Local Deformable Transformers for Text Spotting on Historical Maps
by: Lin, Yijun, et al.
Published: (2025)
by: Lin, Yijun, et al.
Published: (2025)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Unified Open-World Segmentation with Multi-Modal Prompts
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
by: Xu, Jinjin, et al.
Published: (2023)
by: Xu, Jinjin, et al.
Published: (2023)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
MiTREE: Multi-input Transformer Ecoregion Encoder for Species Distribution Modelling
by: Chen, Theresa, et al.
Published: (2024)
by: Chen, Theresa, et al.
Published: (2024)
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
by: Ahn, Daechul, et al.
Published: (2024)
by: Ahn, Daechul, et al.
Published: (2024)
Federated Modality-specific Encoders and Multimodal Anchors for Personalized Brain Tumor Segmentation
by: Dai, Qian, et al.
Published: (2024)
by: Dai, Qian, et al.
Published: (2024)
RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
by: Hu, Huiyang, et al.
Published: (2025)
by: Hu, Huiyang, et al.
Published: (2025)
Beyond Prototypes: Semantic Anchor Regularization for Better Representation Learning
by: Ge, Yanqi, et al.
Published: (2023)
by: Ge, Yanqi, et al.
Published: (2023)
LDTR: Transformer-based Lane Detection with Anchor-chain Representation
by: Yang, Zhongyu, et al.
Published: (2024)
by: Yang, Zhongyu, et al.
Published: (2024)
BoIR: Box-Supervised Instance Representation for Multi-Person Pose Estimation
by: Jeong, Uyoung, et al.
Published: (2023)
by: Jeong, Uyoung, et al.
Published: (2023)
CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation
by: Berian, Alex, et al.
Published: (2025)
by: Berian, Alex, et al.
Published: (2025)
Large Motion Model for Unified Multi-Modal Motion Generation
by: Zhang, Mingyuan, et al.
Published: (2024)
by: Zhang, Mingyuan, et al.
Published: (2024)
UNIV: Unified Foundation Model for Infrared and Visible Modalities
by: Mao, Fangyuan, et al.
Published: (2025)
by: Mao, Fangyuan, et al.
Published: (2025)
Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
by: Shi, Youxu, et al.
Published: (2025)
by: Shi, Youxu, et al.
Published: (2025)
DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR
by: Cho, Hoonhee, et al.
Published: (2026)
by: Cho, Hoonhee, et al.
Published: (2026)
Diffusion Models For Multi-Modal Generative Modeling
by: Chen, Changyou, et al.
Published: (2024)
by: Chen, Changyou, et al.
Published: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration
by: Liao, Chih-Ting, et al.
Published: (2025)
by: Liao, Chih-Ting, et al.
Published: (2025)
Modality-Agnostic Structural Image Representation Learning for Deformable Multi-Modality Medical Image Registration
by: Mok, Tony C. W., et al.
Published: (2024)
by: Mok, Tony C. W., et al.
Published: (2024)
Similar Items
-
WalkCLIP: Multimodal Learning for Urban Walkability Prediction
by: Xiang, Shilong, et al.
Published: (2025) -
Generalizing Supervised Contrastive learning: A Projection Perspective
by: Jeong, Minoh, et al.
Published: (2025) -
Probabilistic Variational Contrastive Learning
by: Jeong, Minoh, et al.
Published: (2025) -
LIGHT: Multi-Modal Text Linking on Historical Maps
by: Lin, Yijun, et al.
Published: (2025) -
II-MMR: Identifying and Improving Multi-modal Multi-hop Reasoning in Visual Question Answering
by: Kil, Jihyung, et al.
Published: (2024)