The Indra Representation Hypothesis for Multimodal Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Jianglin, Wang, Hailing, Yang, Kuo, Zhang, Yitian, Jenni, Simon, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Outlier-Aware Post-Training Quantization for Image Super-Resolution
by: Wang, Hailing, et al.
Published: (2025)
by: Wang, Hailing, et al.
Published: (2025)
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
Trajectory Prediction Meets Large Language Models: A Survey
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
by: Lu, Jianglin, et al.
Published: (2025)
by: Lu, Jianglin, et al.
Published: (2025)
Physically Inspired Gaussian Splatting for HDR Novel View Synthesis
by: Zeng, Huimin, et al.
Published: (2026)
by: Zeng, Huimin, et al.
Published: (2026)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning
by: Ma, Xu, et al.
Published: (2026)
by: Ma, Xu, et al.
Published: (2026)
DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning
by: Majeedi, Abrar, et al.
Published: (2026)
by: Majeedi, Abrar, et al.
Published: (2026)
IndraEye: Infrared Electro-Optical UAV-based Perception Dataset for Robust Downstream Tasks
by: D, Manjunath, et al.
Published: (2024)
by: D, Manjunath, et al.
Published: (2024)
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
by: Zhang, Bin, et al.
Published: (2025)
by: Zhang, Bin, et al.
Published: (2025)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
by: Xiao, Junyuan, et al.
Published: (2026)
by: Xiao, Junyuan, et al.
Published: (2026)
Visual Representation Alignment for Multimodal Large Language Models
by: Yoon, Heeji, et al.
Published: (2025)
by: Yoon, Heeji, et al.
Published: (2025)
UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
by: Shi, Kuo, et al.
Published: (2025)
by: Shi, Kuo, et al.
Published: (2025)
Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
by: Lu, Jianglin, et al.
Published: (2025)
by: Lu, Jianglin, et al.
Published: (2025)
Point-SRA: Self-Representation Alignment for 3D Representation Learning
by: Wei, Lintong, et al.
Published: (2026)
by: Wei, Lintong, et al.
Published: (2026)
Dreaming User Multimodal Representation Guided by The Platonic Representation Hypothesis for Micro-Video Recommendation
by: Lin, Chengzhi, et al.
Published: (2024)
by: Lin, Chengzhi, et al.
Published: (2024)
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
by: Huang, Yiyang, et al.
Published: (2025)
by: Huang, Yiyang, et al.
Published: (2025)
Assessment of Multimodal Large Language Models in Alignment with Human Values
by: Shi, Zhelun, et al.
Published: (2024)
by: Shi, Zhelun, et al.
Published: (2024)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
by: Qian, Chengxuan, et al.
Published: (2025)
by: Qian, Chengxuan, et al.
Published: (2025)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
by: Fan, Weichen, et al.
Published: (2025)
by: Fan, Weichen, et al.
Published: (2025)
GmNet: Revisiting Gating Mechanisms From A Frequency View
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
by: Qi, Daiqing, et al.
Published: (2025)
by: Qi, Daiqing, et al.
Published: (2025)
Memory-based Cross-modal Semantic Alignment Network for Radiology Report Generation
by: Tao, Yitian, et al.
Published: (2024)
by: Tao, Yitian, et al.
Published: (2024)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
by: TV, Sethuraman, et al.
Published: (2025)
by: TV, Sethuraman, et al.
Published: (2025)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
by: Yang, Shengzhu, et al.
Published: (2025)
by: Yang, Shengzhu, et al.
Published: (2025)
Representation Forcing for Bottleneck-Free Unified Multimodal Models
by: Wang, Yuqing, et al.
Published: (2026)
by: Wang, Yuqing, et al.
Published: (2026)
HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
by: Wu, Lin, et al.
Published: (2025)
by: Wu, Lin, et al.
Published: (2025)
Accessing Vision Foundation Models via ImageNet-1K
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Dynamic Snake Upsampling Operater and Boundary-Skeleton Weighted Loss for Tubular Structure Segmentation
by: Chen, Yiqi, et al.
Published: (2025)
by: Chen, Yiqi, et al.
Published: (2025)
Improving Taxonomic Image-based Out-of-distribution Detection With DNA Barcodes
by: Impiö, Mikko, et al.
Published: (2024)
by: Impiö, Mikko, et al.
Published: (2024)
Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport
by: Yu, Qinkai, et al.
Published: (2025)
by: Yu, Qinkai, et al.
Published: (2025)
When Does Perceptual Alignment Benefit Vision Representations?
by: Sundaram, Shobhita, et al.
Published: (2024)
by: Sundaram, Shobhita, et al.
Published: (2024)
CosmicMan: A Text-to-Image Foundation Model for Humans
by: Li, Shikai, et al.
Published: (2024)
by: Li, Shikai, et al.
Published: (2024)
Similar Items
-
Outlier-Aware Post-Training Quantization for Image Super-Resolution
by: Wang, Hailing, et al.
Published: (2025) -
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026) -
Trajectory Prediction Meets Large Language Models: A Survey
by: Xu, Yi, et al.
Published: (2025) -
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026) -
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)