MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Du, Zhicheng, Shi, Qingyang, Lu, Jiasheng, Liang, Yingshan, Zhang, Xinyu, Wang, Yiran, Qin, Peiwu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
by: Liang, Yingshan, et al.
Published: (2025)
by: Liang, Yingshan, et al.
Published: (2025)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024)
by: Du, Zhicheng, et al.
Published: (2024)
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
by: Roy, Subhadeep, et al.
Published: (2026)
by: Roy, Subhadeep, et al.
Published: (2026)
Joint Imaging-ROI Representation Learning via Cross-View Contrastive Alignment for Brain Disorder Classification
by: Liang, Wei, et al.
Published: (2026)
by: Liang, Wei, et al.
Published: (2026)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
by: Lu, Liming, et al.
Published: (2025)
by: Lu, Liming, et al.
Published: (2025)
MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
by: Shi, Kuo, et al.
Published: (2025)
by: Shi, Kuo, et al.
Published: (2025)
Conv-INR: Convolutional Implicit Neural Representation for Multimodal Visual Signals
by: Cai, Zhicheng
Published: (2024)
by: Cai, Zhicheng
Published: (2024)
Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
by: Wang, Xinkun, et al.
Published: (2025)
by: Wang, Xinkun, et al.
Published: (2025)
Robust Multimodal Learning via Representation Decoupling
by: Wei, Shicai, et al.
Published: (2024)
by: Wei, Shicai, et al.
Published: (2024)
OmniCam: Unified Multimodal Video Generation via Camera Control
by: Yang, Xiaoda, et al.
Published: (2025)
by: Yang, Xiaoda, et al.
Published: (2025)
SJTU:Spatial judgments in multimodal models towards unified segmentation through coordinate detection
by: Chae, Joongwon, et al.
Published: (2024)
by: Chae, Joongwon, et al.
Published: (2024)
Implicit Neural Representations for Robust Joint Sparse-View CT Reconstruction
by: Shi, Jiayang, et al.
Published: (2024)
by: Shi, Jiayang, et al.
Published: (2024)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
by: Jung, Woojun, et al.
Published: (2026)
by: Jung, Woojun, et al.
Published: (2026)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
by: Liu, Zichuan, et al.
Published: (2025)
by: Liu, Zichuan, et al.
Published: (2025)
Multimodal Prompt Alignment for Facial Expression Recognition
by: Ma, Fuyan, et al.
Published: (2025)
by: Ma, Fuyan, et al.
Published: (2025)
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
by: Li, Yingshu, et al.
Published: (2023)
by: Li, Yingshu, et al.
Published: (2023)
Evaluating Facial Expression Recognition Datasets for Deep Learning: A Benchmark Study with Novel Similarity Metrics
by: Gaya-Morey, F. Xavier, et al.
Published: (2025)
by: Gaya-Morey, F. Xavier, et al.
Published: (2025)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization
by: Liang, Yue, et al.
Published: (2026)
by: Liang, Yue, et al.
Published: (2026)
EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
by: Qu, Qiang, et al.
Published: (2024)
by: Qu, Qiang, et al.
Published: (2024)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
by: Cho, CH, et al.
Published: (2025)
by: Cho, CH, et al.
Published: (2025)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Novel Memory Forgetting Techniques for Autonomous AI Agents: Balancing Relevance and Efficiency
by: Fofadiya, Payal, et al.
Published: (2026)
by: Fofadiya, Payal, et al.
Published: (2026)
Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
by: Gupta, Kishor Datta, et al.
Published: (2025)
by: Gupta, Kishor Datta, et al.
Published: (2025)
DIFFUMA: High-Fidelity Spatio-Temporal Video Prediction via Dual-Path Mamba and Diffusion Enhancement
by: Xie, Xinyu, et al.
Published: (2025)
by: Xie, Xinyu, et al.
Published: (2025)
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
by: Zhang, Junzhe, et al.
Published: (2025)
by: Zhang, Junzhe, et al.
Published: (2025)
A Novel Adaptive Fine-Tuning Algorithm for Multimodal Models: Self-Optimizing Classification and Selection of High-Quality Datasets in Remote Sensing
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
by: He, Yiran, et al.
Published: (2025)
by: He, Yiran, et al.
Published: (2025)
JointRF: End-to-End Joint Optimization for Dynamic Neural Radiance Field Representation and Compression
by: Zheng, Zihan, et al.
Published: (2024)
by: Zheng, Zihan, et al.
Published: (2024)
Video Representation Learning with Joint-Embedding Predictive Architectures
by: Drozdov, Katrina, et al.
Published: (2024)
by: Drozdov, Katrina, et al.
Published: (2024)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
by: Jiang, Yibo, et al.
Published: (2026)
by: Jiang, Yibo, et al.
Published: (2026)
Exposing the Copycat Problem of Imitation-based Planner: A Novel Closed-Loop Simulator, Causal Benchmark and Joint IL-RL Baseline
by: Zhou, Hui, et al.
Published: (2025)
by: Zhou, Hui, et al.
Published: (2025)
Continual Multimodal Egocentric Activity Recognition via Modality-Aware Novel Detection
by: Lim, Wonseon, et al.
Published: (2026)
by: Lim, Wonseon, et al.
Published: (2026)
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
by: Li, Yuxin, et al.
Published: (2024)
by: Li, Yuxin, et al.
Published: (2024)
RU-AI: A Large Multimodal Dataset for Machine-Generated Content Detection
by: Huang, Liting, et al.
Published: (2024)
by: Huang, Liting, et al.
Published: (2024)
IDRL: An Individual-Aware Multimodal Depression-Related Representation Learning Framework for Depression Diagnosis
by: Wang, Chongxiao, et al.
Published: (2026)
by: Wang, Chongxiao, et al.
Published: (2026)
Learning Noise-Robust Joint Representation for Multimodal Emotion Recognition under Incomplete Data Scenarios
by: Fan, Qi, et al.
Published: (2023)
by: Fan, Qi, et al.
Published: (2023)
Similar Items
-
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
by: Liang, Yingshan, et al.
Published: (2025) -
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024) -
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics
by: Roy, Subhadeep, et al.
Published: (2026) -
Joint Imaging-ROI Representation Learning via Cross-View Contrastive Alignment for Brain Disorder Classification
by: Liang, Wei, et al.
Published: (2026) -
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
by: Lu, Liming, et al.
Published: (2025)