Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
Fuente:
arXiv
Saved in:
| Main Authors: | Ghazanfari, Sara, Garg, Siddharth, Flammarion, Nicolas, Krishnamurthy, Prashanth, Khorrami, Farshad, Croce, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LipSim: A Provably Robust Perceptual Similarity Metric
by: Ghazanfari, Sara, et al.
Published: (2023)
by: Ghazanfari, Sara, et al.
Published: (2023)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024)
by: Ghazanfari, Sara, et al.
Published: (2024)
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026)
by: Ghazanfari, Sara, et al.
Published: (2026)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)
by: Schlarmann, Christian, et al.
Published: (2025)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
by: Neuhaus, Yannic, et al.
Published: (2026)
by: Neuhaus, Yannic, et al.
Published: (2026)
Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics
by: Croce, Francesco, et al.
Published: (2025)
by: Croce, Francesco, et al.
Published: (2025)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
by: Patel, Naman, et al.
Published: (2025)
by: Patel, Naman, et al.
Published: (2025)
Out-of-Distribution Detection with Overlap Index
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
An Upper Bound for the Distribution Overlap Index and Its Applications
by: Fu, Hao, et al.
Published: (2022)
by: Fu, Hao, et al.
Published: (2022)
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring
by: Fu, Hao, et al.
Published: (2024)
by: Fu, Hao, et al.
Published: (2024)
FlashMix: Fast Map-Free LiDAR Localization via Feature Mixing and Contrastive-Constrained Accelerated Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Efficient and Distributed Large-Scale 3D Map Registration using Tomographic Features
by: Unlu, Halil Utku, et al.
Published: (2024)
by: Unlu, Halil Utku, et al.
Published: (2024)
SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
Towards Reliable Evaluation and Fast Training of Robust Semantic Segmentation Models
by: Croce, Francesco, et al.
Published: (2023)
by: Croce, Francesco, et al.
Published: (2023)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
by: Goswami, Raktim Gautam, et al.
Published: (2024)
by: Goswami, Raktim Gautam, et al.
Published: (2024)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
SpotEdit: Evaluating Visually-Guided Image Editing Methods
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
On the Adversarial Robustness of Discrete Image Tokenizers
by: Bhagwatkar, Rishika, et al.
Published: (2026)
by: Bhagwatkar, Rishika, et al.
Published: (2026)
Generalizable Blood Cell Detection via Unified Dataset and Faster R-CNN
by: Sahay, Siddharth
Published: (2025)
by: Sahay, Siddharth
Published: (2025)
Towards Modality Generalization: A Benchmark and Prospective Analysis
by: Liu, Xiaohao, et al.
Published: (2024)
by: Liu, Xiaohao, et al.
Published: (2024)
Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations
by: Jeong, Minoh, et al.
Published: (2024)
by: Jeong, Minoh, et al.
Published: (2024)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
by: Khorrami, Khazar, et al.
Published: (2021)
by: Khorrami, Khazar, et al.
Published: (2021)
MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
by: Nezakati, Niki, et al.
Published: (2024)
by: Nezakati, Niki, et al.
Published: (2024)
Image Statistics Predict the Sensitivity of Perceptual Quality Metrics
by: Hepburn, Alexander, et al.
Published: (2023)
by: Hepburn, Alexander, et al.
Published: (2023)
Building Age Estimation: A New Multi-Modal Benchmark Dataset and Community Challenge
by: Dionelis, Nikolaos, et al.
Published: (2025)
by: Dionelis, Nikolaos, et al.
Published: (2025)
On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem
by: Pichler, Georg, et al.
Published: (2024)
by: Pichler, Georg, et al.
Published: (2024)
Rethinking Perceptual Metrics for Medical Image Translation
by: Konz, Nicholas, et al.
Published: (2024)
by: Konz, Nicholas, et al.
Published: (2024)
DAAL: Density-Aware Adaptive Line Margin Loss for Multi-Modal Deep Metric Learning
by: Gebrerufael, Hadush Hailu, et al.
Published: (2024)
by: Gebrerufael, Hadush Hailu, et al.
Published: (2024)
Modality Unified Attack for Omni-Modality Person Re-Identification
by: Bian, Yuan, et al.
Published: (2025)
by: Bian, Yuan, et al.
Published: (2025)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
by: Schlarmann, Christian, et al.
Published: (2024)
by: Schlarmann, Christian, et al.
Published: (2024)
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP
by: Singh, Naman Deep, et al.
Published: (2024)
by: Singh, Naman Deep, et al.
Published: (2024)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Enhancing Cardiovascular Disease Prediction through Multi-Modal Self-Supervised Learning
by: Girlanda, Francesco, et al.
Published: (2024)
by: Girlanda, Francesco, et al.
Published: (2024)
MGHF: Multi-Granular High-Frequency Perceptual Loss for Image Super-Resolution
by: Sami, Shoaib Meraj, et al.
Published: (2024)
by: Sami, Shoaib Meraj, et al.
Published: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Curia: A Multi-Modal Foundation Model for Radiology
by: Dancette, Corentin, et al.
Published: (2025)
by: Dancette, Corentin, et al.
Published: (2025)
Diffusion Model with Perceptual Loss
by: Lin, Shanchuan, et al.
Published: (2023)
by: Lin, Shanchuan, et al.
Published: (2023)
A Novel Metric for Detecting Memorization in Generative Models for Brain MRI Synthesis
by: Scardace, Antonio, et al.
Published: (2025)
by: Scardace, Antonio, et al.
Published: (2025)
Similar Items
-
LipSim: A Provably Robust Perceptual Similarity Metric
by: Ghazanfari, Sara, et al.
Published: (2023) -
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
by: Ghazanfari, Sara, et al.
Published: (2025) -
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
by: Ghazanfari, Sara, et al.
Published: (2024) -
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026) -
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
by: Schlarmann, Christian, et al.
Published: (2025)