Semantic Leakage from Image Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Yiyi, Xu, Qiongkai, Elliott, Desmond, Li, Qiongxiu, Bjerva, Johannes |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Defending against Backdoor Attacks via Module Switching
por: Li, Weijun, et al.
Publicado: (2025)
por: Li, Weijun, et al.
Publicado: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
por: He, Wei
Publicado: (2026)
por: He, Wei
Publicado: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
por: Ji, Yikun, et al.
Publicado: (2025)
por: Ji, Yikun, et al.
Publicado: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
por: Freitas, Diogo, et al.
Publicado: (2025)
por: Freitas, Diogo, et al.
Publicado: (2025)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
por: Shah, Nisarg A., et al.
Publicado: (2025)
por: Shah, Nisarg A., et al.
Publicado: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
por: Bucher, Martin JJ., et al.
Publicado: (2025)
por: Bucher, Martin JJ., et al.
Publicado: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
por: Cui, Shaoyang, et al.
Publicado: (2026)
por: Cui, Shaoyang, et al.
Publicado: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
por: Anh, Duy Le Dinh, et al.
Publicado: (2024)
por: Anh, Duy Le Dinh, et al.
Publicado: (2024)
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
por: Dehghani, Mahshid, et al.
Publicado: (2024)
por: Dehghani, Mahshid, et al.
Publicado: (2024)
Seeing the Forest through the Trees: Data Leakage from Partial Transformer Gradients
por: Li, Weijun, et al.
Publicado: (2024)
por: Li, Weijun, et al.
Publicado: (2024)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
por: Cao, Jingtao, et al.
Publicado: (2024)
por: Cao, Jingtao, et al.
Publicado: (2024)
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
por: Niu, Yuwei, et al.
Publicado: (2025)
por: Niu, Yuwei, et al.
Publicado: (2025)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
por: Koukounas, Andreas, et al.
Publicado: (2024)
por: Koukounas, Andreas, et al.
Publicado: (2024)
On the Limitations of Vision-Language Models in Understanding Image Transforms
por: Anis, Ahmad Mustafa, et al.
Publicado: (2025)
por: Anis, Ahmad Mustafa, et al.
Publicado: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
por: Chen, Yuangong, et al.
Publicado: (2026)
por: Chen, Yuangong, et al.
Publicado: (2026)
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
por: Parcalabescu, Letitia, et al.
Publicado: (2021)
por: Parcalabescu, Letitia, et al.
Publicado: (2021)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
por: Parcalabescu, Letitia, et al.
Publicado: (2022)
por: Parcalabescu, Letitia, et al.
Publicado: (2022)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
por: Sanders, Kate, et al.
Publicado: (2024)
por: Sanders, Kate, et al.
Publicado: (2024)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
por: Deichler, Anna, et al.
Publicado: (2026)
por: Deichler, Anna, et al.
Publicado: (2026)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
RONA: Pragmatically Diverse Image Captioning with Coherence Relations
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
por: Ji, Yikun, et al.
Publicado: (2025)
por: Ji, Yikun, et al.
Publicado: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
por: Oliveira, Daniel A. P., et al.
Publicado: (2024)
por: Oliveira, Daniel A. P., et al.
Publicado: (2024)
HuMoCon: Concept Discovery for Human Motion Understanding
por: Fang, Qihang, et al.
Publicado: (2025)
por: Fang, Qihang, et al.
Publicado: (2025)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
por: Khurdula, Harsha Vardhan, et al.
Publicado: (2024)
por: Khurdula, Harsha Vardhan, et al.
Publicado: (2024)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
por: Oliveira, Daniel, et al.
Publicado: (2026)
por: Oliveira, Daniel, et al.
Publicado: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
por: Masrourisaadat, Nila, et al.
Publicado: (2024)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
por: Teo, Charlton
Publicado: (2025)
por: Teo, Charlton
Publicado: (2025)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2025)
Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models
por: Kedia, Akhil, et al.
Publicado: (2024)
por: Kedia, Akhil, et al.
Publicado: (2024)
CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model
por: Yeh, Wei-Hsin, et al.
Publicado: (2025)
por: Yeh, Wei-Hsin, et al.
Publicado: (2025)
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
por: Deichler, Anna, et al.
Publicado: (2025)
por: Deichler, Anna, et al.
Publicado: (2025)
ALGEN: Few-shot Inversion Attacks on Textual Embeddings using Alignment and Generation
por: Chen, Yiyi, et al.
Publicado: (2025)
por: Chen, Yiyi, et al.
Publicado: (2025)
Ejemplares similares
-
Defending against Backdoor Attacks via Module Switching
por: Li, Weijun, et al.
Publicado: (2025) -
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
por: He, Wei
Publicado: (2026) -
GroundCap: A Visually Grounded Image Captioning Dataset
por: Oliveira, Daniel A. P., et al.
Publicado: (2025) -
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
por: Ji, Yikun, et al.
Publicado: (2025) -
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
por: Freitas, Diogo, et al.
Publicado: (2025)