Explaining How Visual, Textual and Multimodal Encoders Share Concepts
Fuente:
arXiv
Salvato in:
| Autori principali: | Cornet, Clément, Besançon, Romaric, Borgne, Hervé Le |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automatic Die Studies for Ancient Numismatics
di: Cornet, Clément, et al.
Pubblicazione: (2024)
di: Cornet, Clément, et al.
Pubblicazione: (2024)
CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
di: LIN, Frédéric, et al.
Pubblicazione: (2025)
di: LIN, Frédéric, et al.
Pubblicazione: (2025)
ALPI: Auto-Labeller with Proxy Injection for 3D Object Detection using 2D Labels Only
di: Lahlali, Saad, et al.
Pubblicazione: (2024)
di: Lahlali, Saad, et al.
Pubblicazione: (2024)
The Deleuzian Representation Hypothesis
di: Cornet, Clément, et al.
Pubblicazione: (2025)
di: Cornet, Clément, et al.
Pubblicazione: (2025)
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet
di: Giulivi, Loris, et al.
Pubblicazione: (2024)
di: Giulivi, Loris, et al.
Pubblicazione: (2024)
Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
di: She, Yifei, et al.
Pubblicazione: (2025)
di: She, Yifei, et al.
Pubblicazione: (2025)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
di: Luo, Yang, et al.
Pubblicazione: (2024)
di: Luo, Yang, et al.
Pubblicazione: (2024)
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
di: Ding, Yizhuo, et al.
Pubblicazione: (2025)
How Do Optical Flow and Textual Prompts Collaborate to Assist in Audio-Visual Semantic Segmentation?
di: Lee, Yujian, et al.
Pubblicazione: (2026)
di: Lee, Yujian, et al.
Pubblicazione: (2026)
Token Activation Map to Visually Explain Multimodal LLMs
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation
di: Grimal, Paul, et al.
Pubblicazione: (2023)
di: Grimal, Paul, et al.
Pubblicazione: (2023)
LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks
di: Nguyen, Truong Thanh Hung, et al.
Pubblicazione: (2024)
di: Nguyen, Truong Thanh Hung, et al.
Pubblicazione: (2024)
Improving the Explain-Any-Concept by Introducing Nonlinearity to the Trainable Surrogate Model
di: Zaval, Mounes, et al.
Pubblicazione: (2024)
di: Zaval, Mounes, et al.
Pubblicazione: (2024)
FG-CLIP: Fine-Grained Visual and Textual Alignment
di: Xie, Chunyu, et al.
Pubblicazione: (2025)
di: Xie, Chunyu, et al.
Pubblicazione: (2025)
Enhancing Spatial Reasoning through Visual and Textual Thinking
di: Liang, Xun, et al.
Pubblicazione: (2025)
di: Liang, Xun, et al.
Pubblicazione: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
di: Verma, Gaurav, et al.
Pubblicazione: (2024)
di: Verma, Gaurav, et al.
Pubblicazione: (2024)
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
di: Baniecki, Hubert, et al.
Pubblicazione: (2025)
di: Baniecki, Hubert, et al.
Pubblicazione: (2025)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
di: Chen, Yuan, et al.
Pubblicazione: (2025)
di: Chen, Yuan, et al.
Pubblicazione: (2025)
Patch-wise Auto-Encoder for Visual Anomaly Detection
di: Cui, Yajie, et al.
Pubblicazione: (2023)
di: Cui, Yajie, et al.
Pubblicazione: (2023)
On Explaining Visual Captioning with Hybrid Markov Logic Networks
di: Shah, Monika, et al.
Pubblicazione: (2025)
di: Shah, Monika, et al.
Pubblicazione: (2025)
CoProNN: Concept-based Prototypical Nearest Neighbors for Explaining Vision Models
di: Chiaburu, Teodor, et al.
Pubblicazione: (2024)
di: Chiaburu, Teodor, et al.
Pubblicazione: (2024)
SIEDD: Shared-Implicit Encoder with Discrete Decoders
di: Rangarajan, Vikram, et al.
Pubblicazione: (2025)
di: Rangarajan, Vikram, et al.
Pubblicazione: (2025)
VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment
di: Cheng, Kai, et al.
Pubblicazione: (2026)
di: Cheng, Kai, et al.
Pubblicazione: (2026)
ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet
di: Cheong, Soon Yau, et al.
Pubblicazione: (2023)
di: Cheong, Soon Yau, et al.
Pubblicazione: (2023)
VideoPrism: A Foundational Visual Encoder for Video Understanding
di: Zhao, Long, et al.
Pubblicazione: (2024)
di: Zhao, Long, et al.
Pubblicazione: (2024)
Language Models Can Explain Visual Features via Steering
di: Ferrando, Javier, et al.
Pubblicazione: (2026)
di: Ferrando, Javier, et al.
Pubblicazione: (2026)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
Breaking Language Barriers in Visual Language Models via Multilingual Textual Regularization
di: Pikabea, Iñigo, et al.
Pubblicazione: (2025)
di: Pikabea, Iñigo, et al.
Pubblicazione: (2025)
Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
di: Feng, Zhanbo, et al.
Pubblicazione: (2023)
di: Feng, Zhanbo, et al.
Pubblicazione: (2023)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
di: Wadhawan, Rohan, et al.
Pubblicazione: (2024)
di: Wadhawan, Rohan, et al.
Pubblicazione: (2024)
Deformable Image Registration with Multi-scale Feature Fusion from Shared Encoder, Auxiliary and Pyramid Decoders
di: Zhou, Hongchao, et al.
Pubblicazione: (2024)
di: Zhou, Hongchao, et al.
Pubblicazione: (2024)
ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
di: Choi, Jinho, et al.
Pubblicazione: (2025)
di: Choi, Jinho, et al.
Pubblicazione: (2025)
How to Blend Concepts in Diffusion Models
di: Olearo, Lorenzo, et al.
Pubblicazione: (2024)
di: Olearo, Lorenzo, et al.
Pubblicazione: (2024)
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
di: Gao, Yuan, et al.
Pubblicazione: (2025)
di: Gao, Yuan, et al.
Pubblicazione: (2025)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
di: Cao, Jiajun, et al.
Pubblicazione: (2025)
Explain Before You Answer: A Survey on Compositional Visual Reasoning
di: Ke, Fucai, et al.
Pubblicazione: (2025)
di: Ke, Fucai, et al.
Pubblicazione: (2025)
BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
di: Mei, Guofeng, et al.
Pubblicazione: (2026)
di: Mei, Guofeng, et al.
Pubblicazione: (2026)
Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and VisualAnalysis Strategy
di: Zhang, Hong, et al.
Pubblicazione: (2024)
di: Zhang, Hong, et al.
Pubblicazione: (2024)
Accountable Textual-Visual Chat Learns to Reject Human Instructions in Image Re-creation
di: Zhang, Zhiwei, et al.
Pubblicazione: (2023)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Automatic Die Studies for Ancient Numismatics
di: Cornet, Clément, et al.
Pubblicazione: (2024) -
CaMiT: A Time-Aware Car Model Dataset for Classification and Generation
di: LIN, Frédéric, et al.
Pubblicazione: (2025) -
ALPI: Auto-Labeller with Proxy Injection for 3D Object Detection using 2D Labels Only
di: Lahlali, Saad, et al.
Pubblicazione: (2024) -
The Deleuzian Representation Hypothesis
di: Cornet, Clément, et al.
Pubblicazione: (2025) -
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet
di: Giulivi, Loris, et al.
Pubblicazione: (2024)