Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Verma, Gaurav, Choi, Minje, Sharma, Kartik, Watson-Daniels, Jamelle, Oh, Sejoon, Kumar, Srijan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Algorithmic Fairness and Color-blind Racism: Navigating the Intersection
von: Watson-Daniels, Jamelle
Veröffentlicht: (2024)
von: Watson-Daniels, Jamelle
Veröffentlicht: (2024)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
von: Broomfield, Julius, et al.
Veröffentlicht: (2025)
von: Broomfield, Julius, et al.
Veröffentlicht: (2025)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Adversarial Text Rewriting for Text-aware Recommender Systems
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2025)
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2025)
One-Topic-Doesn't-Fit-All: Transcreating Reading Comprehension Test for Personalized Learning
von: Han, Jieun, et al.
Veröffentlicht: (2025)
von: Han, Jieun, et al.
Veröffentlicht: (2025)
Explorations of the Softmax Space: Knowing When the Neural Network Doesn't Know
von: Sikar, Daniel, et al.
Veröffentlicht: (2025)
von: Sikar, Daniel, et al.
Veröffentlicht: (2025)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
von: He, Qianxi, et al.
Veröffentlicht: (2025)
von: He, Qianxi, et al.
Veröffentlicht: (2025)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024)
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lin, Jiacheng, et al.
Veröffentlicht: (2025)
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
von: Li, Victoria R., et al.
Veröffentlicht: (2024)
von: Li, Victoria R., et al.
Veröffentlicht: (2024)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
von: Jain, Shomik, et al.
Veröffentlicht: (2025)
von: Jain, Shomik, et al.
Veröffentlicht: (2025)
Dead Code Doesn't Talk: Authentic Requirements Elicitation in Introductory Software Engineering
von: Berrezueta-Guzman, Santiago, et al.
Veröffentlicht: (2026)
von: Berrezueta-Guzman, Santiago, et al.
Veröffentlicht: (2026)
Characterizing, Detecting, and Predicting Online Ban Evasion
von: Niverthi, Manoj, et al.
Veröffentlicht: (2022)
von: Niverthi, Manoj, et al.
Veröffentlicht: (2022)
Just Because We Camp, Doesn't Mean We Should: The Ethics of Modelling Queer Voices
von: Sigurgeirsson, Atli, et al.
Veröffentlicht: (2024)
von: Sigurgeirsson, Atli, et al.
Veröffentlicht: (2024)
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
von: Taguchi, Chihiro, et al.
Veröffentlicht: (2024)
von: Taguchi, Chihiro, et al.
Veröffentlicht: (2024)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
von: Gong, Shuzhi, et al.
Veröffentlicht: (2026)
von: Gong, Shuzhi, et al.
Veröffentlicht: (2026)
Testing Autonomous Driving Systems -- What Really Matters and What Doesn't
von: Li, Changwen, et al.
Veröffentlicht: (2025)
von: Li, Changwen, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling
von: Byeon, Junhyeong, et al.
Veröffentlicht: (2026)
von: Byeon, Junhyeong, et al.
Veröffentlicht: (2026)
Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
von: Barančíková, Petra, et al.
Veröffentlicht: (2025)
von: Barančíková, Petra, et al.
Veröffentlicht: (2025)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
von: Wang, Ziqiao, et al.
Veröffentlicht: (2025)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
von: Cooper, A. Feder, et al.
Veröffentlicht: (2024)
von: Cooper, A. Feder, et al.
Veröffentlicht: (2024)
One Size Doesn't Fit All: Age-Aware Gamification Mechanics for Multimedia Learning Environments
von: Kaißer, Sarah, et al.
Veröffentlicht: (2025)
von: Kaißer, Sarah, et al.
Veröffentlicht: (2025)
Something There Is That Doesn't Love a Computer (Nor Hate It Either).
von: Friedman, Fred T.
Veröffentlicht: (1984)
von: Friedman, Fred T.
Veröffentlicht: (1984)
Why Doesn't Microsoft Let Me Sleep? How Automaticity of Windows Updates Impacts User Autonomy
von: Ahuja, Sanju, et al.
Veröffentlicht: (2024)
von: Ahuja, Sanju, et al.
Veröffentlicht: (2024)
Camera Height Doesn't Change: Unsupervised Training for Metric Monocular Road-Scene Depth Estimation
von: Kinoshita, Genki, et al.
Veröffentlicht: (2023)
von: Kinoshita, Genki, et al.
Veröffentlicht: (2023)
ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2024)
IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language
von: Chance, Christina, et al.
Veröffentlicht: (2026)
von: Chance, Christina, et al.
Veröffentlicht: (2026)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
von: Sharma, Kartik, et al.
Veröffentlicht: (2025)
von: Sharma, Kartik, et al.
Veröffentlicht: (2025)
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs
von: Sun, Huaman, et al.
Veröffentlicht: (2023)
von: Sun, Huaman, et al.
Veröffentlicht: (2023)
Consciousness Doesn't Do That
von: Matthias Michel
Veröffentlicht: (2026)
von: Matthias Michel
Veröffentlicht: (2026)
The Category Mistake of Cislunar Time: Why NASA Cannot Synchronize What Doesn't Exist
von: Borrill, Paul
Veröffentlicht: (2026)
von: Borrill, Paul
Veröffentlicht: (2026)
CLEAR: Character Unlearning in Textual and Visual Modalities
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
A Framework for Situating Innovations, Opportunities, and Challenges in Advancing Vertical Systems with Large AI Models
von: Verma, Gaurav, et al.
Veröffentlicht: (2025)
von: Verma, Gaurav, et al.
Veröffentlicht: (2025)
Shoestring Digital Library: If Existing Digital Library Software Doesn't Suit Your Needs, Create Your Own
von: Weber, Jonathan
Veröffentlicht: (2006)
von: Weber, Jonathan
Veröffentlicht: (2006)
Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanism
von: Zhang, Haoxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Haoxiang, et al.
Veröffentlicht: (2026)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
Does Alignment Tuning Really Break LLMs' Internal Confidence?
von: Oh, Hongseok, et al.
Veröffentlicht: (2024)
von: Oh, Hongseok, et al.
Veröffentlicht: (2024)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
von: Yu, Zony, et al.
Veröffentlicht: (2025)
von: Yu, Zony, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Algorithmic Fairness and Color-blind Racism: Navigating the Intersection
von: Watson-Daniels, Jamelle
Veröffentlicht: (2024) -
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024) -
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
von: Broomfield, Julius, et al.
Veröffentlicht: (2025) -
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
von: Oh, Sejoon, et al.
Veröffentlicht: (2024) -
Adversarial Text Rewriting for Text-aware Recommender Systems
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)