Using Images to Find Context-Independent Word Representations in Vector Space
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Kumar, Harsh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
von: Broomfield, Julius, et al.
Veröffentlicht: (2025)
von: Broomfield, Julius, et al.
Veröffentlicht: (2025)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2023)
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2023)
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
Counting Without Numbers and Finding Without Words
von: Patro, Badri Narayana
Veröffentlicht: (2026)
von: Patro, Badri Narayana
Veröffentlicht: (2026)
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
von: He, Zoe Wanying, et al.
Veröffentlicht: (2025)
von: He, Zoe Wanying, et al.
Veröffentlicht: (2025)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
Concept Lancet: Image Editing with Compositional Representation Transplant
von: Luo, Jinqi, et al.
Veröffentlicht: (2025)
von: Luo, Jinqi, et al.
Veröffentlicht: (2025)
Beyond Words: Multimodal LLM Knows When to Speak
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
von: Huang, Brandon, et al.
Veröffentlicht: (2024)
Designing Practical Models for Isolated Word Visual Speech Recognition
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
von: Wang, Junling, et al.
Veröffentlicht: (2026)
von: Wang, Junling, et al.
Veröffentlicht: (2026)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
von: Sharma, Aditya, et al.
Veröffentlicht: (2024)
von: Sharma, Aditya, et al.
Veröffentlicht: (2024)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
von: Zhang, Juntian, et al.
Veröffentlicht: (2025)
von: Zhang, Juntian, et al.
Veröffentlicht: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
von: Hashemi, Mohammad Abuzar, et al.
Veröffentlicht: (2021)
von: Hashemi, Mohammad Abuzar, et al.
Veröffentlicht: (2021)
Visually Descriptive Language Model for Vector Graphics Reasoning
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2024)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2024)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Long-Context Translation for Cultural Image Captioning: Gators submission for AmericasNLP 2026 shared task
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
von: Dhawan, Aashish, et al.
Veröffentlicht: (2026)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
von: Ma, Ziqiao, et al.
Veröffentlicht: (2023)
von: Ma, Ziqiao, et al.
Veröffentlicht: (2023)
GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
von: Omri, Yasmine, et al.
Veröffentlicht: (2025)
von: Omri, Yasmine, et al.
Veröffentlicht: (2025)
Order Is Not Layout: Order-to-Space Bias in Image Generation
von: Zhang, Yongkang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongkang, et al.
Veröffentlicht: (2026)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates
von: Gomes, Gonçalo, et al.
Veröffentlicht: (2025)
von: Gomes, Gonçalo, et al.
Veröffentlicht: (2025)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
von: Jung, Jaeyoon, et al.
Veröffentlicht: (2026)
von: Jung, Jaeyoon, et al.
Veröffentlicht: (2026)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
von: Verma, Gaurav, et al.
Veröffentlicht: (2024)
von: Verma, Gaurav, et al.
Veröffentlicht: (2024)
Decoding Diffusion: A Scalable Framework for Unsupervised Analysis of Latent Space Biases and Representations Using Natural Language Prompts
von: Zeng, E. Zhixuan, et al.
Veröffentlicht: (2024)
von: Zeng, E. Zhixuan, et al.
Veröffentlicht: (2024)
In-Context Meta LoRA Generation
von: Shao, Yihua, et al.
Veröffentlicht: (2025)
von: Shao, Yihua, et al.
Veröffentlicht: (2025)
Evaluating Automated Radiology Report Quality through Fine-Grained Phrasal Grounding of Clinical Findings
von: Mahmood, Razi, et al.
Veröffentlicht: (2024)
von: Mahmood, Razi, et al.
Veröffentlicht: (2024)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
von: Ananthram, Amith, et al.
Veröffentlicht: (2025)
von: Ananthram, Amith, et al.
Veröffentlicht: (2025)
DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2025)
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Scalable Vector Graphics-Driven Image Understanding
von: Cai, Mu, et al.
Veröffentlicht: (2023)
von: Cai, Mu, et al.
Veröffentlicht: (2023)
Retrieving Counterfactuals Improves Visual In-Context Learning
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
von: Xiong, Guangzhi, et al.
Veröffentlicht: (2026)
RingGesture: A Ring-Based Mid-Air Gesture Typing System Powered by a Deep-Learning Word Prediction Framework
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
von: Liang, Mingliang, et al.
Veröffentlicht: (2024)
WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
von: Ma, Jiawei, et al.
Veröffentlicht: (2024)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
von: Lee, Yebin, et al.
Veröffentlicht: (2024)
von: Lee, Yebin, et al.
Veröffentlicht: (2024)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
von: Tong, Tony Cheng, et al.
Veröffentlicht: (2024)
von: Tong, Tony Cheng, et al.
Veröffentlicht: (2024)
VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representation
von: Gehlaut, Tarun, et al.
Veröffentlicht: (2026)
von: Gehlaut, Tarun, et al.
Veröffentlicht: (2026)
Internalized Reasoning for Long-Context Visual Document Understanding
von: Veselka, Austin
Veröffentlicht: (2026)
von: Veselka, Austin
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
von: Broomfield, Julius, et al.
Veröffentlicht: (2025) -
StarVector: Generating Scalable Vector Graphics Code from Images and Text
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2023) -
Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024) -
Counting Without Numbers and Finding Without Words
von: Patro, Badri Narayana
Veröffentlicht: (2026) -
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models
von: He, Zoe Wanying, et al.
Veröffentlicht: (2025)