LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lan, Zhibin, Niu, Liqiang, Meng, Fandong, Zhou, Jie, Su, Jinsong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
von: Yu, Fangxu, et al.
Veröffentlicht: (2026)
von: Yu, Fangxu, et al.
Veröffentlicht: (2026)
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
von: Chen, Wenchao, et al.
Veröffentlicht: (2024)
von: Chen, Wenchao, et al.
Veröffentlicht: (2024)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
von: Chang, Yue, et al.
Veröffentlicht: (2024)
von: Chang, Yue, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Why do LLaVA Vision-Language Models Reply to Images in English?
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
von: Hinck, Musashi, et al.
Veröffentlicht: (2024)
D2C: Unlocking the Potential of Continuous Autoregressive Image Generation with Discrete Tokens
von: Wang, Panpan, et al.
Veröffentlicht: (2025)
von: Wang, Panpan, et al.
Veröffentlicht: (2025)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
von: Jing, Liqiang, et al.
Veröffentlicht: (2025)
SQ-LLaVA: Self-Questioning for Large Vision-Language Assistant
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
von: Sun, Guohao, et al.
Veröffentlicht: (2024)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
von: Lan, Zhibin, et al.
Veröffentlicht: (2025)
von: Lan, Zhibin, et al.
Veröffentlicht: (2025)
PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
von: Liang, Yuci, et al.
Veröffentlicht: (2024)
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
von: Agrawal, Aakriti, et al.
Veröffentlicht: (2025)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
von: You, Zebin, et al.
Veröffentlicht: (2025)
von: You, Zebin, et al.
Veröffentlicht: (2025)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
von: Feng, Jie, et al.
Veröffentlicht: (2025)
von: Feng, Jie, et al.
Veröffentlicht: (2025)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyan, et al.
Veröffentlicht: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
von: Castro, Santiago, et al.
Veröffentlicht: (2024)
von: Castro, Santiago, et al.
Veröffentlicht: (2024)
Dynamic Embedding of Hierarchical Visual Features for Efficient Vision-Language Fine-Tuning
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
von: Wei, Xinyu, et al.
Veröffentlicht: (2025)
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2025)
Lost in Embeddings: Information Loss in Vision-Language Models
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
von: Li, Wenyan, et al.
Veröffentlicht: (2025)
Efficient Few-Shot Medical Image Analysis via Hierarchical Contrastive Vision-Language Learning
von: Fuller, Harrison, et al.
Veröffentlicht: (2025)
von: Fuller, Harrison, et al.
Veröffentlicht: (2025)
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
von: Sun, Kai, et al.
Veröffentlicht: (2025)
von: Sun, Kai, et al.
Veröffentlicht: (2025)
WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
von: Sugiura, Issa, et al.
Veröffentlicht: (2025)
LLaVA-OneVision: Easy Visual Task Transfer
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
Where do Large Vision-Language Models Look at when Answering Questions?
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
von: Xu, Zhiyu, et al.
Veröffentlicht: (2026)
von: Xu, Zhiyu, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models (LVLMs) via Language-Contrastive Decoding (LCD)
von: Manevich, Avshalom, et al.
Veröffentlicht: (2024)
von: Manevich, Avshalom, et al.
Veröffentlicht: (2024)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
von: Zhu, Yichen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024) -
Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation
von: Lan, Zhibin, et al.
Veröffentlicht: (2024) -
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
von: Liu, Juntao, et al.
Veröffentlicht: (2025) -
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
von: Yu, Fangxu, et al.
Veröffentlicht: (2026) -
FGAIF: Aligning Large Vision-Language Models with Fine-grained AI Feedback
von: Jing, Liqiang, et al.
Veröffentlicht: (2024)