Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhan, Yufei, Zhao, Hongyin, Zhu, Yousong, Yang, Fan, Tang, Ming, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
von: Zhan, Yufei, et al.
Veröffentlicht: (2024)
von: Zhan, Yufei, et al.
Veröffentlicht: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2023)
von: Zhan, Yufei, et al.
Veröffentlicht: (2023)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2026)
von: Yang, Fan, et al.
Veröffentlicht: (2026)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
von: Yang, Fan, et al.
Veröffentlicht: (2025)
von: Yang, Fan, et al.
Veröffentlicht: (2025)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models
von: Ma, Kexin, et al.
Veröffentlicht: (2026)
von: Ma, Kexin, et al.
Veröffentlicht: (2026)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation
von: Zhu, Yuanbing, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanbing, et al.
Veröffentlicht: (2024)
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
von: Zhan, Yang, et al.
Veröffentlicht: (2024)
von: Zhan, Yang, et al.
Veröffentlicht: (2024)
Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
von: Jiao, Yang, et al.
Veröffentlicht: (2024)
Efficient Masked Autoencoders with Self-Consistency
von: Li, Zhaowen, et al.
Veröffentlicht: (2023)
von: Li, Zhaowen, et al.
Veröffentlicht: (2023)
Bridging Vision and Language: Optimal Transport-Driven Radiology Report Generation via LLMs
von: Zhao, Haifeng, et al.
Veröffentlicht: (2025)
von: Zhao, Haifeng, et al.
Veröffentlicht: (2025)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025)
von: Zheng, Xiangxi, et al.
Veröffentlicht: (2025)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
von: Zhou, Weijie, et al.
Veröffentlicht: (2025)
Vision-Driven Prompt Optimization for Large Language Models in Multimodal Generative Tasks
von: Franklin, Leo, et al.
Veröffentlicht: (2025)
von: Franklin, Leo, et al.
Veröffentlicht: (2025)
Object-Centric Vision Token Pruning for Vision Language Models
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
von: Wu, Jiulong, et al.
Veröffentlicht: (2025)
von: Wu, Jiulong, et al.
Veröffentlicht: (2025)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
von: Long, Zijun, et al.
Veröffentlicht: (2023)
von: Long, Zijun, et al.
Veröffentlicht: (2023)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
von: Zhao, Libang, et al.
Veröffentlicht: (2026)
von: Zhao, Libang, et al.
Veröffentlicht: (2026)
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
von: Yue, Tongtian, et al.
Veröffentlicht: (2025)
von: Yue, Tongtian, et al.
Veröffentlicht: (2025)
Can We Predict Performance of Large Models across Vision-Language Tasks?
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
von: Zhao, Qinyu, et al.
Veröffentlicht: (2024)
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
von: Schroeder, Philip, et al.
Veröffentlicht: (2025)
von: Schroeder, Philip, et al.
Veröffentlicht: (2025)
Vision-Language Models Unlock Task-Centric Latent Actions
von: Nikulin, Alexander, et al.
Veröffentlicht: (2026)
von: Nikulin, Alexander, et al.
Veröffentlicht: (2026)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
von: Yin, Ziyi, et al.
Veröffentlicht: (2023)
von: Yin, Ziyi, et al.
Veröffentlicht: (2023)
Synthetic Data is an Elegant GIFT for Continual Vision-Language Models
von: Wu, Bin, et al.
Veröffentlicht: (2025)
von: Wu, Bin, et al.
Veröffentlicht: (2025)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Tian, Xiaoyu, et al.
Veröffentlicht: (2024)
Bridging the Indoor-Outdoor Gap: Vision-Centric Instruction-Guided Embodied Navigation for the Last Meters
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Yuxiang, et al.
Veröffentlicht: (2026)
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
von: Zhang, Zilun, et al.
Veröffentlicht: (2025)
von: Zhang, Zilun, et al.
Veröffentlicht: (2025)
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
Vision-Language Models for Vision Tasks: A Survey
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2023)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
von: Zhan, Yufei, et al.
Veröffentlicht: (2024) -
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
von: Zhan, Yufei, et al.
Veröffentlicht: (2025) -
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2023) -
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
von: Yang, Fan, et al.
Veröffentlicht: (2026) -
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
von: Zheng, Shurong, et al.
Veröffentlicht: (2026)