VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Kelly, Chris, Hu, Luhui, Yang, Bang, Tian, Yu, Yang, Deshun, Yang, Cindy, Huang, Zaoshan, Li, Zihao, Hu, Jiayin, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024)
by: Kelly, Chris, et al.
Published: (2024)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
by: Yang, Deshun, et al.
Published: (2024)
by: Yang, Deshun, et al.
Published: (2024)
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
by: Yang, Bang, et al.
Published: (2023)
by: Yang, Bang, et al.
Published: (2023)
Environmental Understanding Vision-Language Model for Embodied Agent
by: Bang, Jinsik, et al.
Published: (2026)
by: Bang, Jinsik, et al.
Published: (2026)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation
by: Li, Jiulin, et al.
Published: (2025)
by: Li, Jiulin, et al.
Published: (2025)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
by: Yang, Bang, et al.
Published: (2024)
by: Yang, Bang, et al.
Published: (2024)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
by: Lu, Xu, et al.
Published: (2025)
by: Lu, Xu, et al.
Published: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
by: Wang, Hanbin, et al.
Published: (2025)
by: Wang, Hanbin, et al.
Published: (2025)
Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained Models
by: Shen, Xuelin, et al.
Published: (2025)
by: Shen, Xuelin, et al.
Published: (2025)
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
by: Yang, Jingcheng, et al.
Published: (2026)
by: Yang, Jingcheng, et al.
Published: (2026)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2023)
by: Tian, Xinyu, et al.
Published: (2023)
Generalized Robot Learning Framework
by: Yan, Jiahuan, et al.
Published: (2024)
by: Yan, Jiahuan, et al.
Published: (2024)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
by: Tang, Lexiang, et al.
Published: (2025)
by: Tang, Lexiang, et al.
Published: (2025)
RegionGPT: Towards Region Understanding Vision Language Model
by: Guo, Qiushan, et al.
Published: (2024)
by: Guo, Qiushan, et al.
Published: (2024)
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
by: Huang, Jialei, et al.
Published: (2025)
by: Huang, Jialei, et al.
Published: (2025)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
by: Nguyen, Phu-Vinh, et al.
Published: (2025)
by: Nguyen, Phu-Vinh, et al.
Published: (2025)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
by: Yang, Lijin, et al.
Published: (2026)
by: Yang, Lijin, et al.
Published: (2026)
Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
by: Li, Yang, et al.
Published: (2026)
by: Li, Yang, et al.
Published: (2026)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
by: Tian, Yuhao, et al.
Published: (2025)
by: Tian, Yuhao, et al.
Published: (2025)
AgentVLN: Towards Agentic Vision-and-Language Navigation
by: Xin, Zihao, et al.
Published: (2026)
by: Xin, Zihao, et al.
Published: (2026)
Syntax-Aware Complex-Valued Neural Machine Translation
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
by: Zou, Shu, et al.
Published: (2025)
by: Zou, Shu, et al.
Published: (2025)
Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
by: Xu, Wanting, et al.
Published: (2024)
by: Xu, Wanting, et al.
Published: (2024)
Técnica para dibujar crisantemo / Wang Deshun
by: Wang Deshun
by: Wang Deshun
Técnica para dibujar peonía / Wang Deshun
by: Wang Deshun
by: Wang Deshun
Understanding Degradation with Vision Language Model
by: Lan, Guanzhou, et al.
Published: (2026)
by: Lan, Guanzhou, et al.
Published: (2026)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
by: Ackermann, Jan, et al.
Published: (2025)
by: Ackermann, Jan, et al.
Published: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
by: Qian, Kangan, et al.
Published: (2025)
by: Qian, Kangan, et al.
Published: (2025)
FinVision: A Multi-Agent Framework for Stock Market Prediction
by: Fatemi, Sorouralsadat, et al.
Published: (2024)
by: Fatemi, Sorouralsadat, et al.
Published: (2024)
Mastering Diverse, Unknown, and Cluttered Tracks for Robust Vision-Based Drone Racing
by: Yu, Feng, et al.
Published: (2025)
by: Yu, Feng, et al.
Published: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
Similar Items
-
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
by: Kelly, Chris, et al.
Published: (2024) -
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
by: Yang, Deshun, et al.
Published: (2024) -
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
by: Wang, Hao, et al.
Published: (2024) -
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
by: Yang, Bang, et al.
Published: (2023) -
Environmental Understanding Vision-Language Model for Embodied Agent
by: Bang, Jinsik, et al.
Published: (2026)