VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kelly, Chris, Hu, Luhui, Yang, Bang, Tian, Yu, Yang, Deshun, Yang, Cindy, Huang, Zaoshan, Li, Zihao, Hu, Jiayin, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
von: Kelly, Chris, et al.
Veröffentlicht: (2024)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
von: Yang, Deshun, et al.
Veröffentlicht: (2024)
von: Yang, Deshun, et al.
Veröffentlicht: (2024)
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
von: Yang, Bang, et al.
Veröffentlicht: (2023)
von: Yang, Bang, et al.
Veröffentlicht: (2023)
Environmental Understanding Vision-Language Model for Embodied Agent
von: Bang, Jinsik, et al.
Veröffentlicht: (2026)
von: Bang, Jinsik, et al.
Veröffentlicht: (2026)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
von: Yang, Bo, et al.
Veröffentlicht: (2025)
von: Yang, Bo, et al.
Veröffentlicht: (2025)
A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation
von: Li, Jiulin, et al.
Veröffentlicht: (2025)
von: Li, Jiulin, et al.
Veröffentlicht: (2025)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
von: Yang, Bang, et al.
Veröffentlicht: (2024)
von: Yang, Bang, et al.
Veröffentlicht: (2024)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
von: Zhuang, Xianwei, et al.
Veröffentlicht: (2025)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
von: He, Mengqi, et al.
Veröffentlicht: (2025)
von: He, Mengqi, et al.
Veröffentlicht: (2025)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
von: Lu, Xu, et al.
Veröffentlicht: (2025)
von: Lu, Xu, et al.
Veröffentlicht: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained Models
von: Shen, Xuelin, et al.
Veröffentlicht: (2025)
von: Shen, Xuelin, et al.
Veröffentlicht: (2025)
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
von: Yang, Jingcheng, et al.
Veröffentlicht: (2026)
von: Yang, Jingcheng, et al.
Veröffentlicht: (2026)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
von: Tian, Xinyu, et al.
Veröffentlicht: (2023)
von: Tian, Xinyu, et al.
Veröffentlicht: (2023)
Generalized Robot Learning Framework
von: Yan, Jiahuan, et al.
Veröffentlicht: (2024)
von: Yan, Jiahuan, et al.
Veröffentlicht: (2024)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
von: Tang, Lexiang, et al.
Veröffentlicht: (2025)
von: Tang, Lexiang, et al.
Veröffentlicht: (2025)
RegionGPT: Towards Region Understanding Vision Language Model
von: Guo, Qiushan, et al.
Veröffentlicht: (2024)
von: Guo, Qiushan, et al.
Veröffentlicht: (2024)
Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
von: Huang, Jialei, et al.
Veröffentlicht: (2025)
von: Huang, Jialei, et al.
Veröffentlicht: (2025)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
Towards Open Environments and Instructions: General Vision-Language Navigation via Fast-Slow Interactive Reasoning
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
von: Tian, Yuhao, et al.
Veröffentlicht: (2025)
AgentVLN: Towards Agentic Vision-and-Language Navigation
von: Xin, Zihao, et al.
Veröffentlicht: (2026)
von: Xin, Zihao, et al.
Veröffentlicht: (2026)
Syntax-Aware Complex-Valued Neural Machine Translation
von: Liu, Yang, et al.
Veröffentlicht: (2023)
von: Liu, Yang, et al.
Veröffentlicht: (2023)
SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
von: Zou, Shu, et al.
Veröffentlicht: (2025)
von: Zou, Shu, et al.
Veröffentlicht: (2025)
Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
von: Xu, Wanting, et al.
Veröffentlicht: (2024)
Técnica para dibujar crisantemo / Wang Deshun
von: Wang Deshun
von: Wang Deshun
Técnica para dibujar peonía / Wang Deshun
von: Wang Deshun
von: Wang Deshun
Understanding Degradation with Vision Language Model
von: Lan, Guanzhou, et al.
Veröffentlicht: (2026)
von: Lan, Guanzhou, et al.
Veröffentlicht: (2026)
Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
von: Ackermann, Jan, et al.
Veröffentlicht: (2025)
von: Ackermann, Jan, et al.
Veröffentlicht: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
FinVision: A Multi-Agent Framework for Stock Market Prediction
von: Fatemi, Sorouralsadat, et al.
Veröffentlicht: (2024)
von: Fatemi, Sorouralsadat, et al.
Veröffentlicht: (2024)
Mastering Diverse, Unknown, and Cluttered Tracks for Robust Vision-Based Drone Racing
von: Yu, Feng, et al.
Veröffentlicht: (2025)
von: Yu, Feng, et al.
Veröffentlicht: (2025)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding
von: Kelly, Chris, et al.
Veröffentlicht: (2024) -
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
von: Yang, Deshun, et al.
Veröffentlicht: (2024) -
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
von: Wang, Hao, et al.
Veröffentlicht: (2024) -
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
von: Yang, Bang, et al.
Veröffentlicht: (2023) -
Environmental Understanding Vision-Language Model for Embodied Agent
von: Bang, Jinsik, et al.
Veröffentlicht: (2026)