POINTS1.5: Building a Vision-Language Model towards Real World Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yuan, Tian, Le, Zhou, Xiao, Gao, Xinyu, Yu, Kavio, Yu, Yang, Zhou, Jie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
POINTS: Improving Your Vision-language Model with Affordable Strategies
di: Liu, Yuan, et al.
Pubblicazione: (2024)
di: Liu, Yuan, et al.
Pubblicazione: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
di: Zhou, Yuchen, et al.
Pubblicazione: (2025)
di: Zhou, Yuchen, et al.
Pubblicazione: (2025)
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026)
di: Song, Shezheng, et al.
Pubblicazione: (2026)
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
di: Liu, Yuan, et al.
Pubblicazione: (2025)
di: Liu, Yuan, et al.
Pubblicazione: (2025)
MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
di: Qiao, Xiangshuo, et al.
Pubblicazione: (2024)
di: Qiao, Xiangshuo, et al.
Pubblicazione: (2024)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
di: Huang, Yuhang, et al.
Pubblicazione: (2024)
di: Huang, Yuhang, et al.
Pubblicazione: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
di: Tang, Hao, et al.
Pubblicazione: (2026)
di: Tang, Hao, et al.
Pubblicazione: (2026)
Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
di: Gao, Ziyuan, et al.
Pubblicazione: (2025)
di: Gao, Ziyuan, et al.
Pubblicazione: (2025)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
di: Tian, Huilin, et al.
Pubblicazione: (2024)
di: Tian, Huilin, et al.
Pubblicazione: (2024)
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
di: Yu, Yanrui, et al.
Pubblicazione: (2025)
di: Yu, Yanrui, et al.
Pubblicazione: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
di: Xiao, Xinyu, et al.
Pubblicazione: (2026)
di: Xiao, Xinyu, et al.
Pubblicazione: (2026)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
di: Zhao, Yumiao, et al.
Pubblicazione: (2024)
di: Zhao, Yumiao, et al.
Pubblicazione: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
VIoTGPT: Learning to Schedule Vision Tools in LLMs towards Intelligent Video Internet of Things
di: Zhong, Yaoyao, et al.
Pubblicazione: (2023)
di: Zhong, Yaoyao, et al.
Pubblicazione: (2023)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
di: Fan, Linfeng, et al.
Pubblicazione: (2026)
di: Fan, Linfeng, et al.
Pubblicazione: (2026)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
di: Guo, Zile, et al.
Pubblicazione: (2026)
di: Guo, Zile, et al.
Pubblicazione: (2026)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
di: Kong, Fanheng, et al.
Pubblicazione: (2025)
di: Kong, Fanheng, et al.
Pubblicazione: (2025)
Unveiling Encoder-Free Vision-Language Models
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025)
di: Watanabe, Mitsuki, et al.
Pubblicazione: (2025)
Grounded Chain-of-Thought for Multimodal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2025)
di: Wu, Qiong, et al.
Pubblicazione: (2025)
Text-Only Data Synthesis for Vision Language Model Training
di: Yu, Xiaomin, et al.
Pubblicazione: (2025)
di: Yu, Xiaomin, et al.
Pubblicazione: (2025)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
ComAlign: Compositional Alignment in Vision-Language Models
di: Abdollah, Ali, et al.
Pubblicazione: (2024)
di: Abdollah, Ali, et al.
Pubblicazione: (2024)
InstructFLIP: Exploring Unified Vision-Language Model for Face Anti-spoofing
di: Lin, Kun-Hsiang, et al.
Pubblicazione: (2025)
di: Lin, Kun-Hsiang, et al.
Pubblicazione: (2025)
Adaptive 3D Gaussian Splatting Video Streaming
di: Gong, Han, et al.
Pubblicazione: (2025)
di: Gong, Han, et al.
Pubblicazione: (2025)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
di: Hao, Jing, et al.
Pubblicazione: (2025)
di: Hao, Jing, et al.
Pubblicazione: (2025)
DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
di: Zhang, Peng-Fei, et al.
Pubblicazione: (2026)
di: Zhang, Peng-Fei, et al.
Pubblicazione: (2026)
Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
di: Zhang, Yiyun, et al.
Pubblicazione: (2023)
di: Zhang, Yiyun, et al.
Pubblicazione: (2023)
DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
di: Yang, Wei, et al.
Pubblicazione: (2025)
di: Yang, Wei, et al.
Pubblicazione: (2025)
Diverse Sign Language Translation
di: Shen, Xin, et al.
Pubblicazione: (2024)
di: Shen, Xin, et al.
Pubblicazione: (2024)
POINTS-GUI-G: GUI-Grounding Journey
di: Zhao, Zhongyin, et al.
Pubblicazione: (2026)
di: Zhao, Zhongyin, et al.
Pubblicazione: (2026)
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2026)
di: Wasi, Azmine Toushik, et al.
Pubblicazione: (2026)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
di: Ma, Jingtian, et al.
Pubblicazione: (2025)
di: Ma, Jingtian, et al.
Pubblicazione: (2025)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
POINTS: Improving Your Vision-language Model with Affordable Strategies
di: Liu, Yuan, et al.
Pubblicazione: (2024) -
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
di: Zhou, Yuchen, et al.
Pubblicazione: (2025) -
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
di: Xu, Jiaqi, et al.
Pubblicazione: (2024) -
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
di: Song, Shezheng, et al.
Pubblicazione: (2026) -
POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
di: Liu, Yuan, et al.
Pubblicazione: (2025)