Visual Large Language Models for Generalized and Specialized Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yifan, Lai, Zhixin, Bao, Wentao, Tan, Zhen, Dao, Anh, Sui, Kewei, Shen, Jiayi, Liu, Dong, Liu, Huan, Kong, Yu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Window Token Concatenation for Efficient Visual Large Language Models
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Facial Affective Behavior Analysis with Instruction Tuning
di: Li, Yifan, et al.
Pubblicazione: (2024)
di: Li, Yifan, et al.
Pubblicazione: (2024)
IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Prompting Language-Informed Distribution for Compositional Zero-Shot Learning
di: Bao, Wentao, et al.
Pubblicazione: (2023)
di: Bao, Wentao, et al.
Pubblicazione: (2023)
Task-Aware Resolution Optimization for Visual Large Language Models
di: Luo, Weiqing, et al.
Pubblicazione: (2025)
di: Luo, Weiqing, et al.
Pubblicazione: (2025)
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model
di: Zhao, Chengshuai, et al.
Pubblicazione: (2026)
di: Zhao, Chengshuai, et al.
Pubblicazione: (2026)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Open Set Face Forgery Detection via Dual-Level Evidence Collection
di: Cai, Zhongyi, et al.
Pubblicazione: (2025)
di: Cai, Zhongyi, et al.
Pubblicazione: (2025)
RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses
di: Nguyen, Minh Anh, et al.
Pubblicazione: (2026)
di: Nguyen, Minh Anh, et al.
Pubblicazione: (2026)
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
di: Yang, Jie, et al.
Pubblicazione: (2025)
di: Yang, Jie, et al.
Pubblicazione: (2025)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
di: Zhang, Jiaxin, et al.
Pubblicazione: (2024)
Advancing Visual Large Language Model for Multi-granular Versatile Perception
di: Xiang, Wentao, et al.
Pubblicazione: (2025)
di: Xiang, Wentao, et al.
Pubblicazione: (2025)
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
di: Chen, Xupeng, et al.
Pubblicazione: (2024)
di: Chen, Xupeng, et al.
Pubblicazione: (2024)
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
di: Zhu, Minghao, et al.
Pubblicazione: (2024)
di: Zhu, Minghao, et al.
Pubblicazione: (2024)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
di: Dao, Quan, et al.
Pubblicazione: (2024)
di: Dao, Quan, et al.
Pubblicazione: (2024)
AnchoredDream: Zero-Shot 360° Indoor Scene Generation from a Single View via Geometric Grounding
di: Yao, Runmao, et al.
Pubblicazione: (2026)
di: Yao, Runmao, et al.
Pubblicazione: (2026)
Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
di: Tao, Keda, et al.
Pubblicazione: (2025)
di: Tao, Keda, et al.
Pubblicazione: (2025)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
di: Du, Yingjun, et al.
Pubblicazione: (2023)
di: Du, Yingjun, et al.
Pubblicazione: (2023)
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
di: Tao, Keda, et al.
Pubblicazione: (2024)
di: Tao, Keda, et al.
Pubblicazione: (2024)
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
di: Fu, Luxuan, et al.
Pubblicazione: (2026)
di: Fu, Luxuan, et al.
Pubblicazione: (2026)
Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models
di: Liu, Zhaoyi, et al.
Pubblicazione: (2025)
di: Liu, Zhaoyi, et al.
Pubblicazione: (2025)
A Survey on Human Interaction Motion Generation
di: Sui, Kewei, et al.
Pubblicazione: (2025)
di: Sui, Kewei, et al.
Pubblicazione: (2025)
HoliTom: Holistic Token Merging for Fast Video Large Language Models
di: Shao, Kele, et al.
Pubblicazione: (2025)
di: Shao, Kele, et al.
Pubblicazione: (2025)
VRSO: Visual-Centric Reconstruction for Static Object Annotation
di: Yu, Chenyao, et al.
Pubblicazione: (2024)
di: Yu, Chenyao, et al.
Pubblicazione: (2024)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
di: Kong, Quan, et al.
Pubblicazione: (2026)
di: Kong, Quan, et al.
Pubblicazione: (2026)
Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
di: Liu, Jie, et al.
Pubblicazione: (2025)
di: Liu, Jie, et al.
Pubblicazione: (2025)
DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation
di: Phung, Hao, et al.
Pubblicazione: (2024)
di: Phung, Hao, et al.
Pubblicazione: (2024)
Causal Debiasing for Visual Commonsense Reasoning
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
di: Zou, Jiayi, et al.
Pubblicazione: (2025)
Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences
di: Li, Bao, et al.
Pubblicazione: (2026)
di: Li, Bao, et al.
Pubblicazione: (2026)
HyperSeg: Towards Universal Visual Segmentation with Large Language Model
di: Wei, Cong, et al.
Pubblicazione: (2024)
di: Wei, Cong, et al.
Pubblicazione: (2024)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
di: Xu, Zhiyang, et al.
Pubblicazione: (2024)
di: Xu, Zhiyang, et al.
Pubblicazione: (2024)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
di: Jiang, Dongsheng, et al.
Pubblicazione: (2023)
di: Jiang, Dongsheng, et al.
Pubblicazione: (2023)
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
di: Tong, Bo, et al.
Pubblicazione: (2024)
di: Tong, Bo, et al.
Pubblicazione: (2024)
Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding
di: Dao, Anh, et al.
Pubblicazione: (2025)
di: Dao, Anh, et al.
Pubblicazione: (2025)
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
di: Dong, Yuhao, et al.
Pubblicazione: (2024)
di: Dong, Yuhao, et al.
Pubblicazione: (2024)
Distilling Specialized Orders for Visual Generation
di: Pramanik, Rishav, et al.
Pubblicazione: (2025)
di: Pramanik, Rishav, et al.
Pubblicazione: (2025)
LiteGPT: Large Vision-Language Model for Joint Chest X-ray Localization and Classification Task
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
di: Liu, Zhongye, et al.
Pubblicazione: (2024)
di: Liu, Zhongye, et al.
Pubblicazione: (2024)
Harmonizing Visual Text Comprehension and Generation
di: Zhao, Zhen, et al.
Pubblicazione: (2024)
di: Zhao, Zhen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Window Token Concatenation for Efficient Visual Large Language Models
di: Li, Yifan, et al.
Pubblicazione: (2025) -
Facial Affective Behavior Analysis with Instruction Tuning
di: Li, Yifan, et al.
Pubblicazione: (2024) -
IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
di: Li, Yifan, et al.
Pubblicazione: (2025) -
Prompting Language-Informed Distribution for Compositional Zero-Shot Learning
di: Bao, Wentao, et al.
Pubblicazione: (2023) -
Task-Aware Resolution Optimization for Visual Large Language Models
di: Luo, Weiqing, et al.
Pubblicazione: (2025)