V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ge, Junqi, Chen, Ziyi, Lin, Jintao, Zhu, Jinguo, Liu, Xihui, Dai, Jifeng, Zhu, Xizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025)
von: Liu, Shi, et al.
Veröffentlicht: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
von: Luo, Gen, et al.
Veröffentlicht: (2024)
von: Luo, Gen, et al.
Veröffentlicht: (2024)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
von: Xu, Weiye, et al.
Veröffentlicht: (2025)
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
von: Fang, Rongyao, et al.
Veröffentlicht: (2025)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
von: Wang, Weiyun, et al.
Veröffentlicht: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
Parameter-Inverted Image Pyramid Networks
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2025)
von: Li, Haoran, et al.
Veröffentlicht: (2025)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
von: Yang, Chenyu, et al.
Veröffentlicht: (2024)
Revisiting Multimodal Positional Encoding in Vision-Language Models
von: Huang, Jie, et al.
Veröffentlicht: (2025)
von: Huang, Jie, et al.
Veröffentlicht: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
von: Meng, Fanqing, et al.
Veröffentlicht: (2024)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
von: Wu, Jiannan, et al.
Veröffentlicht: (2024)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
von: Duan, Yuchen, et al.
Veröffentlicht: (2024)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
von: Yin, Ziyi, et al.
Veröffentlicht: (2023)
von: Yin, Ziyi, et al.
Veröffentlicht: (2023)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
von: Gao, Zhangwei, et al.
Veröffentlicht: (2024)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
An Investigation on The Position Encoding in Vision-Based Dynamics Prediction
von: Zhu, Jiageng, et al.
Veröffentlicht: (2024)
von: Zhu, Jiageng, et al.
Veröffentlicht: (2024)
AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models
von: Lin, Jintao, et al.
Veröffentlicht: (2026)
von: Lin, Jintao, et al.
Veröffentlicht: (2026)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
von: Chen, Zhe, et al.
Veröffentlicht: (2023)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
Learning 1D Causal Visual Representation with De-focus Attention Networks
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
von: Cui, Erfei, et al.
Veröffentlicht: (2023)
big.LITTLE Vision Transformer for Efficient Visual Recognition
von: Guo, He, et al.
Veröffentlicht: (2024)
von: Guo, He, et al.
Veröffentlicht: (2024)
OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
von: Huang, Ruoxiang, et al.
Veröffentlicht: (2025)
von: Huang, Ruoxiang, et al.
Veröffentlicht: (2025)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
von: Yang, Chenyu, et al.
Veröffentlicht: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
von: Wang, Weiyun, et al.
Veröffentlicht: (2024)
OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
von: Lin, Jingli, et al.
Veröffentlicht: (2025)
Improving Visual Storytelling with Multimodal Large Language Models
von: Lin, Xiaochuan, et al.
Veröffentlicht: (2024)
von: Lin, Xiaochuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025) -
CoMemo: LVLMs Need Image Context with Image Memory
von: Liu, Shi, et al.
Veröffentlicht: (2025) -
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
von: Wang, Weiyun, et al.
Veröffentlicht: (2024) -
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
von: Luo, Gen, et al.
Veröffentlicht: (2024) -
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024)