The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Weiyun, Ren, Yiming, Luo, Haowen, Li, Tiantong, Yan, Chenxiang, Chen, Zhe, Wang, Wenhai, Li, Qingyun, Lu, Lewei, Zhu, Xizhou, Qiao, Yu, Dai, Jifeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
di: Duan, Yuchen, et al.
Pubblicazione: (2024)
di: Duan, Yuchen, et al.
Pubblicazione: (2024)
Needle In A Multimodal Haystack
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
di: Liu, Yangzhou, et al.
Pubblicazione: (2024)
di: Liu, Yangzhou, et al.
Pubblicazione: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)
di: Tian, Changyao, et al.
Pubblicazione: (2024)
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
di: Gao, Zhangwei, et al.
Pubblicazione: (2024)
di: Gao, Zhangwei, et al.
Pubblicazione: (2024)
Demystify Transformers & Convolutions in Modern Image Deep Networks
di: Hu, Xiaowei, et al.
Pubblicazione: (2022)
di: Hu, Xiaowei, et al.
Pubblicazione: (2022)
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
di: Xu, Weiye, et al.
Pubblicazione: (2025)
di: Xu, Weiye, et al.
Pubblicazione: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
di: Luo, Gen, et al.
Pubblicazione: (2025)
di: Luo, Gen, et al.
Pubblicazione: (2025)
CoMemo: LVLMs Need Image Context with Image Memory
di: Liu, Shi, et al.
Pubblicazione: (2025)
di: Liu, Shi, et al.
Pubblicazione: (2025)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
di: Ren, Yiming, et al.
Pubblicazione: (2025)
di: Ren, Yiming, et al.
Pubblicazione: (2025)
MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks
di: Su, Shiqian, et al.
Pubblicazione: (2026)
di: Su, Shiqian, et al.
Pubblicazione: (2026)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
di: Wang, Weiyun, et al.
Pubblicazione: (2025)
di: Wang, Weiyun, et al.
Pubblicazione: (2025)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
di: Cui, Erfei, et al.
Pubblicazione: (2023)
di: Cui, Erfei, et al.
Pubblicazione: (2023)
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
di: Luo, Gen, et al.
Pubblicazione: (2025)
di: Luo, Gen, et al.
Pubblicazione: (2025)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
Docopilot: Improving Multimodal Models for Document-Level Understanding
di: Duan, Yuchen, et al.
Pubblicazione: (2025)
di: Duan, Yuchen, et al.
Pubblicazione: (2025)
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
di: Chen, Junting, et al.
Pubblicazione: (2025)
di: Chen, Junting, et al.
Pubblicazione: (2025)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
di: Xiong, Yuwen, et al.
Pubblicazione: (2024)
di: Xiong, Yuwen, et al.
Pubblicazione: (2024)
Parameter-Inverted Image Pyramid Networks
di: Zhu, Xizhou, et al.
Pubblicazione: (2024)
di: Zhu, Xizhou, et al.
Pubblicazione: (2024)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
di: Li, Hao, et al.
Pubblicazione: (2023)
di: Li, Hao, et al.
Pubblicazione: (2023)
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
di: Wu, Jiannan, et al.
Pubblicazione: (2024)
di: Wu, Jiannan, et al.
Pubblicazione: (2024)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
di: Chen, Zhe, et al.
Pubblicazione: (2023)
di: Chen, Zhe, et al.
Pubblicazione: (2023)
Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
di: Yang, Yang, et al.
Pubblicazione: (2024)
di: Yang, Yang, et al.
Pubblicazione: (2024)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
di: Tian, Changyao, et al.
Pubblicazione: (2023)
di: Tian, Changyao, et al.
Pubblicazione: (2023)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
di: Wang, Zhaokai, et al.
Pubblicazione: (2025)
di: Wang, Zhaokai, et al.
Pubblicazione: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
di: Tian, Changyao, et al.
Pubblicazione: (2025)
di: Tian, Changyao, et al.
Pubblicazione: (2025)
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
di: Chen, Zhe, et al.
Pubblicazione: (2024)
di: Chen, Zhe, et al.
Pubblicazione: (2024)
Learning 1D Causal Visual Representation with De-focus Attention Networks
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
Towards a More Generalized Approach in Open Relation Extraction
di: Wang, Qing, et al.
Pubblicazione: (2025)
di: Wang, Qing, et al.
Pubblicazione: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
di: Chen, Zhe, et al.
Pubblicazione: (2024)
di: Chen, Zhe, et al.
Pubblicazione: (2024)
Sequential Diffusion Language Models
di: Liu, Yangzhou, et al.
Pubblicazione: (2025)
di: Liu, Yangzhou, et al.
Pubblicazione: (2025)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
di: Luo, Gen, et al.
Pubblicazione: (2024)
di: Luo, Gen, et al.
Pubblicazione: (2024)
ZeroGUI: Automating Online GUI Learning at Zero Human Cost
di: Yang, Chenyu, et al.
Pubblicazione: (2025)
di: Yang, Chenyu, et al.
Pubblicazione: (2025)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
di: Ge, Junqi, et al.
Pubblicazione: (2024)
di: Ge, Junqi, et al.
Pubblicazione: (2024)
Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-end Oriented Object Detection with Single Point Supervision
di: Yu, Yi, et al.
Pubblicazione: (2023)
di: Yu, Yi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
di: Duan, Yuchen, et al.
Pubblicazione: (2024) -
Needle In A Multimodal Haystack
di: Wang, Weiyun, et al.
Pubblicazione: (2024) -
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
di: Wang, Weiyun, et al.
Pubblicazione: (2024) -
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
di: Liu, Yangzhou, et al.
Pubblicazione: (2024) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)