ParGo: Bridging Vision-Language with Partial and Global Views
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, An-Lan, Shan, Bin, Shi, Wei, Lin, Kun-Yu, Fei, Xiang, Tang, Guozhi, Liao, Lei, Huang, Can, Tang, Jingqun, Zheng, Wei-Shi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
von: Shan, Bin, et al.
Veröffentlicht: (2024)
von: Shan, Bin, et al.
Veröffentlicht: (2024)
Therapeutic strategies for aberrant splicing in cancer and genetic disorders
von: Wenhua Shi, et al.
Veröffentlicht: (2024)
von: Wenhua Shi, et al.
Veröffentlicht: (2024)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
von: Yu, Haiyang, et al.
Veröffentlicht: (2025)
Advancing Sequential Numerical Prediction in Autoregressive Models
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
von: Feng, Hao, et al.
Veröffentlicht: (2025)
von: Feng, Hao, et al.
Veröffentlicht: (2025)
Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting
von: Feng, Hao, et al.
Veröffentlicht: (2026)
von: Feng, Hao, et al.
Veröffentlicht: (2026)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
von: Xue, Wei, et al.
Veröffentlicht: (2026)
von: Xue, Wei, et al.
Veröffentlicht: (2026)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
von: Wang, An-Lan, et al.
Veröffentlicht: (2025)
von: Wang, An-Lan, et al.
Veröffentlicht: (2025)
Partial confinement in a quantum-link simulator
von: Tang, Zheng, et al.
Veröffentlicht: (2024)
von: Tang, Zheng, et al.
Veröffentlicht: (2024)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
Vision as LoRA
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
Can National ESG Inhibit the Impact of Extreme Climate on Global Financial Risks?
von: Haonan Wang, et al.
Veröffentlicht: (2025)
von: Haonan Wang, et al.
Veröffentlicht: (2025)
Geo-EVS: Geometry-Conditioned Extrapolative View Synthesis for Autonomous Driving
von: Lan, Yatong, et al.
Veröffentlicht: (2026)
von: Lan, Yatong, et al.
Veröffentlicht: (2026)
Post-Completion Learning for Language Models
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
von: Fei, Xiang, et al.
Veröffentlicht: (2025)
MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
von: Jia, Weitao, et al.
Veröffentlicht: (2025)
von: Jia, Weitao, et al.
Veröffentlicht: (2025)
PRET: Planning with Directed Fidelity Trajectory for Vision and Language Navigation
von: Lu, Renjie, et al.
Veröffentlicht: (2024)
von: Lu, Renjie, et al.
Veröffentlicht: (2024)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
von: Luo, Chuwei, et al.
Veröffentlicht: (2022)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
ParsTranslit: Truly Versatile Tajik-Farsi Transliteration
von: Merchant, Rayyan, et al.
Veröffentlicht: (2025)
von: Merchant, Rayyan, et al.
Veröffentlicht: (2025)
PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
von: Fei, Fan, et al.
Veröffentlicht: (2025)
von: Fei, Fan, et al.
Veröffentlicht: (2025)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
von: Tang, Jingqun, et al.
Veröffentlicht: (2024)
von: Tang, Jingqun, et al.
Veröffentlicht: (2024)
Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
von: Kong, Zhenglun, et al.
Veröffentlicht: (2025)
von: Kong, Zhenglun, et al.
Veröffentlicht: (2025)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
von: Zhu, Hanshen, et al.
Veröffentlicht: (2026)
von: Zhu, Hanshen, et al.
Veröffentlicht: (2026)
Human-Centric Transformer for Domain Adaptive Action Recognition
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
von: Lin, Kun-Yu, et al.
Veröffentlicht: (2024)
AerialGo: Walking-through City View Generation from Aerial Perspectives
von: Zhao, Fuqiang, et al.
Veröffentlicht: (2024)
von: Zhao, Fuqiang, et al.
Veröffentlicht: (2024)
Fluctuations of topological charges in two-dimensional classical Heisenberg model
von: Tang, Shan-Chang, et al.
Veröffentlicht: (2025)
von: Tang, Shan-Chang, et al.
Veröffentlicht: (2025)
HiGPT: Heterogeneous Graph Language Model
von: Tang, Jiabin, et al.
Veröffentlicht: (2024)
von: Tang, Jiabin, et al.
Veröffentlicht: (2024)
SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly
von: Zhu, Wei, et al.
Veröffentlicht: (2026)
von: Zhu, Wei, et al.
Veröffentlicht: (2026)
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
Global Maximum Principle for Partially Observed Risk-Sensitive Progressive Optimal Control of FBSDE with Poisson Jumps
von: Lin, Jingtao, et al.
Veröffentlicht: (2025)
von: Lin, Jingtao, et al.
Veröffentlicht: (2025)
Non-RS type cyclic MDS codes over finite fields via cyclotomic field reduction
von: Xiang, Can, et al.
Veröffentlicht: (2026)
von: Xiang, Can, et al.
Veröffentlicht: (2026)
Discontinuous Galerkin methods for the Laplace-Beltrami operator on point cloud
von: Dong, Guozhi, et al.
Veröffentlicht: (2020)
von: Dong, Guozhi, et al.
Veröffentlicht: (2020)
Task-Oriented 6-DoF Grasp Pose Detection in Clutters
von: Wang, An-Lan, et al.
Veröffentlicht: (2025)
von: Wang, An-Lan, et al.
Veröffentlicht: (2025)
TechCoach: Towards Technical-Point-Aware Descriptive Action Coaching
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2024)
Single-View Scene Point Cloud Human Grasp Generation
von: Wang, Yan-Kang, et al.
Veröffentlicht: (2024)
von: Wang, Yan-Kang, et al.
Veröffentlicht: (2024)
GoLF-NRT: Integrating Global Context and Local Geometry for Few-Shot View Synthesis
von: Wang, You, et al.
Veröffentlicht: (2025)
von: Wang, You, et al.
Veröffentlicht: (2025)
Harmonizing Visual Text Comprehension and Generation
von: Zhao, Zhen, et al.
Veröffentlicht: (2024)
von: Zhao, Zhen, et al.
Veröffentlicht: (2024)
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
von: Li, Zhuohao, et al.
Veröffentlicht: (2026)
von: Li, Zhuohao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
von: Shan, Bin, et al.
Veröffentlicht: (2024) -
Therapeutic strategies for aberrant splicing in cancer and genetic disorders
von: Wenhua Shi, et al.
Veröffentlicht: (2024) -
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
von: Yu, Haiyang, et al.
Veröffentlicht: (2025) -
Advancing Sequential Numerical Prediction in Autoregressive Models
von: Fei, Xiang, et al.
Veröffentlicht: (2025) -
Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
von: Feng, Hao, et al.
Veröffentlicht: (2025)