Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Xuchen, Hu, Shiyu, Feng, Xiaokun, Zhang, Dailing, Wu, Meiqi, Zhang, Jing, Huang, Kaiqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
di: Li, Xuchen, et al.
Pubblicazione: (2024)
di: Li, Xuchen, et al.
Pubblicazione: (2024)
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
di: Li, Xuchen, et al.
Pubblicazione: (2025)
di: Li, Xuchen, et al.
Pubblicazione: (2025)
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
di: Wu, Meiqi, et al.
Pubblicazione: (2025)
di: Wu, Meiqi, et al.
Pubblicazione: (2025)
Customizing Visual-Language Foundation Models for Multi-modal Anomaly Detection and Reasoning
di: Xu, Xiaohao, et al.
Pubblicazione: (2024)
di: Xu, Xiaohao, et al.
Pubblicazione: (2024)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
di: Li, Xuchen, et al.
Pubblicazione: (2025)
di: Li, Xuchen, et al.
Pubblicazione: (2025)
Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness
di: Chen, Honghao, et al.
Pubblicazione: (2024)
di: Chen, Honghao, et al.
Pubblicazione: (2024)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
di: Zhang, Ming, et al.
Pubblicazione: (2024)
di: Zhang, Ming, et al.
Pubblicazione: (2024)
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
di: Wu, Siwei, et al.
Pubblicazione: (2024)
di: Wu, Siwei, et al.
Pubblicazione: (2024)
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
di: Wu, Meiqi, et al.
Pubblicazione: (2024)
di: Wu, Meiqi, et al.
Pubblicazione: (2024)
Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V
di: Wu, Meiqi, et al.
Pubblicazione: (2026)
di: Wu, Meiqi, et al.
Pubblicazione: (2026)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
di: Zou, Chengke, et al.
Pubblicazione: (2024)
di: Zou, Chengke, et al.
Pubblicazione: (2024)
Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
di: Villa, Andrés, et al.
Pubblicazione: (2023)
di: Villa, Andrés, et al.
Pubblicazione: (2023)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
di: Li, Zhaowei, et al.
Pubblicazione: (2024)
di: Li, Zhaowei, et al.
Pubblicazione: (2024)
DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
di: Li, Xuzhao, et al.
Pubblicazione: (2025)
di: Li, Xuzhao, et al.
Pubblicazione: (2025)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
di: Sun, Hao, et al.
Pubblicazione: (2026)
di: Sun, Hao, et al.
Pubblicazione: (2026)
Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
di: Chen, Meiqi, et al.
Pubblicazione: (2024)
di: Chen, Meiqi, et al.
Pubblicazione: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
di: Wang, Jiuniu, et al.
Pubblicazione: (2025)
di: Wang, Jiuniu, et al.
Pubblicazione: (2025)
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
di: Chen, Honghao, et al.
Pubblicazione: (2025)
di: Chen, Honghao, et al.
Pubblicazione: (2025)
Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
di: Ye, Weihao, et al.
Pubblicazione: (2024)
di: Ye, Weihao, et al.
Pubblicazione: (2024)
Bootstrapping Referring Multi-Object Tracking
di: Zhang, Yani, et al.
Pubblicazione: (2024)
di: Zhang, Yani, et al.
Pubblicazione: (2024)
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval
di: Wu, Siwei, et al.
Pubblicazione: (2024)
di: Wu, Siwei, et al.
Pubblicazione: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
di: Yang, Rui, et al.
Pubblicazione: (2025)
di: Yang, Rui, et al.
Pubblicazione: (2025)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
di: Wu, Chengyue, et al.
Pubblicazione: (2024)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
di: Kang, Caixin, et al.
Pubblicazione: (2025)
di: Kang, Caixin, et al.
Pubblicazione: (2025)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
di: Ma, Yubo, et al.
Pubblicazione: (2024)
di: Ma, Yubo, et al.
Pubblicazione: (2024)
Rethinking Patient Education as Multi-turn Multi-modal Interaction
di: Yao, Zonghai, et al.
Pubblicazione: (2026)
di: Yao, Zonghai, et al.
Pubblicazione: (2026)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
di: Chen, Junjie, et al.
Pubblicazione: (2024)
di: Chen, Junjie, et al.
Pubblicazione: (2024)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
di: Li, Kaixin, et al.
Pubblicazione: (2024)
di: Li, Kaixin, et al.
Pubblicazione: (2024)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
di: Zhao, Yi, et al.
Pubblicazione: (2026)
di: Zhao, Yi, et al.
Pubblicazione: (2026)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
di: Wu, Yin, et al.
Pubblicazione: (2025)
di: Wu, Yin, et al.
Pubblicazione: (2025)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
di: Gong, Meiqi, et al.
Pubblicazione: (2025)
di: Gong, Meiqi, et al.
Pubblicazione: (2025)
Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
di: Zhang, Yichi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
di: Li, Xuchen, et al.
Pubblicazione: (2024) -
DTLLM-VLT: Diverse Text Generation for Visual Language Tracking Based on LLM
di: Li, Xuchen, et al.
Pubblicazione: (2024) -
How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
di: Li, Xuchen, et al.
Pubblicazione: (2024) -
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
di: Li, Xuchen, et al.
Pubblicazione: (2025) -
VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
di: Wu, Meiqi, et al.
Pubblicazione: (2025)