VidText: Towards Comprehensive Evaluation for Video Text Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhoufaran, Shu, Yan, Wang, Jing, Yang, Zhifei, Zhang, Yan, Li, Yu, Lu, Keyang, Zeng, Gangyan, Liu, Shaohui, Zhou, Yu, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
by: Zhang, Yan, et al.
Published: (2024)
by: Zhang, Yan, et al.
Published: (2024)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
by: Zhong, Yangyang, et al.
Published: (2025)
by: Zhong, Yangyang, et al.
Published: (2025)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
by: Zeng, Weichao, et al.
Published: (2024)
by: Zeng, Weichao, et al.
Published: (2024)
First Creating Backgrounds Then Rendering Texts: A New Paradigm for Visual Text Blending
by: Li, Zhenhang, et al.
Published: (2024)
by: Li, Zhenhang, et al.
Published: (2024)
Vision+X: A Survey on Multimodal Learning in the Light of Data
by: Zhu, Ye, et al.
Published: (2022)
by: Zhu, Ye, et al.
Published: (2022)
Visual Text Meets Low-level Vision: A Comprehensive Survey on Visual Text Processing
by: Shu, Yan, et al.
Published: (2024)
by: Shu, Yan, et al.
Published: (2024)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Towards Training-Free Scene Text Editing
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
Video-Browser: Towards Agentic Open-web Video Browsing
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
by: Zeng, Gangyan, et al.
Published: (2024)
by: Zeng, Gangyan, et al.
Published: (2024)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
SafeVid: Toward Safety Aligned Video Large Multimodal Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
by: Yang, Yiming, et al.
Published: (2025)
by: Yang, Yiming, et al.
Published: (2025)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
by: Zanchetta, Luca, et al.
Published: (2025)
by: Zanchetta, Luca, et al.
Published: (2025)
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
by: D'Incà, Moreno, et al.
Published: (2024)
by: D'Incà, Moreno, et al.
Published: (2024)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
VideoTetris: Towards Compositional Text-to-Video Generation
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
by: Rigo, Andrea, et al.
Published: (2025)
by: Rigo, Andrea, et al.
Published: (2025)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
by: Couairon, Paul, et al.
Published: (2023)
by: Couairon, Paul, et al.
Published: (2023)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
by: Yu, Xueqing, et al.
Published: (2026)
by: Yu, Xueqing, et al.
Published: (2026)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
by: Ma, Qi, et al.
Published: (2025)
by: Ma, Qi, et al.
Published: (2025)
Inflation with Diffusion: Efficient Temporal Adaptation for Text-to-Video Super-Resolution
by: Yuan, Xin, et al.
Published: (2024)
by: Yuan, Xin, et al.
Published: (2024)
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
by: Cheng, Sijie, et al.
Published: (2024)
by: Cheng, Sijie, et al.
Published: (2024)
Similar Items
-
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
by: Shu, Yan, et al.
Published: (2025) -
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025) -
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
by: Zhang, Yan, et al.
Published: (2024) -
TextVidBench: A Benchmark for Long Video Scene Text Understanding
by: Zhong, Yangyang, et al.
Published: (2025) -
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
by: Lyu, Jiahao, et al.
Published: (2024)