TextVidBench: A Benchmark for Long Video Scene Text Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhong, Yangyang, Qi, Ji, Yao, Yuan, Luo, Pengxin, Yan, Yunfeng, Qi, Donglian, Liu, Zhiyuan, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
par: He, Weizhen, et autres
Publié: (2025)
par: He, Weizhen, et autres
Publié: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
par: Qi, Ji, et autres
Publié: (2025)
par: Qi, Ji, et autres
Publié: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
par: Zhou, Sheng, et autres
Publié: (2024)
par: Zhou, Sheng, et autres
Publié: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
par: Qin, Bosheng, et autres
Publié: (2023)
par: Qin, Bosheng, et autres
Publié: (2023)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
par: Zhou, Sheng, et autres
Publié: (2025)
par: Zhou, Sheng, et autres
Publié: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
par: Yang, Zhoufaran, et autres
Publié: (2025)
par: Yang, Zhoufaran, et autres
Publié: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
par: Chu, Meng, et autres
Publié: (2025)
par: Chu, Meng, et autres
Publié: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
par: Jin, Zhe, et autres
Publié: (2025)
par: Jin, Zhe, et autres
Publié: (2025)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
par: Fei, Hao, et autres
Publié: (2023)
par: Fei, Hao, et autres
Publié: (2023)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
par: Chu, Meng, et autres
Publié: (2023)
par: Chu, Meng, et autres
Publié: (2023)
Universal Scene Graph Generation
par: Wu, Shengqiong, et autres
Publié: (2025)
par: Wu, Shengqiong, et autres
Publié: (2025)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
par: Chen, Yiyang, et autres
Publié: (2022)
par: Chen, Yiyang, et autres
Publié: (2022)
TextSculptor: Training and Benchmarking Scene Text Editing
par: Lin, Yiheng, et autres
Publié: (2026)
par: Lin, Yiheng, et autres
Publié: (2026)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
par: Yang, Yiming, et autres
Publié: (2025)
par: Yang, Yiming, et autres
Publié: (2025)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
par: Ni, Zhenliang, et autres
Publié: (2025)
par: Ni, Zhenliang, et autres
Publié: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
par: Wang, Yi, et autres
Publié: (2023)
par: Wang, Yi, et autres
Publié: (2023)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
par: Zhu, Fengbin, et autres
Publié: (2024)
par: Zhu, Fengbin, et autres
Publié: (2024)
Text-Pass Filter: An Efficient Scene Text Detector
par: Yang, Chuang, et autres
Publié: (2026)
par: Yang, Chuang, et autres
Publié: (2026)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
par: Chen, Seng Nam, et autres
Publié: (2026)
par: Chen, Seng Nam, et autres
Publié: (2026)
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
par: Liu, Zhiyuan, et autres
Publié: (2024)
par: Liu, Zhiyuan, et autres
Publié: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
par: Fang, Xiang, et autres
Publié: (2026)
par: Fang, Xiang, et autres
Publié: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
par: Xu, Yicheng, et autres
Publié: (2025)
par: Xu, Yicheng, et autres
Publié: (2025)
Can I Trust Your Answer? Visually Grounded Video Question Answering
par: Xiao, Junbin, et autres
Publié: (2023)
par: Xiao, Junbin, et autres
Publié: (2023)
Discriminative Probing and Tuning for Text-to-Image Generation
par: Qu, Leigang, et autres
Publié: (2024)
par: Qu, Leigang, et autres
Publié: (2024)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
par: Li, Qi, et autres
Publié: (2025)
par: Li, Qi, et autres
Publié: (2025)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
par: Tzachor, Issar, et autres
Publié: (2026)
par: Tzachor, Issar, et autres
Publié: (2026)
OmniVid: A Generative Framework for Universal Video Understanding
par: Wang, Junke, et autres
Publié: (2024)
par: Wang, Junke, et autres
Publié: (2024)
ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
par: Zhou, Zhenglin, et autres
Publié: (2025)
par: Zhou, Zhenglin, et autres
Publié: (2025)
Text-Guided Mixup Towards Long-Tailed Image Categorization
par: Franklin, Richard, et autres
Publié: (2024)
par: Franklin, Richard, et autres
Publié: (2024)
LET-US: Long Event-Text Understanding of Scenes
par: Chen, Rui, et autres
Publié: (2025)
par: Chen, Rui, et autres
Publié: (2025)
Extending Visual Dynamics for Video-to-Music Generation
par: Liu, Xiaohao, et autres
Publié: (2025)
par: Liu, Xiaohao, et autres
Publié: (2025)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
par: Zhou, Ziwei, et autres
Publié: (2026)
par: Zhou, Ziwei, et autres
Publié: (2026)
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
par: Jin, Kaiming, et autres
Publié: (2026)
par: Jin, Kaiming, et autres
Publié: (2026)
VidHal: Benchmarking Temporal Hallucinations in Vision LLMs
par: Choong, Wey Yeh, et autres
Publié: (2024)
par: Choong, Wey Yeh, et autres
Publié: (2024)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
par: Yuan, Shenghai, et autres
Publié: (2024)
par: Yuan, Shenghai, et autres
Publié: (2024)
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
par: Shi, Enyi, et autres
Publié: (2026)
par: Shi, Enyi, et autres
Publié: (2026)
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
par: Zanchetta, Luca, et autres
Publié: (2025)
par: Zanchetta, Luca, et autres
Publié: (2025)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
par: Lan, Xing, et autres
Publié: (2024)
par: Lan, Xing, et autres
Publié: (2024)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
par: Feng, Weixi, et autres
Publié: (2025)
par: Feng, Weixi, et autres
Publié: (2025)
Towards 3D Molecule-Text Interpretation in Language Models
par: Li, Sihang, et autres
Publié: (2024)
par: Li, Sihang, et autres
Publié: (2024)
Documents similaires
-
Adept: Annotation-Denoising Auxiliary Tasks with Discrete Cosine Transform Map and Keypoint for Human-Centric Pretraining
par: He, Weizhen, et autres
Publié: (2025) -
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
par: Qi, Ji, et autres
Publié: (2025) -
Scene-Text Grounding for Text-Based Video Question Answering
par: Zhou, Sheng, et autres
Publié: (2024) -
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
par: Qin, Bosheng, et autres
Publié: (2023) -
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
par: Zhou, Sheng, et autres
Publié: (2025)