Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yan, Zeng, Gangyan, Shen, Huawen, Wu, Daiqing, Zhou, Yu, Ma, Can |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
von: Zhang, Yan, et al.
Veröffentlicht: (2026)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026)
von: He, Haibin, et al.
Veröffentlicht: (2026)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
von: Xie, Enze, et al.
Veröffentlicht: (2024)
von: Xie, Enze, et al.
Veröffentlicht: (2024)
Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text Pairs
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
Resolving Sentiment Discrepancy for Multimodal Sentiment Detection via Semantics Completion and Decomposition
von: Wu, Daiqing, et al.
Veröffentlicht: (2024)
von: Wu, Daiqing, et al.
Veröffentlicht: (2024)
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
von: Li, Gengluo, et al.
Veröffentlicht: (2026)
EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
Winner Team Mia at TextVQA Challenge 2021: Vision-and-Language Representation Learning with Pre-trained Sequence-to-Sequence Model
von: Qiao, Yixuan, et al.
Veröffentlicht: (2021)
von: Qiao, Yixuan, et al.
Veröffentlicht: (2021)
Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
von: Yu, Chenmin, et al.
Veröffentlicht: (2026)
von: Yu, Chenmin, et al.
Veröffentlicht: (2026)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts
von: Li, Gengluo, et al.
Veröffentlicht: (2025)
von: Li, Gengluo, et al.
Veröffentlicht: (2025)
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
von: Du, Henghui, et al.
Veröffentlicht: (2025)
von: Du, Henghui, et al.
Veröffentlicht: (2025)
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
STCMOT: Spatio-Temporal Cohesion Learning for UAV-Based Multiple Object Tracking
von: Ma, Jianbo, et al.
Veröffentlicht: (2024)
von: Ma, Jianbo, et al.
Veröffentlicht: (2024)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2023)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2023)
SpatioTemporal Difference Network for Video Depth Super-Resolution
von: Wang, Zhengxue, et al.
Veröffentlicht: (2025)
von: Wang, Zhengxue, et al.
Veröffentlicht: (2025)
Towards Training-Free Scene Text Editing
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition
von: Chen, Tongjia, et al.
Veröffentlicht: (2023)
von: Chen, Tongjia, et al.
Veröffentlicht: (2023)
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
von: zhang, Kaixin, et al.
Veröffentlicht: (2026)
von: zhang, Kaixin, et al.
Veröffentlicht: (2026)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
von: Yu, Li, et al.
Veröffentlicht: (2025)
von: Yu, Li, et al.
Veröffentlicht: (2025)
Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices
von: Cohen, Nathaniel, et al.
Veröffentlicht: (2024)
von: Cohen, Nathaniel, et al.
Veröffentlicht: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
von: Shen, Hao, et al.
Veröffentlicht: (2024)
von: Shen, Hao, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
von: Cao, Tri, et al.
Veröffentlicht: (2026)
von: Cao, Tri, et al.
Veröffentlicht: (2026)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
von: Wu, Peng, et al.
Veröffentlicht: (2024)
von: Wu, Peng, et al.
Veröffentlicht: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026)
von: Gu, Xin, et al.
Veröffentlicht: (2026)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
von: Li, Xinhao, et al.
Veröffentlicht: (2025)
von: Li, Xinhao, et al.
Veröffentlicht: (2025)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
Transformer RGBT Tracking with Spatio-Temporal Multimodal Tokens
von: Sun, Dengdi, et al.
Veröffentlicht: (2024)
von: Sun, Dengdi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
von: Zhang, Yan, et al.
Veröffentlicht: (2025) -
Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding
von: Zhang, Yan, et al.
Veröffentlicht: (2026) -
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026) -
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2025) -
Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
von: Xie, Enze, et al.
Veröffentlicht: (2024)