Saved in:
| Main Authors: | Fei, Yulin, Gao, Yuhui, Xian, Xingyuan, Zhang, Xiaojin, Wu, Tao, Chen, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.20613 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
by: Zhang, Yulin, et al.
Published: (2026)
by: Zhang, Yulin, et al.
Published: (2026)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025)
by: Chen, Song, et al.
Published: (2025)
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024)
by: Liu, Yuanxin, et al.
Published: (2024)
CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
by: Li, Chenglin, et al.
Published: (2024)
by: Li, Chenglin, et al.
Published: (2024)
GOGS: High-Fidelity Geometry and Relighting for Glossy Objects via Gaussian Surfels
by: Yang, Xingyuan, et al.
Published: (2025)
by: Yang, Xingyuan, et al.
Published: (2025)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
by: Zhang, Pu, et al.
Published: (2025)
by: Zhang, Pu, et al.
Published: (2025)
OmniOCR: Generalist OCR for Ethnic Minority Languages
by: Liu, Bonan, et al.
Published: (2026)
by: Liu, Bonan, et al.
Published: (2026)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
The Dawn of Video Generation: Preliminary Explorations with SORA-like Models
by: Zeng, Ailing, et al.
Published: (2024)
by: Zeng, Ailing, et al.
Published: (2024)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
by: Fei, Hao, et al.
Published: (2023)
by: Fei, Hao, et al.
Published: (2023)
OCR-Agent: Agentic OCR with Capability and Memory Reflection
by: Wen, Shimin, et al.
Published: (2026)
by: Wen, Shimin, et al.
Published: (2026)
An Empirical Study of Scaling Law for OCR
by: Rang, Miao, et al.
Published: (2023)
by: Rang, Miao, et al.
Published: (2023)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
SCBench: A Sports Commentary Benchmark for Video LLMs
by: Ge, Kuangzhi, et al.
Published: (2024)
by: Ge, Kuangzhi, et al.
Published: (2024)
Agentar-Fin-OCR
by: Qian, Siyi, et al.
Published: (2026)
by: Qian, Siyi, et al.
Published: (2026)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
by: Zhang, Junyuan, et al.
Published: (2024)
by: Zhang, Junyuan, et al.
Published: (2024)
When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation
by: Sun, Lin, et al.
Published: (2026)
by: Sun, Lin, et al.
Published: (2026)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
by: Wang, Hui, et al.
Published: (2026)
by: Wang, Hui, et al.
Published: (2026)
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
by: Wu, Yuhui, et al.
Published: (2025)
by: Wu, Yuhui, et al.
Published: (2025)
PaddleOCR 3.0 Technical Report
by: Cui, Cheng, et al.
Published: (2025)
by: Cui, Cheng, et al.
Published: (2025)
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
by: Zhang, Yulin, et al.
Published: (2025)
by: Zhang, Yulin, et al.
Published: (2025)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
DRFusion: Drift-Resilient Temporally Consistent Infrared-Visible Video Fusion
by: Li, Xingyuan, et al.
Published: (2026)
by: Li, Xingyuan, et al.
Published: (2026)
DeepSeek-OCR: Contexts Optical Compression
by: Wei, Haoran, et al.
Published: (2025)
by: Wei, Haoran, et al.
Published: (2025)
Anomize: Better Open Vocabulary Video Anomaly Detection
by: Li, Fei, et al.
Published: (2025)
by: Li, Fei, et al.
Published: (2025)
KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing
by: Jiang, Siyu, et al.
Published: (2026)
by: Jiang, Siyu, et al.
Published: (2026)
Context-Independent OCR with Multimodal LLMs: Effects of Image Resolution and Visual Complexity
by: Inoue, Kotaro
Published: (2025)
by: Inoue, Kotaro
Published: (2025)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
by: Zhang, Yulong
Published: (2025)
by: Zhang, Yulong
Published: (2025)
A Closer Look at Edema Area Segmentation in SD-OCT Images Using Adversarial Framework
by: Tao, Yuhui, et al.
Published: (2025)
by: Tao, Yuhui, et al.
Published: (2025)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
by: Wu, Xian, et al.
Published: (2025)
by: Wu, Xian, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Prototypical Contrastive Learning-based CLIP Fine-tuning for Object Re-identification
by: Li, Jiachen, et al.
Published: (2023)
by: Li, Jiachen, et al.
Published: (2023)
Tuning-free Universally-Supervised Semantic Segmentation
by: Yang, Xiaobo, et al.
Published: (2024)
by: Yang, Xiaobo, et al.
Published: (2024)
Unleashing the Potential of Pre-Trained Diffusion Models for Generalizable Person Re-Identification
by: Li, Jiachen, et al.
Published: (2025)
by: Li, Jiachen, et al.
Published: (2025)
Similar Items
-
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025) -
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
by: Zhang, Yulin, et al.
Published: (2026) -
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025) -
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
by: Chen, Song, et al.
Published: (2025) -
TempCompass: Do Video LLMs Really Understand Videos?
by: Liu, Yuanxin, et al.
Published: (2024)