VimTS: A Unified Video and Image Text Spotter for Enhancing the Cross-domain Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuliang, Huang, Mingxin, Yan, Hao, Deng, Linger, Wu, Weijia, Lu, Hao, Shen, Chunhua, Jin, Lianwen, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023)
by: Deng, Linger, et al.
Published: (2023)
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023)
by: Peng, Dezhi, et al.
Published: (2023)
GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving
by: Deng, Linger, et al.
Published: (2026)
by: Deng, Linger, et al.
Published: (2026)
LEGO: Self-Supervised Representation Learning for Scene Text Images
by: Ren, Yujin, et al.
Published: (2024)
by: Ren, Yujin, et al.
Published: (2024)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
by: Zhai, Yukun, et al.
Published: (2023)
by: Zhai, Yukun, et al.
Published: (2023)
Omni-IML: Towards Unified Image Manipulation Localization
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
TiCLS : Tightly Coupled Language Text Spotter
by: Jang, Leeje, et al.
Published: (2026)
by: Jang, Leeje, et al.
Published: (2026)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
by: Liu, Yuliang, et al.
Published: (2023)
by: Liu, Yuliang, et al.
Published: (2023)
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Apresentação
by: Carlos Linger
Published: (2016)
by: Carlos Linger
Published: (2016)
Vim4Path: Self-Supervised Vision Mamba for Histopathology Images
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
Theorem-Validated Reverse Chain-of-Thought Problem Generation for Geometric Reasoning
by: Deng, Linger, et al.
Published: (2024)
by: Deng, Linger, et al.
Published: (2024)
UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation
by: Zhong, Siru, et al.
Published: (2024)
by: Zhong, Siru, et al.
Published: (2024)
Arbitrary Reading Order Scene Text Spotter with Local Semantics Guidance
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
by: Li, Liyang, et al.
Published: (2026)
by: Li, Liyang, et al.
Published: (2026)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
FiCoTS: Fine-to-Coarse LLM-Enhanced Hierarchical Cross-Modality Interaction for Time Series Forecasting
by: Lyu, Yafei, et al.
Published: (2025)
by: Lyu, Yafei, et al.
Published: (2025)
TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering
by: Zhu, Hanshen, et al.
Published: (2026)
by: Zhu, Hanshen, et al.
Published: (2026)
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
by: Fu, Ling, et al.
Published: (2024)
by: Fu, Ling, et al.
Published: (2024)
Puzzle Pieces Picker: Deciphering Ancient Chinese Characters with Radical Reconstruction
by: Wang, Pengjie, et al.
Published: (2024)
by: Wang, Pengjie, et al.
Published: (2024)
Deciphering Oracle Bone Language with Diffusion Models
by: Guan, Haisu, et al.
Published: (2024)
by: Guan, Haisu, et al.
Published: (2024)
BadVim: Unveiling Backdoor Threats in Visual State Space Model
by: Lee, Cheng-Yi, et al.
Published: (2024)
by: Lee, Cheng-Yi, et al.
Published: (2024)
DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
EdgeSpotter: Multi-Scale Dense Text Spotting for Industrial Panel Monitoring
by: Fu, Changhong, et al.
Published: (2025)
by: Fu, Changhong, et al.
Published: (2025)
Privacy-Preserving Biometric Verification with Handwritten Random Digit String
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
by: Li, Zhang, et al.
Published: (2025)
by: Li, Zhang, et al.
Published: (2025)
DualTSR: Unified Dual-Diffusion Transformer for Scene Text Image Super-Resolution
by: Niu, Axi, et al.
Published: (2026)
by: Niu, Axi, et al.
Published: (2026)
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models
by: Wang, Wen, et al.
Published: (2023)
by: Wang, Wen, et al.
Published: (2023)
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
by: Xie, Yu, et al.
Published: (2024)
by: Xie, Yu, et al.
Published: (2024)
Hierarchical Side-Tuning for Vision Transformers
by: Lin, Weifeng, et al.
Published: (2023)
by: Lin, Weifeng, et al.
Published: (2023)
An open dataset for the evolution of oracle bone characters: EVOBC
by: Guan, Haisu, et al.
Published: (2024)
by: Guan, Haisu, et al.
Published: (2024)
HotSpotter - Patterned Species Instance Recognition
by: Crall, Jonathan P., et al.
Published: (2025)
by: Crall, Jonathan P., et al.
Published: (2025)
Revisiting Tampered Scene Text Detection in the Era of Generative AI
by: Qu, Chenfan, et al.
Published: (2024)
by: Qu, Chenfan, et al.
Published: (2024)
DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
by: Wu, Weijia, et al.
Published: (2023)
by: Wu, Weijia, et al.
Published: (2023)
Phosphorus Limitation Constrains Global Forest Productivity Directly and Indirectly via Forest Community Structural Attributes: Meta‐Analysis
by: Ewuketu Linger, et al.
Published: (2025)
by: Ewuketu Linger, et al.
Published: (2025)
Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
by: Zhang, Peirong, et al.
Published: (2024)
by: Zhang, Peirong, et al.
Published: (2024)
Similar Items
-
Progressive Evolution from Single-Point to Polygon for Scene Text
by: Deng, Linger, et al.
Published: (2023) -
SwinTextSpotter v2: Towards Better Synergy for Scene Text Spotting
by: Huang, Mingxin, et al.
Published: (2024) -
Bridging the Gap Between End-to-End and Two-Step Text Spotting
by: Huang, Mingxin, et al.
Published: (2024) -
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024) -
ViTEraser: Harnessing the Power of Vision Transformers for Scene Text Removal with SegMIM Pretraining
by: Peng, Dezhi, et al.
Published: (2023)