Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bi, Tianci, Zhang, Xiaoyi, Zhang, Zhizheng, Xie, Wenxuan, Lan, Cuiling, Lu, Yan, Zheng, Nanning |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
von: Bi, Tianci, et al.
Veröffentlicht: (2025)
von: Bi, Tianci, et al.
Veröffentlicht: (2025)
Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement
von: Yang, Tao, et al.
Veröffentlicht: (2024)
von: Yang, Tao, et al.
Veröffentlicht: (2024)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Slot-VLM: SlowFast Slots for Video-Language Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
von: Guan, Tongkun, et al.
Veröffentlicht: (2023)
von: Guan, Tongkun, et al.
Veröffentlicht: (2023)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images
von: Cui, Ziteng, et al.
Veröffentlicht: (2024)
von: Cui, Ziteng, et al.
Veröffentlicht: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
von: Fhima, Jonathan, et al.
Veröffentlicht: (2024)
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark
von: Cui, Ziteng, et al.
Veröffentlicht: (2025)
von: Cui, Ziteng, et al.
Veröffentlicht: (2025)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Self-Supervised Pre-training with Symmetric Superimposition Modeling for Scene Text Recognition
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
von: Gao, Zuan, et al.
Veröffentlicht: (2024)
Deciphering Functions of Neurons in Vision-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2025)
Multitwine: Multi-Object Compositing with Text and Layout Control
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2025)
von: Tarrés, Gemma Canet, et al.
Veröffentlicht: (2025)
GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval
von: Liu, Yating, et al.
Veröffentlicht: (2025)
von: Liu, Yating, et al.
Veröffentlicht: (2025)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
von: Jiang, Yue, et al.
Veröffentlicht: (2026)
von: Jiang, Yue, et al.
Veröffentlicht: (2026)
TextMamba: Scene Text Detector with Mamba
von: Zhao, Qiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Qiyan, et al.
Veröffentlicht: (2025)
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
von: Lyu, Jiahao, et al.
Veröffentlicht: (2024)
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
von: Zhang, Zijian, et al.
Veröffentlicht: (2024)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training
von: Jiang, Zhouqiang, et al.
Veröffentlicht: (2024)
von: Jiang, Zhouqiang, et al.
Veröffentlicht: (2024)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
von: Zhang, Xike, et al.
Veröffentlicht: (2026)
von: Zhang, Xike, et al.
Veröffentlicht: (2026)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
von: Jin, Xiaojie, et al.
Veröffentlicht: (2023)
von: Jin, Xiaojie, et al.
Veröffentlicht: (2023)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024)
von: Ye, Wei, et al.
Veröffentlicht: (2024)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
von: Wu, Shihan, et al.
Veröffentlicht: (2024)
von: Wu, Shihan, et al.
Veröffentlicht: (2024)
Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learning
von: Wang, Huiyi, et al.
Veröffentlicht: (2024)
von: Wang, Huiyi, et al.
Veröffentlicht: (2024)
Generating Animated Layouts as Structured Text Representations
von: Shin, Yeonsang, et al.
Veröffentlicht: (2025)
von: Shin, Yeonsang, et al.
Veröffentlicht: (2025)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
Text-Pass Filter: An Efficient Scene Text Detector
von: Yang, Chuang, et al.
Veröffentlicht: (2026)
von: Yang, Chuang, et al.
Veröffentlicht: (2026)
ODM: A Text-Image Further Alignment Pre-training Approach for Scene Text Detection and Spotting
von: Duan, Chen, et al.
Veröffentlicht: (2024)
von: Duan, Chen, et al.
Veröffentlicht: (2024)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
von: Shen, Guibao, et al.
Veröffentlicht: (2024)
von: Shen, Guibao, et al.
Veröffentlicht: (2024)
TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting
von: Xie, Liangbin, et al.
Veröffentlicht: (2025)
von: Xie, Liangbin, et al.
Veröffentlicht: (2025)
Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
von: Jin, Jiayun, et al.
Veröffentlicht: (2026)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
von: Zhang, Beichen, et al.
Veröffentlicht: (2024)
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
von: SR, Nikitha, et al.
Veröffentlicht: (2024)
von: SR, Nikitha, et al.
Veröffentlicht: (2024)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
von: Ma, Shuailei, et al.
Veröffentlicht: (2023)
von: Ma, Shuailei, et al.
Veröffentlicht: (2023)
Leveraging Anchor-based LiDAR 3D Object Detection via Point Assisted Sample Selection
von: Chen, Shitao, et al.
Veröffentlicht: (2024)
von: Chen, Shitao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
von: Bi, Tianci, et al.
Veröffentlicht: (2025) -
Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement
von: Yang, Tao, et al.
Veröffentlicht: (2024) -
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023) -
Slot-VLM: SlowFast Slots for Video-Language Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024) -
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
von: Mao, Zhiming, et al.
Veröffentlicht: (2024)