Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Maracani, Andrea, Ozkan, Savas, Cho, Sijun, Kim, Hyowon, Noh, Eunchung, Min, Jeongwon, Min, Cho Jung, Park, Dookun, Ozay, Mete |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient and Accurate Scene Text Recognition with Cascaded-Transformers
by: Ozkan, Savas, et al.
Published: (2025)
by: Ozkan, Savas, et al.
Published: (2025)
Decoding Text Spans for Efficient and Accurate Named-Entity Recognition
by: Maracani, Andrea, et al.
Published: (2026)
by: Maracani, Andrea, et al.
Published: (2026)
Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
by: Zhu, Junyi, et al.
Published: (2025)
by: Zhu, Junyi, et al.
Published: (2025)
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
by: Bourouis, Ahmed, et al.
Published: (2026)
by: Bourouis, Ahmed, et al.
Published: (2026)
Guided Model Merging for Hybrid Data Learning: Leveraging Centralized Data to Refine Decentralized Models
by: Zhu, Junyi, et al.
Published: (2025)
by: Zhu, Junyi, et al.
Published: (2025)
Efficient 3D Full-Body Motion Generation from Sparse Tracking Inputs with Temporal Windows
by: Angelis, Georgios Fotios, et al.
Published: (2025)
by: Angelis, Georgios Fotios, et al.
Published: (2025)
Responsible Federated LLMs via Safety Filtering and Constitutional AI
by: Noh, Eunchung, et al.
Published: (2025)
by: Noh, Eunchung, et al.
Published: (2025)
Mem-MLP: Real-Time 3D Human Motion Generation from Sparse Inputs
by: Mutlu, Sinan, et al.
Published: (2025)
by: Mutlu, Sinan, et al.
Published: (2025)
HOP to the Next Tasks and Domains for Continual Learning in NLP
by: Michieli, Umberto, et al.
Published: (2024)
by: Michieli, Umberto, et al.
Published: (2024)
Self-Knowledge Distillation for Learning Ambiguity
by: Park, Hancheol, et al.
Published: (2024)
by: Park, Hancheol, et al.
Published: (2024)
Object-conditioned Bag of Instances for Few-Shot Personalized Instance Recognition
by: Michieli, Umberto, et al.
Published: (2024)
by: Michieli, Umberto, et al.
Published: (2024)
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
FFT-based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
by: Camuffo, Elena, et al.
Published: (2024)
by: Camuffo, Elena, et al.
Published: (2024)
A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
by: Fish, Edward, et al.
Published: (2023)
by: Fish, Edward, et al.
Published: (2023)
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
by: Chen, Zhuohao, et al.
Published: (2026)
by: Chen, Zhuohao, et al.
Published: (2026)
Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
by: Barbato, Francesco, et al.
Published: (2024)
by: Barbato, Francesco, et al.
Published: (2024)
Dataset Distillation for Super-Resolution without Class Labels and Pre-trained Models
by: Cho, Sunwoo, et al.
Published: (2025)
by: Cho, Sunwoo, et al.
Published: (2025)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
Trust And Balance: Few Trusted Samples Pseudo-Labeling and Temperature Scaled Loss for Effective Source-Free Unsupervised Domain Adaptation
by: Maracani, Andrea, et al.
Published: (2024)
by: Maracani, Andrea, et al.
Published: (2024)
Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-Resolution
by: Park, Karam, et al.
Published: (2025)
by: Park, Karam, et al.
Published: (2025)
Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
by: Nair, Sathvik, et al.
Published: (2026)
by: Nair, Sathvik, et al.
Published: (2026)
Fast and Accurate Neural Rendering Using Semi-Gradients
by: Cho, In-Young, et al.
Published: (2024)
by: Cho, In-Young, et al.
Published: (2024)
Long-tailed Adversarial Training with Self-Distillation
by: Cho, Seungju, et al.
Published: (2025)
by: Cho, Seungju, et al.
Published: (2025)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
Improving Factual Error Correction for Abstractive Summarization via Data Distillation and Conditional-generation Cloze
by: Li, Yiyang, et al.
Published: (2024)
by: Li, Yiyang, et al.
Published: (2024)
Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search
by: Paramonov, Kirill, et al.
Published: (2024)
by: Paramonov, Kirill, et al.
Published: (2024)
Burst Image Super-Resolution with Base Frame Selection
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025)
by: Lee, Junhong, et al.
Published: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
DALDA: Data Augmentation Leveraging Diffusion Model and LLM with Adaptive Guidance Scaling
by: Jung, Kyuheon, et al.
Published: (2024)
by: Jung, Kyuheon, et al.
Published: (2024)
Feature-Space Generative Models for One-Shot Class-Incremental Learning
by: Foster, Jack, et al.
Published: (2026)
by: Foster, Jack, et al.
Published: (2026)
Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes
by: Choi, JunYong, et al.
Published: (2025)
by: Choi, JunYong, et al.
Published: (2025)
ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Why Latent Actions Fail, and How to Prevent It
by: Lee, Jung Min, et al.
Published: (2026)
by: Lee, Jung Min, et al.
Published: (2026)
Efficient Compositional Multi-tasking for On-device Large Language Models
by: Bohdal, Ondrej, et al.
Published: (2025)
by: Bohdal, Ondrej, et al.
Published: (2025)
Dynamic Guidance Adversarial Distillation with Enhanced Teacher Knowledge
by: Park, Hyejin, et al.
Published: (2024)
by: Park, Hyejin, et al.
Published: (2024)
iConFormer: Dynamic Parameter-Efficient Tuning with Input-Conditioned Adaptation
by: Jo, Hayeon, et al.
Published: (2024)
by: Jo, Hayeon, et al.
Published: (2024)
Scene Abstraction for Lexical Semantics: Structured Representations of Situated Meaning
by: Cho, Yejin, et al.
Published: (2026)
by: Cho, Yejin, et al.
Published: (2026)
Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
by: Camuffo, Elena, et al.
Published: (2024)
by: Camuffo, Elena, et al.
Published: (2024)
Similar Items
-
Efficient and Accurate Scene Text Recognition with Cascaded-Transformers
by: Ozkan, Savas, et al.
Published: (2025) -
Decoding Text Spans for Efficient and Accurate Named-Entity Recognition
by: Maracani, Andrea, et al.
Published: (2026) -
Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER
by: Zhu, Junyi, et al.
Published: (2025) -
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
by: Bourouis, Ahmed, et al.
Published: (2026) -
Guided Model Merging for Hybrid Data Learning: Leveraging Centralized Data to Refine Decentralized Models
by: Zhu, Junyi, et al.
Published: (2025)