DOTA: Deformable Optimized Transformer Architecture for End-to-End Text Recognition with Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Nithisopa, Naphat, Panboonyuen, Teerapong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Foundations and Architectures of Artificial Intelligence for Motor Insurance
by: Panboonyuen, Teerapong
Published: (2026)
by: Panboonyuen, Teerapong
Published: (2026)
REG: Refined Generalized Focal Loss for Road Asset Detection on Thai Highways Using Vision-Based Detection and Segmentation Models
by: Panboonyuen, Teerapong
Published: (2024)
by: Panboonyuen, Teerapong
Published: (2024)
MARS: Mask Attention Refinement with Sequential Quadtree Nodes for Car Damage Instance Segmentation
by: Panboonyuen, Teerapong, et al.
Published: (2023)
by: Panboonyuen, Teerapong, et al.
Published: (2023)
ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation
by: Panboonyuen, Teerapong
Published: (2025)
by: Panboonyuen, Teerapong
Published: (2025)
KAO: Kernel-Adaptive Optimization in Diffusion for Satellite Image
by: Panboonyuen, Teerapong
Published: (2025)
by: Panboonyuen, Teerapong
Published: (2025)
SLICK: Selective Localization and Instance Calibration for Knowledge-Enhanced Car Damage Segmentation in Automotive Insurance
by: Panboonyuen, Teerapong
Published: (2025)
by: Panboonyuen, Teerapong
Published: (2025)
Seeing Isn't Always Believing: Analysis of Grad-CAM Faithfulness and Localization Reliability in Lung Cancer CT Classification
by: Panboonyuen, Teerapong
Published: (2026)
by: Panboonyuen, Teerapong
Published: (2026)
HOMEY: Heuristic Object Masking with Enhanced YOLO for Property Insurance Risk Detection
by: Panboonyuen, Teerapong
Published: (2026)
by: Panboonyuen, Teerapong
Published: (2026)
HERS: Hidden-Pattern Expert Learning for Risk-Specific Vehicle Damage Adaptation in Diffusion Models
by: Panboonyuen, Teerapong
Published: (2026)
by: Panboonyuen, Teerapong
Published: (2026)
End-to-End Optimized Image Compression with the Frequency-Oriented Transform
by: Zhang, Yuefeng, et al.
Published: (2024)
by: Zhang, Yuefeng, et al.
Published: (2024)
LP-LLM: End-to-End Real-World Degraded License Plate Text Recognition via Large Multimodal Models
by: Gong, Haoyan, et al.
Published: (2026)
by: Gong, Haoyan, et al.
Published: (2026)
An End-to-End, Segmentation-Free, Arabic Handwritten Recognition Model on KHATT
by: Aabed, Sondos, et al.
Published: (2024)
by: Aabed, Sondos, et al.
Published: (2024)
Augmenting End-to-End Steering Angle Prediction with CAN Bus Data
by: Singh, Amit
Published: (2023)
by: Singh, Amit
Published: (2023)
SGTR+: End-to-end Scene Graph Generation with Transformer
by: Li, Rongjie, et al.
Published: (2024)
by: Li, Rongjie, et al.
Published: (2024)
USAD: End-to-End Human Activity Recognition via Diffusion Model with Spatiotemporal Attention
by: Xiao, Hang, et al.
Published: (2025)
by: Xiao, Hang, et al.
Published: (2025)
End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction
by: Li, Xi, et al.
Published: (2025)
by: Li, Xi, et al.
Published: (2025)
An End-to-End Deep Learning Generative Framework for Refinable Shape Matching and Generation
by: Kalaie, Soodeh, et al.
Published: (2024)
by: Kalaie, Soodeh, et al.
Published: (2024)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
An End-to-End Two-Stream Network Based on RGB Flow and Representation Flow for Human Action Recognition
by: Lai, Song-Jiang, et al.
Published: (2024)
by: Lai, Song-Jiang, et al.
Published: (2024)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
by: Molino, Daniele, et al.
Published: (2026)
by: Molino, Daniele, et al.
Published: (2026)
End-to-End Human Instance Matting
by: Liu, Qinglin, et al.
Published: (2024)
by: Liu, Qinglin, et al.
Published: (2024)
MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer For Efficient Medical Image Segmentation
by: Perera, Shehan, et al.
Published: (2024)
by: Perera, Shehan, et al.
Published: (2024)
TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
by: Qin, Chunxia, et al.
Published: (2026)
by: Qin, Chunxia, et al.
Published: (2026)
JointRF: End-to-End Joint Optimization for Dynamic Neural Radiance Field Representation and Compression
by: Zheng, Zihan, et al.
Published: (2024)
by: Zheng, Zihan, et al.
Published: (2024)
Guiding Attention in End-to-End Driving Models
by: Porres, Diego, et al.
Published: (2024)
by: Porres, Diego, et al.
Published: (2024)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
by: Ahn, Young Jin, et al.
Published: (2024)
by: Ahn, Young Jin, et al.
Published: (2024)
RAG-HAR: Retrieval Augmented Generation-based Human Activity Recognition
by: Sivaroopan, Nirhoshan, et al.
Published: (2025)
by: Sivaroopan, Nirhoshan, et al.
Published: (2025)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction
by: Halinkovic, Matej, et al.
Published: (2026)
by: Halinkovic, Matej, et al.
Published: (2026)
AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
by: Wang, Zile, et al.
Published: (2025)
by: Wang, Zile, et al.
Published: (2025)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
by: Du, Yongkun, et al.
Published: (2025)
by: Du, Yongkun, et al.
Published: (2025)
iPad: Iterative Proposal-centric End-to-End Autonomous Driving
by: Guo, Ke, et al.
Published: (2025)
by: Guo, Ke, et al.
Published: (2025)
Fose: Fusion of One-Step Diffusion and End-to-End Network for Pansharpening
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
Beyond Hungarian: Match-Free Supervision for End-to-End Object Detection
by: Qiu, Shoumeng, et al.
Published: (2026)
by: Qiu, Shoumeng, et al.
Published: (2026)
Hint-AD: Holistically Aligned Interpretability in End-to-End Autonomous Driving
by: Ding, Kairui, et al.
Published: (2024)
by: Ding, Kairui, et al.
Published: (2024)
Adversarial Flow Matching for Imperceptible Attacks on End-to-End Autonomous Driving
by: Zeng, Xinyu, et al.
Published: (2026)
by: Zeng, Xinyu, et al.
Published: (2026)
Similar Items
-
Foundations and Architectures of Artificial Intelligence for Motor Insurance
by: Panboonyuen, Teerapong
Published: (2026) -
REG: Refined Generalized Focal Loss for Road Asset Detection on Thai Highways Using Vision-Based Detection and Segmentation Models
by: Panboonyuen, Teerapong
Published: (2024) -
MARS: Mask Attention Refinement with Sequential Quadtree Nodes for Car Damage Instance Segmentation
by: Panboonyuen, Teerapong, et al.
Published: (2023) -
ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation
by: Panboonyuen, Teerapong
Published: (2025) -
KAO: Kernel-Adaptive Optimization in Diffusion for Satellite Image
by: Panboonyuen, Teerapong
Published: (2025)