DocSynthv2: A Practical Autoregressive Modeling for Document Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Biswas, Sanket, Jain, Rajiv, Morariu, Vlad I., Gu, Jiuxiang, Mathur, Puneet, Wigington, Curtis, Sun, Tong, Lladós, Josep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
by: Tiwari, Adarsh, et al.
Published: (2024)
by: Tiwari, Adarsh, et al.
Published: (2024)
DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
by: Suri, Manan, et al.
Published: (2024)
by: Suri, Manan, et al.
Published: (2024)
LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach
by: Pilligua, Maria, et al.
Published: (2024)
by: Pilligua, Maria, et al.
Published: (2024)
Handheld Video Document Scanning: A Robust On-Device Model for Multi-Page Document Scanning
by: Wigington, Curtis
Published: (2024)
by: Wigington, Curtis
Published: (2024)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
by: Biescas, Nil, et al.
Published: (2024)
by: Biescas, Nil, et al.
Published: (2024)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
by: Bakkali, Souhail, et al.
Published: (2023)
by: Bakkali, Souhail, et al.
Published: (2023)
AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
by: Lin, Jiawei, et al.
Published: (2026)
by: Lin, Jiawei, et al.
Published: (2026)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
by: Zhang, Jianyi, et al.
Published: (2024)
by: Zhang, Jianyi, et al.
Published: (2024)
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)
by: Purkayastha, Kunal, et al.
Published: (2026)
FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
by: Lin, Zihao, et al.
Published: (2026)
by: Lin, Zihao, et al.
Published: (2026)
Customization Assistant for Text-to-image Generation
by: Zhou, Yufan, et al.
Published: (2023)
by: Zhou, Yufan, et al.
Published: (2023)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
Text-Conditioned Background Generation for Editable Multi-Layer Documents
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
by: Molina, Adrià, et al.
Published: (2024)
by: Molina, Adrià, et al.
Published: (2024)
ImageFolder: Autoregressive Image Generation with Folded Tokens
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
by: Banerjee, Ayan, et al.
Published: (2025)
by: Banerjee, Ayan, et al.
Published: (2025)
CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models
by: Banerjee, Ayan, et al.
Published: (2025)
by: Banerjee, Ayan, et al.
Published: (2025)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing
by: Rodríguez, Adrià Molina, et al.
Published: (2025)
by: Rodríguez, Adrià Molina, et al.
Published: (2025)
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
XQ-GAN: An Open-source Image Tokenization Framework for Autoregressive Generation
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
DocParseNet: Advanced Semantic Segmentation and OCR Embeddings for Efficient Scanned Document Annotation
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
by: Mohammadshirazi, Ahmad, et al.
Published: (2024)
Recurrent Few-Shot model for Document Verification
by: Talarmain, Maxime, et al.
Published: (2024)
by: Talarmain, Maxime, et al.
Published: (2024)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
by: Zhao, Haozhe, et al.
Published: (2025)
by: Zhao, Haozhe, et al.
Published: (2025)
SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
by: Chen, Jian, et al.
Published: (2024)
by: Chen, Jian, et al.
Published: (2024)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
by: Mondal, Anindya, et al.
Published: (2025)
by: Mondal, Anindya, et al.
Published: (2025)
Agentic Design Review System
by: Nag, Sayan, et al.
Published: (2025)
by: Nag, Sayan, et al.
Published: (2025)
GAN-based Content-Conditioned Generation of Handwritten Musical Symbols
by: Asbert, Gerard, et al.
Published: (2025)
by: Asbert, Gerard, et al.
Published: (2025)
DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning
by: Gu, Wenhao, et al.
Published: (2025)
by: Gu, Wenhao, et al.
Published: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
by: Zhou, Yufan, et al.
Published: (2024)
by: Zhou, Yufan, et al.
Published: (2024)
La cadena global de valor en la industria electrónica
by: Josep Lladós Masllorens
Published: (2018)
by: Josep Lladós Masllorens
Published: (2018)
Similar Items
-
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
by: Tiwari, Adarsh, et al.
Published: (2024) -
DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
by: Suri, Manan, et al.
Published: (2024) -
LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach
by: Pilligua, Maria, et al.
Published: (2024) -
Handheld Video Document Scanning: A Robust On-Device Model for Multi-Page Document Scanning
by: Wigington, Curtis
Published: (2024) -
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)