AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Jiawei, Zhu, Wanrong, Morariu, Vlad I, Tensmeyer, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
Text-Conditioned Background Generation for Editable Multi-Layer Documents
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
by: Lin, Zihao, et al.
Published: (2026)
by: Lin, Zihao, et al.
Published: (2026)
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024)
by: Biswas, Sanket, et al.
Published: (2024)
DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
by: Suri, Manan, et al.
Published: (2024)
by: Suri, Manan, et al.
Published: (2024)
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding
by: Al-Homoud, Haneen, et al.
Published: (2025)
by: Al-Homoud, Haneen, et al.
Published: (2025)
TutoAI: A Cross-domain Framework for AI-assisted Mixed-media Tutorial Creation on Physical Tasks
by: Chen, Yuexi, et al.
Published: (2024)
by: Chen, Yuexi, et al.
Published: (2024)
Guided Mini Website Simulator (HTML & CSS): A Didactic Environment for Introductory Web Design
by: Capano, Domenico, et al.
Published: (2026)
by: Capano, Domenico, et al.
Published: (2026)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
by: Fallah, Forouzan, et al.
Published: (2025)
by: Fallah, Forouzan, et al.
Published: (2025)
El gran libro de HTML5, CSS3 y JavaScript / J.D. Gauchat
by: Gauchat, J.D
Published: (2019)
by: Gauchat, J.D
Published: (2019)
ScaleDoc: Scaling LLM-based Predicates over Large Document Collections
by: Zhang, Hengrui, et al.
Published: (2025)
by: Zhang, Hengrui, et al.
Published: (2025)
AniDoc: Animation Creation Made Easier
by: Meng, Yihao, et al.
Published: (2024)
by: Meng, Yihao, et al.
Published: (2024)
Making It Work for Everyone: HTML5 and CSS Level 3 for Responsive, Accessible Design on Your Library's Web Site
by: Baker, Stewart C.
Published: (2014)
by: Baker, Stewart C.
Published: (2014)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction
by: Sun, Ting, et al.
Published: (2025)
by: Sun, Ting, et al.
Published: (2025)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
by: Chen, Ketong, et al.
Published: (2025)
by: Chen, Ketong, et al.
Published: (2025)
Online Statistical Inference of Constrained Stochastic Optimization via Random Scaling
by: Du, Xinchen, et al.
Published: (2025)
by: Du, Xinchen, et al.
Published: (2025)
Agentic Design Review System
by: Nag, Sayan, et al.
Published: (2025)
by: Nag, Sayan, et al.
Published: (2025)
DocDeshadower: Frequency-Aware Transformer for Document Shadow Removal
by: Zhou, Ziyang, et al.
Published: (2023)
by: Zhou, Ziyang, et al.
Published: (2023)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
CogDoc: Towards Unified thinking in Documents
by: Xu, Qixin, et al.
Published: (2025)
by: Xu, Qixin, et al.
Published: (2025)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
Structure of CSS and CSS-T Quantum Codes
by: Berardini, Elena, et al.
Published: (2023)
by: Berardini, Elena, et al.
Published: (2023)
AnySR: Realizing Image Super-Resolution as Any-Scale, Any-Resource
by: Zhan, Wengyi, et al.
Published: (2024)
by: Zhan, Wengyi, et al.
Published: (2024)
DocCGen: Document-based Controlled Code Generation
by: Pimparkhede, Sameer, et al.
Published: (2024)
by: Pimparkhede, Sameer, et al.
Published: (2024)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
by: Feng, Xiang, et al.
Published: (2026)
by: Feng, Xiang, et al.
Published: (2026)
Context-Aware or Context-Insensitive? Assessing LLMs' Performance in Document-Level Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
by: Huang, Jiaxi, et al.
Published: (2025)
by: Huang, Jiaxi, et al.
Published: (2025)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
by: Hu, Xuran, et al.
Published: (2026)
by: Hu, Xuran, et al.
Published: (2026)
Doc To The Future: Infomorphs for Interactive, Multimodal Document Transformation and Generation
by: Kumaravel, Balasaravanan Thoravi
Published: (2025)
by: Kumaravel, Balasaravanan Thoravi
Published: (2025)
DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering
by: Lin, Teng, et al.
Published: (2026)
by: Lin, Teng, et al.
Published: (2026)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
by: Ouyang, Linke, et al.
Published: (2024)
by: Ouyang, Linke, et al.
Published: (2024)
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
by: wu, Weihao, et al.
Published: (2025)
by: wu, Weihao, et al.
Published: (2025)
DocTER: Evaluating Document-based Knowledge Editing
by: Wu, Suhang, et al.
Published: (2023)
by: Wu, Suhang, et al.
Published: (2023)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
by: Li, Jierui, et al.
Published: (2023)
by: Li, Jierui, et al.
Published: (2023)
PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
by: Jaisankar, Vijay, et al.
Published: (2024)
by: Jaisankar, Vijay, et al.
Published: (2024)
AnyTrans: Translate AnyText in the Image with Large Scale Models
by: Qian, Zhipeng, et al.
Published: (2024)
by: Qian, Zhipeng, et al.
Published: (2024)
TreeCSS: An Efficient Framework for Vertical Federated Learning
by: Zhang, Qinbo, et al.
Published: (2024)
by: Zhang, Qinbo, et al.
Published: (2024)
Similar Items
-
FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
by: Jiang, Yue, et al.
Published: (2024) -
Text-Conditioned Background Generation for Editable Multi-Layer Documents
by: Kang, Taewon, et al.
Published: (2025) -
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
by: Lin, Zihao, et al.
Published: (2026) -
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024) -
DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
by: Suri, Manan, et al.
Published: (2024)