Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
Fuente:
arXiv
Saved in:
| Main Author: | Duan, Changxu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation
by: Duan, Changxu
Published: (2025)
by: Duan, Changxu
Published: (2025)
Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
by: Shi, Xiang, et al.
Published: (2024)
by: Shi, Xiang, et al.
Published: (2024)
Differentiating Emigration from Return Migration of Scholars Using Name-Based Nationality Detection Models
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
SciCom Wiki: Fact-Checking and FAIR Knowledge Distribution for Scientific Videos and Podcasts
by: Wittenborg, Tim, et al.
Published: (2025)
by: Wittenborg, Tim, et al.
Published: (2025)
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
by: Pan, Yiming, et al.
Published: (2026)
by: Pan, Yiming, et al.
Published: (2026)
HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
M$^3$Face: A Unified Multi-Modal Multilingual Framework for Human Face Generation and Editing
by: Mofayezi, Mohammadreza, et al.
Published: (2024)
by: Mofayezi, Mohammadreza, et al.
Published: (2024)
Order Is Not Layout: Order-to-Space Bias in Image Generation
by: Zhang, Yongkang, et al.
Published: (2026)
by: Zhang, Yongkang, et al.
Published: (2026)
Semantically Orthogonal Framework for Citation Classification: Disentangling Intent and Content
by: Duan, Changxu, et al.
Published: (2026)
by: Duan, Changxu, et al.
Published: (2026)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
An HTR-LLM Workflow for High-Accuracy Transcription and Analysis of Abbreviated Latin Court Hand
by: Isom, Joshua D.
Published: (2025)
by: Isom, Joshua D.
Published: (2025)
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
by: Lan, Jing, et al.
Published: (2026)
by: Lan, Jing, et al.
Published: (2026)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
by: Chen, Yijing, et al.
Published: (2025)
by: Chen, Yijing, et al.
Published: (2025)
DreamArtist++: Controllable One-Shot Text-to-Image Generation via Positive-Negative Adapter
by: Dong, Ziyi, et al.
Published: (2022)
by: Dong, Ziyi, et al.
Published: (2022)
AI Blob! LLM-Driven Recontextualization of Italian Television Archives
by: Balestri, Roberto
Published: (2025)
by: Balestri, Roberto
Published: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
by: Han, Wei, et al.
Published: (2023)
by: Han, Wei, et al.
Published: (2023)
Quo Vadis Handwritten Text Generation for Handwritten Text Recognition?
by: Pippi, Vittorio, et al.
Published: (2025)
by: Pippi, Vittorio, et al.
Published: (2025)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
by: Geng, Tiantian, et al.
Published: (2024)
by: Geng, Tiantian, et al.
Published: (2024)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
by: Liang, Zhengyang, et al.
Published: (2024)
by: Liang, Zhengyang, et al.
Published: (2024)
RemEdit: Efficient Diffusion Editing with Riemannian Geometry
by: Adhikarla, Eashan, et al.
Published: (2026)
by: Adhikarla, Eashan, et al.
Published: (2026)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Self-supervised Photographic Image Layout Representation Learning
by: Zhao, Zhaoran, et al.
Published: (2024)
by: Zhao, Zhaoran, et al.
Published: (2024)
Evolving Thematic Map Design in Academic Cartography: A Thirty-Year Study Based on Multilingual Journals
by: Wei, Zhiwei, et al.
Published: (2026)
by: Wei, Zhiwei, et al.
Published: (2026)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations
by: Jiang, Bowen, et al.
Published: (2025)
by: Jiang, Bowen, et al.
Published: (2025)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
by: Sun, Qianqian, et al.
Published: (2025)
by: Sun, Qianqian, et al.
Published: (2025)
Multi-Disciplinary Dataset Discovery from Citation-Verified Literature Contexts
by: Tan, Zhiyin, et al.
Published: (2026)
by: Tan, Zhiyin, et al.
Published: (2026)
Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation
by: Vaidya, Shreyas, et al.
Published: (2023)
by: Vaidya, Shreyas, et al.
Published: (2023)
CK-Transformer: Commonsense Knowledge Enhanced Transformers for Referring Expression Comprehension
by: Zhang, Zhi, et al.
Published: (2023)
by: Zhang, Zhi, et al.
Published: (2023)
Med-Banana-50K: A Cross-modality Large-Scale Dataset for Text-guided Medical Image Editing
by: Chen, Zhihui, et al.
Published: (2025)
by: Chen, Zhihui, et al.
Published: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024)
by: He, Yingqing, et al.
Published: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
by: Wang, Ziteng, et al.
Published: (2026)
by: Wang, Ziteng, et al.
Published: (2026)
Hyperbolic Safety-Aware Vision-Language Models
by: Poppi, Tobia, et al.
Published: (2025)
by: Poppi, Tobia, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Similar Items
-
Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation
by: Duan, Changxu
Published: (2025) -
Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
by: Shi, Xiang, et al.
Published: (2024) -
Differentiating Emigration from Return Migration of Scholars Using Name-Based Nationality Detection Models
by: Ghorbanpour, Faeze, et al.
Published: (2025) -
SciCom Wiki: Fact-Checking and FAIR Knowledge Distribution for Scientific Videos and Podcasts
by: Wittenborg, Tim, et al.
Published: (2025) -
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
by: Pan, Yiming, et al.
Published: (2026)