LAPDoc: Layout-Aware Prompting for Documents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lamott, Marcel, Weweler, Yves-Noel, Ulges, Adrian, Shafait, Faisal, Krechel, Dirk, Obradovic, Darko |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
von: Wang, Baode, et al.
Veröffentlicht: (2025)
von: Wang, Baode, et al.
Veröffentlicht: (2025)
Benchmarking Graph Neural Networks for Document Layout Analysis in Public Affairs
von: Lopez-Duran, Miguel, et al.
Veröffentlicht: (2025)
von: Lopez-Duran, Miguel, et al.
Veröffentlicht: (2025)
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
von: Fujitake, Masato
Veröffentlicht: (2024)
von: Fujitake, Masato
Veröffentlicht: (2024)
Leveraging Distillation Techniques for Document Understanding: A Case Study with FLAN-T5
von: Lamott, Marcel, et al.
Veröffentlicht: (2024)
von: Lamott, Marcel, et al.
Veröffentlicht: (2024)
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
von: Liu, Xinyang, et al.
Veröffentlicht: (2023)
von: Liu, Xinyang, et al.
Veröffentlicht: (2023)
PromptTA: Prompt-driven Text Adapter for Source-free Domain Generalization
von: Zhang, Haoran, et al.
Veröffentlicht: (2024)
von: Zhang, Haoran, et al.
Veröffentlicht: (2024)
Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
von: Cho, Jaemin, et al.
Veröffentlicht: (2023)
Prompt as Free Lunch: Enhancing Diversity in Source-Free Cross-domain Few-shot Learning through Semantic-Guided Prompting
von: Zhuo, Linhai, et al.
Veröffentlicht: (2024)
von: Zhuo, Linhai, et al.
Veröffentlicht: (2024)
Semantic Prompt Learning for Weakly-Supervised Semantic Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
Answering Questions in Stages: Prompt Chaining for Contract QA
von: Roegiest, Adam, et al.
Veröffentlicht: (2024)
von: Roegiest, Adam, et al.
Veröffentlicht: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
von: Du, Yingjun, et al.
Veröffentlicht: (2024)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
von: Cheng, Jiale, et al.
Veröffentlicht: (2025)
Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2025)
Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
von: Beliaev, Mark, et al.
Veröffentlicht: (2025)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
von: Jie, Shibo, et al.
Veröffentlicht: (2024)
One Category One Prompt: Dataset Distillation using Diffusion Models
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
von: Abbasi, Ali, et al.
Veröffentlicht: (2024)
Context-Aware Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
von: Luo, Jun, et al.
Veröffentlicht: (2024)
von: Luo, Jun, et al.
Veröffentlicht: (2024)
PQPP: A Joint Benchmark for Text-to-Image Prompt and Query Performance Prediction
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
von: Poesina, Eduard, et al.
Veröffentlicht: (2024)
LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
DocDjinn: Controllable Synthetic Document Generation with VLMs and Handwriting Diffusion
von: Lamott, Marcel, et al.
Veröffentlicht: (2026)
von: Lamott, Marcel, et al.
Veröffentlicht: (2026)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
von: Meng, Debin, et al.
Veröffentlicht: (2025)
von: Meng, Debin, et al.
Veröffentlicht: (2025)
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
von: Chen, Menglan, et al.
Veröffentlicht: (2025)
von: Chen, Menglan, et al.
Veröffentlicht: (2025)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
von: Luo, Chuwei, et al.
Veröffentlicht: (2024)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
Which Client is Reliable?: A Reliable and Personalized Prompt-based Federated Learning for Medical Image Question Answering
von: Zhu, He, et al.
Veröffentlicht: (2024)
von: Zhu, He, et al.
Veröffentlicht: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
Robustness of Structured Data Extraction from Perspectively Distorted Documents
von: Nakada, Hyakka, et al.
Veröffentlicht: (2025)
von: Nakada, Hyakka, et al.
Veröffentlicht: (2025)
Enhancing Post-Training Quantization via Future Activation Awareness
von: Lv, Zheqi, et al.
Veröffentlicht: (2026)
von: Lv, Zheqi, et al.
Veröffentlicht: (2026)
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
von: Hamed, Omar, et al.
Veröffentlicht: (2024)
von: Hamed, Omar, et al.
Veröffentlicht: (2024)
PALP: Prompt Aligned Personalization of Text-to-Image Models
von: Arar, Moab, et al.
Veröffentlicht: (2024)
von: Arar, Moab, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
von: Du, Yao, et al.
Veröffentlicht: (2026)
von: Du, Yao, et al.
Veröffentlicht: (2026)
Bengali Document Layout Analysis -- A YOLOV8 Based Ensembling Approach
von: Ahmed, Nazmus Sakib, et al.
Veröffentlicht: (2023)
von: Ahmed, Nazmus Sakib, et al.
Veröffentlicht: (2023)
Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
von: Khayatan, Pegah, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
von: Wang, Baode, et al.
Veröffentlicht: (2025) -
Benchmarking Graph Neural Networks for Document Layout Analysis in Public Affairs
von: Lopez-Duran, Miguel, et al.
Veröffentlicht: (2025) -
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
von: Fujitake, Masato
Veröffentlicht: (2024) -
Leveraging Distillation Techniques for Document Understanding: A Case Study with FLAN-T5
von: Lamott, Marcel, et al.
Veröffentlicht: (2024) -
Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)