PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Jaisankar, Vijay, Bandyopadhyay, Sambaran, Vyas, Kalp, Chaitanya, Varre, Somasundaram, Shwetha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
by: Somasundaram, Shwetha, et al.
Published: (2024)
by: Somasundaram, Shwetha, et al.
Published: (2024)
Presentations are not always linear! GNN meets LLM for Document-to-Presentation Transformation with Attribution
by: Maheshwari, Himanshu, et al.
Published: (2024)
by: Maheshwari, Himanshu, et al.
Published: (2024)
Enhancing Presentation Slide Generation by LLMs with a Multi-Staged End-to-End Approach
by: Bandyopadhyay, Sambaran, et al.
Published: (2024)
by: Bandyopadhyay, Sambaran, et al.
Published: (2024)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
by: Shin, Joongmin, et al.
Published: (2026)
by: Shin, Joongmin, et al.
Published: (2026)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
by: Anand, Nikhil, et al.
Published: (2026)
by: Anand, Nikhil, et al.
Published: (2026)
Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
by: Muppidi, Ananth, et al.
Published: (2025)
by: Muppidi, Ananth, et al.
Published: (2025)
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
by: Wang, Minzheng, et al.
Published: (2024)
by: Wang, Minzheng, et al.
Published: (2024)
DocCGen: Document-based Controlled Code Generation
by: Pimparkhede, Sameer, et al.
Published: (2024)
by: Pimparkhede, Sameer, et al.
Published: (2024)
Demonstrating ViviDoc: Generating Interactive Documents through Human-Agent Collaboration
by: Tang, Yinghao, et al.
Published: (2026)
by: Tang, Yinghao, et al.
Published: (2026)
LongDocFACTScore: Evaluating the Factuality of Long Document Abstractive Summarisation
by: Bishop, Jennifer A, et al.
Published: (2023)
by: Bishop, Jennifer A, et al.
Published: (2023)
DocMEdit: Towards Document-Level Model Editing
by: Zeng, Li, et al.
Published: (2025)
by: Zeng, Li, et al.
Published: (2025)
DocTER: Evaluating Document-based Knowledge Editing
by: Wu, Suhang, et al.
Published: (2023)
by: Wu, Suhang, et al.
Published: (2023)
Submodular Evaluation Subset Selection in Automatic Prompt Optimization
by: Nian, Jinming, et al.
Published: (2026)
by: Nian, Jinming, et al.
Published: (2026)
Doc-Guided Sent2Sent++: A Sent2Sent++ Agent with Doc-Guided memory for Document-level Machine Translation
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
DocMamba: Efficient Document Pre-training with State Space Model
by: Hu, Pengfei, et al.
Published: (2024)
by: Hu, Pengfei, et al.
Published: (2024)
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
by: Giovannini, Simone, et al.
Published: (2025)
by: Giovannini, Simone, et al.
Published: (2025)
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks
by: Ni, Feng, et al.
Published: (2025)
by: Ni, Feng, et al.
Published: (2025)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
by: Huang, Kui, et al.
Published: (2025)
by: Huang, Kui, et al.
Published: (2025)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
DocReLM: Mastering Document Retrieval with Language Model
by: Wei, Gengchen, et al.
Published: (2024)
by: Wei, Gengchen, et al.
Published: (2024)
DocAgent: A Multi-Agent System for Automated Code Documentation Generation
by: Yang, Dayu, et al.
Published: (2025)
by: Yang, Dayu, et al.
Published: (2025)
VaeDiff-DocRE: End-to-end Data Augmentation Framework for Document-level Relation Extraction
by: Tran, Khai Phan, et al.
Published: (2024)
by: Tran, Khai Phan, et al.
Published: (2024)
Code2Doc: A Quality-First Curated Dataset for Code Documentation
by: Karaman, Recep Kaan, et al.
Published: (2025)
by: Karaman, Recep Kaan, et al.
Published: (2025)
DocReward: A Document Reward Model for Structuring and Stylizing
by: Liu, Junpeng, et al.
Published: (2025)
by: Liu, Junpeng, et al.
Published: (2025)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025)
by: Bandyopadhyay, Dibyanayan, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
Doc-to-LoRA: Learning to Instantly Internalize Contexts
by: Charakorn, Rujikorn, et al.
Published: (2026)
by: Charakorn, Rujikorn, et al.
Published: (2026)
Is This a Bad Table? A Closer Look at the Evaluation of Table Generation from Text
by: Ramu, Pritika, et al.
Published: (2024)
by: Ramu, Pritika, et al.
Published: (2024)
Infogen: Generating Complex Statistical Infographics from Documents
by: Ghosh, Akash, et al.
Published: (2025)
by: Ghosh, Akash, et al.
Published: (2025)
DocEDA: Automated Extraction and Design of Analog Circuits from Documents with Large Language Model
by: Chen, Hong Cai, et al.
Published: (2024)
by: Chen, Hong Cai, et al.
Published: (2024)
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
by: Zhu, Fengbin, et al.
Published: (2023)
by: Zhu, Fengbin, et al.
Published: (2023)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
by: Pang, Wei, et al.
Published: (2025)
by: Pang, Wei, et al.
Published: (2025)
Synthetic Multimodal Question Generation
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents
by: Galarnyk, Michael, et al.
Published: (2026)
by: Galarnyk, Michael, et al.
Published: (2026)
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024)
by: Wu, Yiheng, et al.
Published: (2024)
Docs2KG: Unified Knowledge Graph Construction from Heterogeneous Documents Assisted by Large Language Models
by: Sun, Qiang, et al.
Published: (2024)
by: Sun, Qiang, et al.
Published: (2024)
Similar Items
-
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
by: Somasundaram, Shwetha, et al.
Published: (2024) -
Presentations are not always linear! GNN meets LLM for Document-to-Presentation Transformation with Attribution
by: Maheshwari, Himanshu, et al.
Published: (2024) -
Enhancing Presentation Slide Generation by LLMs with a Multi-Staged End-to-End Approach
by: Bandyopadhyay, Sambaran, et al.
Published: (2024) -
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024) -
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
by: Shin, Joongmin, et al.
Published: (2026)