DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Suri, Manan, Mathur, Puneet, Dernoncourt, Franck, Jain, Rajiv, Morariu, Vlad I, Sawhney, Ramit, Nakov, Preslav, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
by: Suri, Manan, et al.
Published: (2024)
by: Suri, Manan, et al.
Published: (2024)
ChartLens: Fine-grained Visual Attribution in Charts
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
Structured Uncertainty guided Clarification for LLM Agents
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024)
by: Biswas, Sanket, et al.
Published: (2024)
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
AnyDoc: Enhancing Document Generation via Large-Scale HTML/CSS Data Synthesis and Height-Aware Reinforcement Optimization
by: Lin, Jiawei, et al.
Published: (2026)
by: Lin, Jiawei, et al.
Published: (2026)
Learning Illumination Control in Diffusion Models
by: Anand, Nishit, et al.
Published: (2026)
by: Anand, Nishit, et al.
Published: (2026)
Charts Are Not Images: On the Challenges of Scientific Chart Editing
by: Li, Shawn, et al.
Published: (2025)
by: Li, Shawn, et al.
Published: (2025)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
by: Glockner, Max, et al.
Published: (2024)
by: Glockner, Max, et al.
Published: (2024)
Multimodal Large Language Models to Support Real-World Fact-Checking
by: Geng, Jiahui, et al.
Published: (2024)
by: Geng, Jiahui, et al.
Published: (2024)
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing
by: Agarwal, Siddhant, et al.
Published: (2024)
by: Agarwal, Siddhant, et al.
Published: (2024)
Cluster-R1: Large Reasoning Models Are Instruction-following Clustering Agents
by: Qing, Peijun, et al.
Published: (2026)
by: Qing, Peijun, et al.
Published: (2026)
MiLDEdit: Reasoning-Based Multi-Layer Design Document Editing
by: Lin, Zihao, et al.
Published: (2026)
by: Lin, Zihao, et al.
Published: (2026)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
by: Geng, Jiahui, et al.
Published: (2025)
by: Geng, Jiahui, et al.
Published: (2025)
DocMEdit: Towards Document-Level Model Editing
by: Zeng, Li, et al.
Published: (2025)
by: Zeng, Li, et al.
Published: (2025)
DocTER: Evaluating Document-based Knowledge Editing
by: Wu, Suhang, et al.
Published: (2023)
by: Wu, Suhang, et al.
Published: (2023)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
by: Ren, Kaixuan, et al.
Published: (2025)
by: Ren, Kaixuan, et al.
Published: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
by: Sundriyal, Megha, et al.
Published: (2023)
by: Sundriyal, Megha, et al.
Published: (2023)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
by: Alhindi, Tariq, et al.
Published: (2023)
by: Alhindi, Tariq, et al.
Published: (2023)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
DenoiseRank: Learning to Rank by Diffusion Models
by: Wang, Ying, et al.
Published: (2026)
by: Wang, Ying, et al.
Published: (2026)
Adapting Fake News Detection to the Era of Large Language Models
by: Su, Jinyan, et al.
Published: (2023)
by: Su, Jinyan, et al.
Published: (2023)
Corpus Poisoning via Approximate Greedy Gradient Descent
by: Su, Jinyan, et al.
Published: (2024)
by: Su, Jinyan, et al.
Published: (2024)
DocuBits: VR Document Decomposition for Procedural Task Completion
by: Lee, Geonsun, et al.
Published: (2024)
by: Lee, Geonsun, et al.
Published: (2024)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
TutoAI: A Cross-domain Framework for AI-assisted Mixed-media Tutorial Creation on Physical Tasks
by: Chen, Yuexi, et al.
Published: (2024)
by: Chen, Yuexi, et al.
Published: (2024)
Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
by: Fang, Jiangnan, et al.
Published: (2026)
by: Fang, Jiangnan, et al.
Published: (2026)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
by: Tomar, Raj Vardhan, et al.
Published: (2025)
by: Tomar, Raj Vardhan, et al.
Published: (2025)
How Does Prefix Matter in Reasoning Model Tuning?
by: Tomar, Raj Vardhan, et al.
Published: (2026)
by: Tomar, Raj Vardhan, et al.
Published: (2026)
Non-Invasive Qualitative Vibration Analysis using Event Camera
by: Bane, Dwijay, et al.
Published: (2024)
by: Bane, Dwijay, et al.
Published: (2024)
Mitigating Memorization in LLMs using Activation Steering
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
Similar Items
-
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
by: Suri, Manan, et al.
Published: (2024) -
ChartLens: Fine-grained Visual Attribution in Charts
by: Suri, Manan, et al.
Published: (2025) -
Structured Uncertainty guided Clarification for LLM Agents
by: Suri, Manan, et al.
Published: (2025) -
DocSynthv2: A Practical Autoregressive Modeling for Document Generation
by: Biswas, Sanket, et al.
Published: (2024) -
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
by: Goswami, Kanika, et al.
Published: (2025)