DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Sungnyun, Liao, Haofu, Appalaraju, Srikar, Tang, Peng, Tu, Zhuowen, Satzoda, Ravi Kumar, Manmatha, R., Mahadevan, Vijay, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Turbocharging Web Automation: The Impact of Compressed History States
von: Zhu, Xiyue, et al.
Veröffentlicht: (2025)
von: Zhu, Xiyue, et al.
Veröffentlicht: (2025)
Enhancing Vision-Language Pre-training with Rich Supervisions
von: Gao, Yuan, et al.
Veröffentlicht: (2024)
von: Gao, Yuan, et al.
Veröffentlicht: (2024)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024)
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024)
VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
von: Abramovich, Ofir, et al.
Veröffentlicht: (2024)
von: Abramovich, Ofir, et al.
Veröffentlicht: (2024)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
Non-autoregressive Sequence-to-Sequence Vision-Language Models
von: Shi, Kunyu, et al.
Veröffentlicht: (2024)
von: Shi, Kunyu, et al.
Veröffentlicht: (2024)
On the Scalability of Diffusion-based Text-to-Image Generation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024)
von: Van Landeghem, Jordy, et al.
Veröffentlicht: (2024)
Scalable Frameworks for Real-World Audio-Visual Speech Recognition
von: Kim, Sungnyun
Veröffentlicht: (2025)
von: Kim, Sungnyun
Veröffentlicht: (2025)
Mixed-Query Transformer: A Unified Image Segmentation Architecture
von: Wang, Pei, et al.
Veröffentlicht: (2024)
von: Wang, Pei, et al.
Veröffentlicht: (2024)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
von: Ding, Chuanghao, et al.
Veröffentlicht: (2024)
TopKD: Top-scaled Knowledge Distillation
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
von: Huang, Kui, et al.
Veröffentlicht: (2025)
von: Huang, Kui, et al.
Veröffentlicht: (2025)
MoKD: Multi-Task Optimization for Knowledge Distillation
von: Hayder, Zeeshan, et al.
Veröffentlicht: (2025)
von: Hayder, Zeeshan, et al.
Veröffentlicht: (2025)
EA-KD: Entropy-based Adaptive Knowledge Distillation
von: Su, Chi-Ping, et al.
Veröffentlicht: (2023)
von: Su, Chi-Ping, et al.
Veröffentlicht: (2023)
BD-KD: Balancing the Divergences for Online Knowledge Distillation
von: Amara, Ibtihel, et al.
Veröffentlicht: (2022)
von: Amara, Ibtihel, et al.
Veröffentlicht: (2022)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling
von: Wang, Yu, et al.
Veröffentlicht: (2022)
von: Wang, Yu, et al.
Veröffentlicht: (2022)
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
von: Zhao, Yue, et al.
Veröffentlicht: (2022)
von: Zhao, Yue, et al.
Veröffentlicht: (2022)
KD4MT: A Survey of Knowledge Distillation for Machine Translation
von: de Gibert, Ona, et al.
Veröffentlicht: (2026)
von: de Gibert, Ona, et al.
Veröffentlicht: (2026)
ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
von: Lan, Qizhen, et al.
Veröffentlicht: (2025)
von: Lan, Qizhen, et al.
Veröffentlicht: (2025)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
von: Wang, Jiabao, et al.
Veröffentlicht: (2023)
von: Wang, Jiabao, et al.
Veröffentlicht: (2023)
FreeKD: Knowledge Distillation via Semantic Frequency Prompt
von: Zhang, Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Yuan, et al.
Veröffentlicht: (2023)
Enhancing Knowledge Distillation for LLMs with Response-Priming Prompting
von: Goyal, Vijay, et al.
Veröffentlicht: (2024)
von: Goyal, Vijay, et al.
Veröffentlicht: (2024)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
von: Banerjee, Ayan, et al.
Veröffentlicht: (2024)
von: Banerjee, Ayan, et al.
Veröffentlicht: (2024)
DocAtlas: Multilingual Document Understanding Across 80+ Languages
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
von: Feng, Xiang, et al.
Veröffentlicht: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
von: Tan, Jing, et al.
Veröffentlicht: (2026)
von: Tan, Jing, et al.
Veröffentlicht: (2026)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2023)
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2023)
$\mathcal{X}$-KD: General Experiential Knowledge Distillation for Large Language Models
von: Cai, Yuang, et al.
Veröffentlicht: (2026)
von: Cai, Yuang, et al.
Veröffentlicht: (2026)
DocTER: Evaluating Document-based Knowledge Editing
von: Wu, Suhang, et al.
Veröffentlicht: (2023)
von: Wu, Suhang, et al.
Veröffentlicht: (2023)
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models
von: Sun, Haoyi, et al.
Veröffentlicht: (2026)
von: Sun, Haoyi, et al.
Veröffentlicht: (2026)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
von: Becker, Evan, et al.
Veröffentlicht: (2024)
von: Becker, Evan, et al.
Veröffentlicht: (2024)
ContraDoc: Understanding Self-Contradictions in Documents with Large Language Models
von: Li, Jierui, et al.
Veröffentlicht: (2023)
von: Li, Jierui, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Turbocharging Web Automation: The Impact of Compressed History States
von: Zhu, Xiyue, et al.
Veröffentlicht: (2025) -
Enhancing Vision-Language Pre-training with Rich Supervisions
von: Gao, Yuan, et al.
Veröffentlicht: (2024) -
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026) -
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
von: Rao, Varun Nagaraj, et al.
Veröffentlicht: (2024) -
VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
von: Abramovich, Ofir, et al.
Veröffentlicht: (2024)