ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yan, Han, Soyeon Caren, Dai, Yue, Cao, Feiqi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024)
by: Ng, Thye Shan, et al.
Published: (2024)
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
by: Han, Soyeon Caren, et al.
Published: (2024)
by: Han, Soyeon Caren, et al.
Published: (2024)
In-game Toxic Language Detection: Shared Task and Attention Residuals
by: Jia, Yuanzhe, et al.
Published: (2022)
by: Jia, Yuanzhe, et al.
Published: (2022)
PEACH: Pretrained-embedding Explanation Across Contextual and Hierarchical Structure
by: Cao, Feiqi, et al.
Published: (2024)
by: Cao, Feiqi, et al.
Published: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025)
by: Dai, Yue, et al.
Published: (2025)
MSG-Chart: Multimodal Scene Graph for ChartQA
by: Dai, Yue, et al.
Published: (2024)
by: Dai, Yue, et al.
Published: (2024)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)
by: Akasaka, Muku, et al.
Published: (2026)
Game-MUG: Multimodal Oriented Game Situation Understanding and Commentary Generation Dataset
by: Zhang, Zhihao, et al.
Published: (2024)
by: Zhang, Zhihao, et al.
Published: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
by: Xiang, Biao, et al.
Published: (2026)
by: Xiang, Biao, et al.
Published: (2026)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Graph Neural Networks for Text Classification: A Survey
by: Wang, Kunze, et al.
Published: (2023)
by: Wang, Kunze, et al.
Published: (2023)
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Enhancing Document Key Information Localization Through Data Augmentation
by: Dai, Yue
Published: (2025)
by: Dai, Yue
Published: (2025)
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
by: Li, Haoyi, et al.
Published: (2025)
by: Li, Haoyi, et al.
Published: (2025)
EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues
by: Yuan, Angela Yifei, et al.
Published: (2025)
by: Yuan, Angela Yifei, et al.
Published: (2025)
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
by: Park, Jiwon, et al.
Published: (2025)
by: Park, Jiwon, et al.
Published: (2025)
SynJAC: Synthetic-data-driven Joint-granular Adaptation and Calibration for Domain Specific Scanned Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
A Survey of Large Language Models in Finance (FinLLMs)
by: Lee, Jean, et al.
Published: (2024)
by: Lee, Jean, et al.
Published: (2024)
SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2026)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2026)
NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long Documents
by: Czinczoll, Tamara, et al.
Published: (2024)
by: Czinczoll, Tamara, et al.
Published: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
by: Chen, Huiyao, et al.
Published: (2025)
by: Chen, Huiyao, et al.
Published: (2025)
TriG-NER: Triplet-Grid Framework for Discontinuous Named Entity Recognition
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
by: Günther, Michael, et al.
Published: (2024)
by: Günther, Michael, et al.
Published: (2024)
Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers
by: Xie, Jiawen, et al.
Published: (2023)
by: Xie, Jiawen, et al.
Published: (2023)
R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
ChuXin: 1.6B Technical Report
by: Zhuang, Xiaomin, et al.
Published: (2024)
by: Zhuang, Xiaomin, et al.
Published: (2024)
MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models
by: Lee, Suchan, et al.
Published: (2025)
by: Lee, Suchan, et al.
Published: (2025)
3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
by: Wu, Kaifeng, et al.
Published: (2025)
by: Wu, Kaifeng, et al.
Published: (2025)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
by: Shin, Joongmin, et al.
Published: (2026)
by: Shin, Joongmin, et al.
Published: (2026)
Docopilot: Improving Multimodal Models for Document-Level Understanding
by: Duan, Yuchen, et al.
Published: (2025)
by: Duan, Yuchen, et al.
Published: (2025)
Similar Items
-
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024) -
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
by: Han, Soyeon Caren, et al.
Published: (2024) -
In-game Toxic Language Detection: Shared Task and Attention Residuals
by: Jia, Yuanzhe, et al.
Published: (2022) -
PEACH: Pretrained-embedding Explanation Across Contextual and Hierarchical Structure
by: Cao, Feiqi, et al.
Published: (2024) -
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025)