SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Jiefeng, Wang, Yan, Liu, Chenyu, Du, Jun, Hu, Yu, Zhang, Zhenrong, Hu, Pengfei, Wang, Qing, Zhang, Jianshu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DocMamba: Efficient Document Pre-training with State Space Model
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
von: Chang, Qikai, et al.
Veröffentlicht: (2025)
von: Chang, Qikai, et al.
Veröffentlicht: (2025)
See then Tell: Enhancing Key Information Extraction with Vision Grounding
von: Liu, Shuhang, et al.
Veröffentlicht: (2024)
von: Liu, Shuhang, et al.
Veröffentlicht: (2024)
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2024)
Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
von: Pan, Yicheng, et al.
Veröffentlicht: (2025)
von: Pan, Yicheng, et al.
Veröffentlicht: (2025)
Bidirectional Trained Tree-Structured Decoder for Handwritten Mathematical Expression Recognition
von: Cheng, Hanbo, et al.
Veröffentlicht: (2023)
von: Cheng, Hanbo, et al.
Veröffentlicht: (2023)
SEMv3: A Fast and Robust Approach to Table Separation Line Detection
von: Qin, Chunxia, et al.
Veröffentlicht: (2024)
von: Qin, Chunxia, et al.
Veröffentlicht: (2024)
PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
von: Chang, Qikai, et al.
Veröffentlicht: (2026)
von: Chang, Qikai, et al.
Veröffentlicht: (2026)
Skeleton and Font Generation Network for Zero-shot Chinese Character Generation
von: Xue, Mobai, et al.
Veröffentlicht: (2025)
von: Xue, Mobai, et al.
Veröffentlicht: (2025)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
von: Liu, Shuhang, et al.
Veröffentlicht: (2025)
von: Liu, Shuhang, et al.
Veröffentlicht: (2025)
Step Potential Advantage Estimation: Harnessing Intermediate Confidence and Correctness for Efficient Mathematical Reasoning
von: Wu, Fei, et al.
Veröffentlicht: (2026)
von: Wu, Fei, et al.
Veröffentlicht: (2026)
SEMv2: Table Separation Line Detection Based on Instance Segmentation
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2023)
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2023)
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
RFL: Simplifying Chemical Structure Recognition with Ring-Free Language
von: Chang, Qikai, et al.
Veröffentlicht: (2024)
von: Chang, Qikai, et al.
Veröffentlicht: (2024)
Event-Centric Human Value Understanding in News-Domain Texts: An Actor-Conditioned, Multi-Granularity Benchmark
von: Wang, Yao, et al.
Veröffentlicht: (2026)
von: Wang, Yao, et al.
Veröffentlicht: (2026)
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
MGSA: Multi-Granularity Graph Structure Attention for Knowledge Graph-to-Text Generation
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
von: Su, Jiamin, et al.
Veröffentlicht: (2025)
WebCoT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback
von: Hu, Minda, et al.
Veröffentlicht: (2025)
von: Hu, Minda, et al.
Veröffentlicht: (2025)
DTELS: Towards Dynamic Granularity of Timeline Summarization
von: Zhang, Chenlong, et al.
Veröffentlicht: (2024)
von: Zhang, Chenlong, et al.
Veröffentlicht: (2024)
Multi-Granularity Semantic Revision for Large Language Model Distillation
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
Multi-Granular Multimodal Clue Fusion for Meme Understanding
von: Zheng, Li, et al.
Veröffentlicht: (2025)
von: Zheng, Li, et al.
Veröffentlicht: (2025)
KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
von: Chen, Junhao, et al.
Veröffentlicht: (2025)
Keywords and Instances: A Hierarchical Contrastive Learning Framework Unifying Hybrid Granularities for Text Generation
von: Li, Mingzhe, et al.
Veröffentlicht: (2022)
von: Li, Mingzhe, et al.
Veröffentlicht: (2022)
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
von: Li, Zizhong, et al.
Veröffentlicht: (2026)
Multi-Granularity Information Interaction Framework for Incomplete Utterance Rewriting
von: Du, Haowei, et al.
Veröffentlicht: (2023)
von: Du, Haowei, et al.
Veröffentlicht: (2023)
Large Language Model for Multi-Domain Translation: Benchmarking and Domain CoT Fine-tuning
von: Hu, Tianxiang, et al.
Veröffentlicht: (2024)
von: Hu, Tianxiang, et al.
Veröffentlicht: (2024)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Multi-Granularity Open Intent Classification via Adaptive Granular-Ball Decision Boundary
von: Li, Yanhua, et al.
Veröffentlicht: (2024)
von: Li, Yanhua, et al.
Veröffentlicht: (2024)
Multi-Modal Multi-Granularity Tokenizer for Chu Bamboo Slip Scripts
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
von: Chen, Yingfa, et al.
Veröffentlicht: (2024)
Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
Investigating the Impact of Rationales for LLMs on Natural Language Understanding
von: Shi, Wenhang, et al.
Veröffentlicht: (2025)
von: Shi, Wenhang, et al.
Veröffentlicht: (2025)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
Boosting Disfluency Detection with Large Language Model as Disfluency Generator
von: Cheng, Zhenrong, et al.
Veröffentlicht: (2024)
von: Cheng, Zhenrong, et al.
Veröffentlicht: (2024)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
von: Hu, Jinwu, et al.
Veröffentlicht: (2025)
SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding
von: Hou, Shuyang, et al.
Veröffentlicht: (2026)
von: Hou, Shuyang, et al.
Veröffentlicht: (2026)
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
von: Cheng, Hanbo, et al.
Veröffentlicht: (2024)
von: Cheng, Hanbo, et al.
Veröffentlicht: (2024)
Continual Few-shot Event Detection via Hierarchical Augmentation Networks
von: Zhang, Chenlong, et al.
Veröffentlicht: (2024)
von: Zhang, Chenlong, et al.
Veröffentlicht: (2024)
MoPS: Modular Story Premise Synthesis for Open-Ended Automatic Story Generation
von: Ma, Yan, et al.
Veröffentlicht: (2024)
von: Ma, Yan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DocMamba: Efficient Document Pre-training with State Space Model
von: Hu, Pengfei, et al.
Veröffentlicht: (2024) -
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
von: Chang, Qikai, et al.
Veröffentlicht: (2025) -
See then Tell: Enhancing Key Information Extraction with Vision Grounding
von: Liu, Shuhang, et al.
Veröffentlicht: (2024) -
UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition
von: Zhang, Zhenrong, et al.
Veröffentlicht: (2024) -
Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
von: Pan, Yicheng, et al.
Veröffentlicht: (2025)