What to Format and How: A Benchmark and Workflow Approach for Document Formatting
Fuente:
arXiv
Saved in:
| Main Authors: | Rao, Shihao, Li, Liang, Liu, Jiapeng, Lin, Tong, Li, Bing, Gao, Xiyan, Fu, Peng, Huang, Jing, Ma, Can |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Reliable is Multilingual LLM-as-a-Judge?
by: Fu, Xiyan, et al.
Published: (2025)
by: Fu, Xiyan, et al.
Published: (2025)
EXaMCaP: Subset Selection with Entropy Gain Maximization for Probing Capability Gains of Large Chart Understanding Training Sets
by: Liu, Jiapeng, et al.
Published: (2026)
by: Liu, Jiapeng, et al.
Published: (2026)
Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
by: Guan, Weixin, et al.
Published: (2026)
by: Guan, Weixin, et al.
Published: (2026)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
Exploring Format Consistency for Instruction Tuning
by: Liang, Shihao, et al.
Published: (2023)
by: Liang, Shihao, et al.
Published: (2023)
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
by: Fu, Xiyan, et al.
Published: (2026)
by: Fu, Xiyan, et al.
Published: (2026)
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
by: Wang, Dingzirui, et al.
Published: (2025)
by: Wang, Dingzirui, et al.
Published: (2025)
From Lists to Emojis: How Format Bias Affects Model Alignment
by: Zhang, Xuanchang, et al.
Published: (2024)
by: Zhang, Xuanchang, et al.
Published: (2024)
Exploring Continual Learning of Compositional Generalization in NLI
by: Fu, Xiyan, et al.
Published: (2024)
by: Fu, Xiyan, et al.
Published: (2024)
The Mystery of Compositional Generalization in Graph-based Generative Commonsense Reasoning
by: Fu, Xiyan, et al.
Published: (2024)
by: Fu, Xiyan, et al.
Published: (2024)
An Extensive Study on Text Serialization Formats and Methods
by: Wei, Wang, et al.
Published: (2025)
by: Wei, Wang, et al.
Published: (2025)
Open-domain Implicit Format Control for Large Language Model Generation
by: Yao, Yiqun, et al.
Published: (2024)
by: Yao, Yiqun, et al.
Published: (2024)
Policy-driven Knowledge Selection and Response Generation for Document-grounded Dialogue
by: Ma, Longxuan, et al.
Published: (2024)
by: Ma, Longxuan, et al.
Published: (2024)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
by: Xiao, Ruixuan, et al.
Published: (2024)
by: Xiao, Ruixuan, et al.
Published: (2024)
What if you said that differently?: How Explanation Formats Affect Human Feedback Efficacy and User Perception
by: Malaviya, Chaitanya, et al.
Published: (2023)
by: Malaviya, Chaitanya, et al.
Published: (2023)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
by: Zhang, Manyi, et al.
Published: (2026)
by: Zhang, Manyi, et al.
Published: (2026)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
by: Zhao, Yibo, et al.
Published: (2024)
by: Zhao, Yibo, et al.
Published: (2024)
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction
by: Lin, Zening, et al.
Published: (2024)
by: Lin, Zening, et al.
Published: (2024)
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
by: Dong, Xuan, et al.
Published: (2026)
by: Dong, Xuan, et al.
Published: (2026)
Automating Date Format Detection for Data Visualization
by: Liang, Zixuan
Published: (2025)
by: Liang, Zixuan
Published: (2025)
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
by: Guo, Zhicheng, et al.
Published: (2024)
by: Guo, Zhicheng, et al.
Published: (2024)
How Does the Disclosure of AI Assistance Affect the Perceptions of Writing?
by: Li, Zhuoyan, et al.
Published: (2024)
by: Li, Zhuoyan, et al.
Published: (2024)
URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding
by: Shi, Yongxin, et al.
Published: (2025)
by: Shi, Yongxin, et al.
Published: (2025)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
by: Xia, Congying, et al.
Published: (2024)
by: Xia, Congying, et al.
Published: (2024)
Benchmarking Real-Time Question Answering via Executable Code Workflows
by: Zhou, Wenjie, et al.
Published: (2026)
by: Zhou, Wenjie, et al.
Published: (2026)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
by: Li, Yaoming, et al.
Published: (2026)
by: Li, Yaoming, et al.
Published: (2026)
RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking Format
by: Huang, Zhehao, et al.
Published: (2026)
by: Huang, Zhehao, et al.
Published: (2026)
What Are They Talking About? A Benchmark of Knowledge-Grounded Discussion Summarization
by: Zhou, Weixiao, et al.
Published: (2025)
by: Zhou, Weixiao, et al.
Published: (2025)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
by: Zhang, Xingjian, et al.
Published: (2024)
by: Zhang, Xingjian, et al.
Published: (2024)
Tree of Reviews: A Tree-based Dynamic Iterative Retrieval Framework for Multi-hop Question Answering
by: Jiapeng, Li, et al.
Published: (2024)
by: Jiapeng, Li, et al.
Published: (2024)
The Format Tax
by: Lee, Ivan Yee, et al.
Published: (2026)
by: Lee, Ivan Yee, et al.
Published: (2026)
Decoupling Task-Solving and Output Formatting in LLM Generation
by: Deng, Haikang, et al.
Published: (2025)
by: Deng, Haikang, et al.
Published: (2025)
Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach
by: Li, Xingyu, et al.
Published: (2025)
by: Li, Xingyu, et al.
Published: (2025)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
by: Ma, Longxuan, et al.
Published: (2024)
by: Ma, Longxuan, et al.
Published: (2024)
Structured Document Translation via Format Reinforcement Learning
by: Song, Haiyue, et al.
Published: (2025)
by: Song, Haiyue, et al.
Published: (2025)
WorkTeam: Constructing Workflows from Natural Language with Multi-Agents
by: Liu, Hanchao, et al.
Published: (2025)
by: Liu, Hanchao, et al.
Published: (2025)
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
by: Sun, Haoyu, et al.
Published: (2025)
by: Sun, Haoyu, et al.
Published: (2025)
Benchmarking and Learning Real-World Customer Service Dialogue
by: Gao, Tianhong, et al.
Published: (2025)
by: Gao, Tianhong, et al.
Published: (2025)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
by: Wang, Jin, et al.
Published: (2026)
by: Wang, Jin, et al.
Published: (2026)
Underutilization of Syntactic Processing by Chinese Learners of English in Comprehending English Sentences, Evidenced from Adapted Garden-Path Ambiguity Experiment
by: Xu, Jiapeng
Published: (2024)
by: Xu, Jiapeng
Published: (2024)
Similar Items
-
How Reliable is Multilingual LLM-as-a-Judge?
by: Fu, Xiyan, et al.
Published: (2025) -
EXaMCaP: Subset Selection with Entropy Gain Maximization for Probing Capability Gains of Large Chart Understanding Training Sets
by: Liu, Jiapeng, et al.
Published: (2026) -
Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
by: Guan, Weixin, et al.
Published: (2026) -
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025) -
Exploring Format Consistency for Instruction Tuning
by: Liang, Shihao, et al.
Published: (2023)