On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Zhitao, Yang, Haolin, Min, Rui, Qin, Zeyu, Fung, Yi R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
di: He, Zhitao, et al.
Pubblicazione: (2025)
di: He, Zhitao, et al.
Pubblicazione: (2025)
Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
di: Ju, Feng, et al.
Pubblicazione: (2025)
di: Ju, Feng, et al.
Pubblicazione: (2025)
MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL
di: Yang, Haolin, et al.
Pubblicazione: (2025)
di: Yang, Haolin, et al.
Pubblicazione: (2025)
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
di: He, Zhitao, et al.
Pubblicazione: (2026)
di: He, Zhitao, et al.
Pubblicazione: (2026)
MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?
di: He, Zhitao, et al.
Pubblicazione: (2025)
di: He, Zhitao, et al.
Pubblicazione: (2025)
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization
di: He, Zhitao, et al.
Pubblicazione: (2025)
di: He, Zhitao, et al.
Pubblicazione: (2025)
MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
di: He, Zhitao, et al.
Pubblicazione: (2025)
di: He, Zhitao, et al.
Pubblicazione: (2025)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
di: Tian, Yufei, et al.
Pubblicazione: (2024)
di: Tian, Yufei, et al.
Pubblicazione: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness
di: Huang, Junsheng, et al.
Pubblicazione: (2025)
di: Huang, Junsheng, et al.
Pubblicazione: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
LaPA$^2$: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
di: Yang, Jiabing, et al.
Pubblicazione: (2025)
di: Yang, Jiabing, et al.
Pubblicazione: (2025)
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
di: Ruan, Jie, et al.
Pubblicazione: (2025)
di: Ruan, Jie, et al.
Pubblicazione: (2025)
What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation
di: Yang, Dingyi, et al.
Pubblicazione: (2025)
di: Yang, Dingyi, et al.
Pubblicazione: (2025)
UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios
di: Luo, Haotian, et al.
Pubblicazione: (2025)
di: Luo, Haotian, et al.
Pubblicazione: (2025)
Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
di: Wang, Chengbing, et al.
Pubblicazione: (2025)
di: Wang, Chengbing, et al.
Pubblicazione: (2025)
Integrating Planning into Single-Turn Long-Form Text Generation
di: Liang, Yi, et al.
Pubblicazione: (2024)
di: Liang, Yi, et al.
Pubblicazione: (2024)
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
di: Geng, Xinyu, et al.
Pubblicazione: (2026)
di: Geng, Xinyu, et al.
Pubblicazione: (2026)
LongStory: Coherent, Complete and Length Controlled Long story Generation
di: Park, Kyeongman, et al.
Pubblicazione: (2023)
di: Park, Kyeongman, et al.
Pubblicazione: (2023)
How Does Response Length Affect Long-Form Factuality
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
di: Zhao, James Xu, et al.
Pubblicazione: (2025)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
di: Costa-jussà, Marta R., et al.
Pubblicazione: (2024)
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
di: Xiao, Zikai, et al.
Pubblicazione: (2025)
di: Xiao, Zikai, et al.
Pubblicazione: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Precise Information Control in Long-Form Text Generation
di: He, Jacqueline, et al.
Pubblicazione: (2025)
di: He, Jacqueline, et al.
Pubblicazione: (2025)
Atomic Calibration of LLMs in Long-Form Generations
di: Zhang, Caiqi, et al.
Pubblicazione: (2024)
di: Zhang, Caiqi, et al.
Pubblicazione: (2024)
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
di: Tu, Lifu, et al.
Pubblicazione: (2024)
di: Tu, Lifu, et al.
Pubblicazione: (2024)
KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
di: Chen, Junhao, et al.
Pubblicazione: (2025)
di: Chen, Junhao, et al.
Pubblicazione: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
LongProc: Benchmarking Long-Context Language Models on Long Procedural Generation
di: Ye, Xi, et al.
Pubblicazione: (2025)
di: Ye, Xi, et al.
Pubblicazione: (2025)
The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
di: Jacovi, Alon, et al.
Pubblicazione: (2025)
di: Jacovi, Alon, et al.
Pubblicazione: (2025)
A Benchmark for Long-Form Medical Question Answering
di: Hosseini, Pedram, et al.
Pubblicazione: (2024)
di: Hosseini, Pedram, et al.
Pubblicazione: (2024)
Learning to Reason for Long-Form Story Generation
di: Gurung, Alexander, et al.
Pubblicazione: (2025)
di: Gurung, Alexander, et al.
Pubblicazione: (2025)
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K
di: Yuan, Tao, et al.
Pubblicazione: (2024)
di: Yuan, Tao, et al.
Pubblicazione: (2024)
A Cognitive Writing Perspective for Constrained Long-Form Text Generation
di: Wan, Kaiyang, et al.
Pubblicazione: (2025)
di: Wan, Kaiyang, et al.
Pubblicazione: (2025)
Massively Multi-Cultural Knowledge Acquisition & LM Benchmarking
di: Fung, Yi, et al.
Pubblicazione: (2024)
di: Fung, Yi, et al.
Pubblicazione: (2024)
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
di: Deng, Haoran, et al.
Pubblicazione: (2025)
di: Deng, Haoran, et al.
Pubblicazione: (2025)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
di: Cheng, Zicong, et al.
Pubblicazione: (2026)
di: Cheng, Zicong, et al.
Pubblicazione: (2026)
Atomic Self-Consistency for Better Long Form Generations
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2024)
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
di: He, Zhitao, et al.
Pubblicazione: (2025) -
Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
di: Ju, Feng, et al.
Pubblicazione: (2025) -
MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL
di: Yang, Haolin, et al.
Pubblicazione: (2025) -
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
di: He, Zhitao, et al.
Pubblicazione: (2026) -
MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?
di: He, Zhitao, et al.
Pubblicazione: (2025)