LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Siwei, Li, Yizhi, Qu, Xingwei, Ravikumar, Rishi, Li, Yucheng, Loakman, Tyler, Quan, Shanghaoran, Wei, Xiaoyong, Batista-Navarro, Riza, Lin, Chenghua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
di: Qu, Xingwei, et al.
Pubblicazione: (2024)
di: Qu, Xingwei, et al.
Pubblicazione: (2024)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
di: Cancellieri, Matteo, et al.
Pubblicazione: (2025)
di: Cancellieri, Matteo, et al.
Pubblicazione: (2025)
Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
Train & Constrain: Phonologically Informed Tongue-Twister Generation from Topics and Paraphrases
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
di: Loakman, Tyler, et al.
Pubblicazione: (2024)
Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
di: Loakman, Tyler, et al.
Pubblicazione: (2025)
LongEval at CLEF 2025: Longitudinal Evaluation of IR Systems on Web and Scientific Data
di: Cancellieri, Matteo, et al.
Pubblicazione: (2025)
di: Cancellieri, Matteo, et al.
Pubblicazione: (2025)
DS@GT at LongEval: Evaluating Temporal Performance in Web Search Systems and Topics with Two-Stage Retrieval
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2025)
di: Miyaguchi, Anthony, et al.
Pubblicazione: (2025)
DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling
di: Quan, Shanghaoran
Pubblicazione: (2024)
di: Quan, Shanghaoran
Pubblicazione: (2024)
Automatically Generating Numerous Context-Driven SFT Data for LLMs across Diverse Granularity
di: Quan, Shanghaoran
Pubblicazione: (2024)
di: Quan, Shanghaoran
Pubblicazione: (2024)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
di: Li, Yucheng, et al.
Pubblicazione: (2023)
di: Li, Yucheng, et al.
Pubblicazione: (2023)
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
di: Ren, Jincheng, et al.
Pubblicazione: (2026)
di: Ren, Jincheng, et al.
Pubblicazione: (2026)
Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments
di: Wu, Siwei, et al.
Pubblicazione: (2026)
di: Wu, Siwei, et al.
Pubblicazione: (2026)
Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study
di: Sun, Yizheng, et al.
Pubblicazione: (2025)
di: Sun, Yizheng, et al.
Pubblicazione: (2025)
Language Models can Self-Lengthen to Generate Long Texts
di: Quan, Shanghaoran, et al.
Pubblicazione: (2024)
di: Quan, Shanghaoran, et al.
Pubblicazione: (2024)
CADGE: Context-Aware Dialogue Generation Enhanced with Graph-Structured Knowledge Aggregation
di: Zhang, Hongbo, et al.
Pubblicazione: (2023)
di: Zhang, Hongbo, et al.
Pubblicazione: (2023)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
di: Wang, Shun, et al.
Pubblicazione: (2024)
di: Wang, Shun, et al.
Pubblicazione: (2024)
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
di: Wu, Yulong, et al.
Pubblicazione: (2025)
di: Wu, Yulong, et al.
Pubblicazione: (2025)
Investigating a Benchmark for Training-set free Evaluation of Linguistic Capabilities in Machine Reading Comprehension
di: Schlegel, Viktor, et al.
Pubblicazione: (2024)
di: Schlegel, Viktor, et al.
Pubblicazione: (2024)
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
di: Sun, Yizheng, et al.
Pubblicazione: (2025)
di: Sun, Yizheng, et al.
Pubblicazione: (2025)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
di: Li, Zirui, et al.
Pubblicazione: (2025)
di: Li, Zirui, et al.
Pubblicazione: (2025)
Aspect-based Sentiment Evaluation of Chess Moves (ASSESS): an NLP-based Method for Evaluating Chess Strategies from Textbooks
di: Alrdahi, Haifa, et al.
Pubblicazione: (2024)
di: Alrdahi, Haifa, et al.
Pubblicazione: (2024)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
di: Ma, David, et al.
Pubblicazione: (2025)
di: Ma, David, et al.
Pubblicazione: (2025)
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
di: Wang, Yang, et al.
Pubblicazione: (2025)
di: Wang, Yang, et al.
Pubblicazione: (2025)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
di: Wang, Shun, et al.
Pubblicazione: (2025)
di: Wang, Shun, et al.
Pubblicazione: (2025)
An Open Source Data Contamination Report for Large Language Models
di: Li, Yucheng, et al.
Pubblicazione: (2023)
di: Li, Yucheng, et al.
Pubblicazione: (2023)
Finding Challenging Metaphors that Confuse Pretrained Language Models
di: Li, Yucheng, et al.
Pubblicazione: (2024)
di: Li, Yucheng, et al.
Pubblicazione: (2024)
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models
di: Wu, Siwei, et al.
Pubblicazione: (2024)
di: Wu, Siwei, et al.
Pubblicazione: (2024)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
di: Wu, Yulong, et al.
Pubblicazione: (2025)
di: Wu, Yulong, et al.
Pubblicazione: (2025)
Observing Micromotives and Macrobehavior of Large Language Models
di: Cheng, Yuyang, et al.
Pubblicazione: (2024)
di: Cheng, Yuyang, et al.
Pubblicazione: (2024)
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
di: Madusanka, Tharindu, et al.
Pubblicazione: (2025)
di: Madusanka, Tharindu, et al.
Pubblicazione: (2025)
LLM-State: Open World State Representation for Long-horizon Task Planning with Large Language Model
di: Chen, Siwei, et al.
Pubblicazione: (2023)
di: Chen, Siwei, et al.
Pubblicazione: (2023)
Large Language Models in Argument Mining: A Survey
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Integrating Planning into Single-Turn Long-Form Text Generation
di: Liang, Yi, et al.
Pubblicazione: (2024)
di: Liang, Yi, et al.
Pubblicazione: (2024)
Evaluating Large Language Models for Generalization and Robustness via Data Compression
di: Li, Yucheng, et al.
Pubblicazione: (2024)
di: Li, Yucheng, et al.
Pubblicazione: (2024)
On the Rigour of Scientific Writing: Criteria, Analysis, and Insights
di: James, Joseph, et al.
Pubblicazione: (2024)
di: James, Joseph, et al.
Pubblicazione: (2024)
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
di: Liang, Yiming, et al.
Pubblicazione: (2024)
di: Liang, Yiming, et al.
Pubblicazione: (2024)
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
di: Li, Yizhi, et al.
Pubblicazione: (2025)
di: Li, Yizhi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
di: Loakman, Tyler, et al.
Pubblicazione: (2024) -
ReproHum #0087-01: Human Evaluation Reproduction Report for Generating Fact Checking Explanations
di: Loakman, Tyler, et al.
Pubblicazione: (2024) -
Overview of the NLPCC 2024 Shared Task on Chinese Metaphor Generation
di: Qu, Xingwei, et al.
Pubblicazione: (2024) -
LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
di: Cancellieri, Matteo, et al.
Pubblicazione: (2025) -
Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
di: Loakman, Tyler, et al.
Pubblicazione: (2025)