LFQA-E: Carefully Benchmarking Long-form QA Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Yuchen, Lin, Chen, Zhong, Xin, Zhang, Shuo, Zhou, Heng, Zhang, Yuchen, Liang, Mingyu, Xie, Chengxing, Hua, Ermo, Chen, Gang, He, Zhizhou, Huang, Cheng, Ding, Ning, Zhou, Bowen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EVA-Score: Evaluating Abstractive Long-form Summarization on Informativeness through Extraction and Validation
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
von: Fan, Yuchen, et al.
Veröffentlicht: (2024)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
von: Hua, Ermo, et al.
Veröffentlicht: (2024)
von: Hua, Ermo, et al.
Veröffentlicht: (2024)
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models
von: Lv, Xingtai, et al.
Veröffentlicht: (2025)
von: Lv, Xingtai, et al.
Veröffentlicht: (2025)
SSRL: Self-Search Reinforcement Learning
von: Fan, Yuchen, et al.
Veröffentlicht: (2025)
von: Fan, Yuchen, et al.
Veröffentlicht: (2025)
Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
von: Hua, Ermo, et al.
Veröffentlicht: (2024)
von: Hua, Ermo, et al.
Veröffentlicht: (2024)
FinTextQA: A Dataset for Long-form Financial Question Answering
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
von: Long, Xinwei, et al.
Veröffentlicht: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
An Impulse-formed Navier-Stokes Solver based on Long-range Particle Flow Maps
von: Li, Zhiqi, et al.
Veröffentlicht: (2026)
von: Li, Zhiqi, et al.
Veröffentlicht: (2026)
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information
von: Ping, Bowen, et al.
Veröffentlicht: (2025)
von: Ping, Bowen, et al.
Veröffentlicht: (2025)
Deflated HeteroPCA: Overcoming the curse of ill-conditioning in heteroskedastic PCA
von: Zhou, Yuchen, et al.
Veröffentlicht: (2023)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2023)
A-MapReduce: Executing Wide Search via Agentic MapReduce
von: Chen, Mingju, et al.
Veröffentlicht: (2026)
von: Chen, Mingju, et al.
Veröffentlicht: (2026)
How to Synthesize Text Data without Model Collapse?
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2024)
Inverse Flow and Consistency Models
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhang, Yuchen, et al.
Veröffentlicht: (2025)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
Unveiling Stripe-shaped Charge Density Modulations in Doped Mott Insulators
von: Xia, Ning, et al.
Veröffentlicht: (2024)
von: Xia, Ning, et al.
Veröffentlicht: (2024)
Long Arithmetic Progressions in Sumsets and Subset Sums: Constructive Proofs and Efficient Witnesses
von: Chen, Lin, et al.
Veröffentlicht: (2025)
von: Chen, Lin, et al.
Veröffentlicht: (2025)
Deep Unfolding with Kernel-based Quantization in MIMO Detection
von: Ren, Zeyi, et al.
Veröffentlicht: (2025)
von: Ren, Zeyi, et al.
Veröffentlicht: (2025)
TTRL: Test-Time Reinforcement Learning
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
von: Wu, Yuchen, et al.
Veröffentlicht: (2025)
Primes of the form $ax+by$ in certain intervals with small solutions
von: Ding, Yuchen, et al.
Veröffentlicht: (2025)
von: Ding, Yuchen, et al.
Veröffentlicht: (2025)
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2024)
Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
von: Qi, Biqing, et al.
Veröffentlicht: (2024)
A New Framework for Quantum Phases in Open Systems: Steady State of Imaginary-Time Lindbladian Evolution
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
von: Guo, Yuchen, et al.
Veröffentlicht: (2024)
Optimal Convergence Analysis of DDPM for General Distributions
von: Jiao, Yuchen, et al.
Veröffentlicht: (2025)
von: Jiao, Yuchen, et al.
Veröffentlicht: (2025)
LaMP-QA: A Benchmark for Personalized Long-form Question Answering
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
von: Salemi, Alireza, et al.
Veröffentlicht: (2025)
Exploring Data-Free LoRA Transferability for Video Diffusion Models
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
von: Wang, Yuchen, et al.
Veröffentlicht: (2026)
Non-contractible closed geodesics on compact Finsler space forms without self-intersections
von: Wang, Yuchen
Veröffentlicht: (2024)
von: Wang, Yuchen
Veröffentlicht: (2024)
Solutions to a Romanoff type problem
von: Ding, Yuchen
Veröffentlicht: (2024)
von: Ding, Yuchen
Veröffentlicht: (2024)
Solutions to some problems on unique representation bases
von: Ding, Yuchen
Veröffentlicht: (2024)
von: Ding, Yuchen
Veröffentlicht: (2024)
On a Romanoff type problem of Erdős and Kalmár
von: Ding, Yuchen
Veröffentlicht: (2025)
von: Ding, Yuchen
Veröffentlicht: (2025)
On a conjecture on shifted primes with large prime factors, II
von: Ding, Yuchen
Veröffentlicht: (2024)
von: Ding, Yuchen
Veröffentlicht: (2024)
Note on a problem of Sárközy on multiplicative representation functions
von: Ding, Yuchen
Veröffentlicht: (2025)
von: Ding, Yuchen
Veröffentlicht: (2025)
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering
von: Jahan, Rafid Ishrak, et al.
Veröffentlicht: (2026)
von: Jahan, Rafid Ishrak, et al.
Veröffentlicht: (2026)
A Claim Decomposition Benchmark for Long-form Answer Verification
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhihao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EVA-Score: Evaluating Abstractive Long-form Summarization on Informativeness through Extraction and Validation
von: Fan, Yuchen, et al.
Veröffentlicht: (2024) -
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025) -
Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization
von: Hua, Ermo, et al.
Veröffentlicht: (2024) -
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024) -
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)