ReIFE: Re-evaluating Instruction-Following Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yixin, Shi, Kejian, Fabbri, Alexander R., Zhao, Yilun, Wang, Peifeng, Wu, Chien-Sheng, Joty, Shafiq, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
von: Phanse, Rohan, et al.
Veröffentlicht: (2025)
von: Phanse, Rohan, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
On Learning to Summarize with Large Language Models as References
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
von: Liu, Yixin, et al.
Veröffentlicht: (2023)
Unsupervised Summarization Re-ranking
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2022)
von: Ravaut, Mathieu, et al.
Veröffentlicht: (2022)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
Direct Judgement Preference Optimization
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
von: Wang, Peifeng, et al.
Veröffentlicht: (2024)
Prompt Leakage effect and defense strategies for multi-turn LLM interactions
von: Agarwal, Divyansh, et al.
Veröffentlicht: (2024)
von: Agarwal, Divyansh, et al.
Veröffentlicht: (2024)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2023)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2023)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
von: Liu, Gabrielle Kaili-May, et al.
Veröffentlicht: (2024)
von: Liu, Gabrielle Kaili-May, et al.
Veröffentlicht: (2024)
Understanding Reference Policies in Direct Preference Optimization
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2024)
SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
von: Wang, Chengye, et al.
Veröffentlicht: (2025)
von: Wang, Chengye, et al.
Veröffentlicht: (2025)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
Table-R1: Inference-Time Scaling for Table Reasoning
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research Papers
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
von: Xu, Zhijian, et al.
Veröffentlicht: (2025)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
von: Weller, Orion, et al.
Veröffentlicht: (2024)
von: Weller, Orion, et al.
Veröffentlicht: (2024)
Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
von: Hu, Tiansheng, et al.
Veröffentlicht: (2025)
von: Hu, Tiansheng, et al.
Veröffentlicht: (2025)
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
von: Zhao, Yilun, et al.
Veröffentlicht: (2026)
Z1: Efficient Test-time Scaling with Code
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
von: Yu, Zhaojian, et al.
Veröffentlicht: (2025)
LimRank: Less is More for Reasoning-Intensive Information Reranking
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
von: Hu, Tiansheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiansheng, et al.
Veröffentlicht: (2026)
Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL
von: Shen, Yifei, et al.
Veröffentlicht: (2026)
von: Shen, Yifei, et al.
Veröffentlicht: (2026)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
von: Zhang, Nan, et al.
Veröffentlicht: (2024)
von: Zhang, Nan, et al.
Veröffentlicht: (2024)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
von: Wadden, David, et al.
Veröffentlicht: (2024)
von: Wadden, David, et al.
Veröffentlicht: (2024)
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature
von: Ding, Hang, et al.
Veröffentlicht: (2025)
von: Ding, Hang, et al.
Veröffentlicht: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
von: Han, Simeng, et al.
Veröffentlicht: (2024)
von: Han, Simeng, et al.
Veröffentlicht: (2024)
ChartInstruct: Instruction Tuning for Chart Comprehension and Reasoning
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
von: Masry, Ahmed, et al.
Veröffentlicht: (2024)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
von: Li, Alan, et al.
Veröffentlicht: (2025)
von: Li, Alan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026) -
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization
von: Liu, Yixin, et al.
Veröffentlicht: (2023) -
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
von: Phanse, Rohan, et al.
Veröffentlicht: (2025) -
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025) -
On Learning to Summarize with Large Language Models as References
von: Liu, Yixin, et al.
Veröffentlicht: (2023)