Gespeichert in:
| Hauptverfasser: | Gao, Yicheng, Xu, Gonghan, Wang, Zhe, Cohan, Arman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2411.04424 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Calibrating Long-form Generations from Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2024)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2024)
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
von: Lee, Jinu, et al.
Veröffentlicht: (2025)
von: Lee, Jinu, et al.
Veröffentlicht: (2025)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
von: Zhao, Yilun, et al.
Veröffentlicht: (2025)
SUCEA: Reasoning-Intensive Retrieval for Adversarial Fact-checking through Claim Decomposition and Editing
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
Survey on Evaluation of LLM-based Agents
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
von: Garikaparthi, Aniketh, et al.
Veröffentlicht: (2025)
von: Garikaparthi, Aniketh, et al.
Veröffentlicht: (2025)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
von: Wang, Benlu, et al.
Veröffentlicht: (2025)
von: Wang, Benlu, et al.
Veröffentlicht: (2025)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
von: Wu, Sihong, et al.
Veröffentlicht: (2026)
von: Wu, Sihong, et al.
Veröffentlicht: (2026)
LocAgent: Graph-Guided LLM Agents for Code Localization
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
von: Garikaparthi, Aniketh, et al.
Veröffentlicht: (2025)
von: Garikaparthi, Aniketh, et al.
Veröffentlicht: (2025)
ReIFE: Re-evaluating Instruction-Following Evaluation
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2024)
SciMDR: Advancing Scientific Multimodal Document Reasoning
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
von: Chen, Ziyu, et al.
Veröffentlicht: (2026)
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
von: Wu, Sihong, et al.
Veröffentlicht: (2026)
von: Wu, Sihong, et al.
Veröffentlicht: (2026)
ToolACE: Winning the Points of LLM Function Calling
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2024)
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
von: Liu, Yixin, et al.
Veröffentlicht: (2024)
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
von: Tang, Xiangru, et al.
Veröffentlicht: (2023)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
von: Zhao, Haochen, et al.
Veröffentlicht: (2024)
von: Zhao, Haochen, et al.
Veröffentlicht: (2024)
MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments
von: Gao, Yicheng, et al.
Veröffentlicht: (2026)
von: Gao, Yicheng, et al.
Veröffentlicht: (2026)
Step-Back Profiling: Distilling User History for Personalized Scientific Writing
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
von: Tang, Xiangru, et al.
Veröffentlicht: (2024)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
von: Long, Yitao, et al.
Veröffentlicht: (2025)
von: Long, Yitao, et al.
Veröffentlicht: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
When do Generative Query and Document Expansions Fail? A Comprehensive Study Across Methods, Retrievers, and Datasets
von: Weller, Orion, et al.
Veröffentlicht: (2023)
von: Weller, Orion, et al.
Veröffentlicht: (2023)
Graphical Reasoning: LLM-based Semi-Open Relation Extraction
von: Tao, Yicheng, et al.
Veröffentlicht: (2024)
von: Tao, Yicheng, et al.
Veröffentlicht: (2024)
MemSim: A Bayesian Simulator for Evaluating Memory of LLM-based Personal Assistants
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains
von: Han, Simeng, et al.
Veröffentlicht: (2024)
von: Han, Simeng, et al.
Veröffentlicht: (2024)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
von: Zhao, Junjie, et al.
Veröffentlicht: (2026)
von: Zhao, Junjie, et al.
Veröffentlicht: (2026)
Mini-Giants: "Small" Language Models and Open Source Win-Win
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2024)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
von: Chen, Sanxing, et al.
Veröffentlicht: (2025)
von: Chen, Sanxing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025) -
Calibrating Long-form Generations from Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2024) -
On the Benefits of Fine-Grained Loss Truncation: A Case Study on Factuality in Summarization
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2024) -
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
von: Lee, Jinu, et al.
Veröffentlicht: (2025) -
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
von: Gao, Mingqi, et al.
Veröffentlicht: (2024)