FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Youquan, Zheng, Miao, Yang, Fan, Dong, Guosheng, Cui, Bin, Chen, Weipeng, Zhou, Zenan, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
von: Lu, Keer, et al.
Veröffentlicht: (2024)
von: Lu, Keer, et al.
Veröffentlicht: (2024)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
Data Proportion Detection for Optimized Data Management for Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Liu, Zhiqiang, et al.
Veröffentlicht: (2025)
SysBench: Can Large Language Models Follow System Messages?
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
von: Bai, Ge, et al.
Veröffentlicht: (2024)
von: Bai, Ge, et al.
Veröffentlicht: (2024)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs
von: Lu, Keer, et al.
Veröffentlicht: (2024)
von: Lu, Keer, et al.
Veröffentlicht: (2024)
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
von: Fouda, Aya E., et al.
Veröffentlicht: (2025)
von: Fouda, Aya E., et al.
Veröffentlicht: (2025)
RuozhiBench: Evaluating LLMs with Logical Fallacies and Misleading Premises
von: Zhai, Zenan, et al.
Veröffentlicht: (2025)
von: Zhai, Zenan, et al.
Veröffentlicht: (2025)
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
von: Zheng, Miao, et al.
Veröffentlicht: (2024)
von: Zheng, Miao, et al.
Veröffentlicht: (2024)
Baichuan Alignment Technical Report
von: Lin, Mingan, et al.
Veröffentlicht: (2024)
von: Lin, Mingan, et al.
Veröffentlicht: (2024)
ArgBench: Benchmarking LLMs on Computational Argumentation Tasks
von: Ajjour, Yamen, et al.
Veröffentlicht: (2026)
von: Ajjour, Yamen, et al.
Veröffentlicht: (2026)
Improving Multi-turn Task Completion in Task-Oriented Dialog Systems via Prompt Chaining and Fine-Grained Feedback
von: Fereidouni, Moghis, et al.
Veröffentlicht: (2025)
von: Fereidouni, Moghis, et al.
Veröffentlicht: (2025)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
von: Cao, Hongye, et al.
Veröffentlicht: (2025)
von: Cao, Hongye, et al.
Veröffentlicht: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
TDR: Task-Decoupled Retrieval with Fine-Grained LLM Feedback for In-Context Learning
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
von: Niu, Yadong, et al.
Veröffentlicht: (2025)
von: Niu, Yadong, et al.
Veröffentlicht: (2025)
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation
von: Gkoumas, Dimitris, et al.
Veröffentlicht: (2024)
von: Gkoumas, Dimitris, et al.
Veröffentlicht: (2024)
FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs
von: Labroo, Arya, et al.
Veröffentlicht: (2026)
von: Labroo, Arya, et al.
Veröffentlicht: (2026)
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
von: Ying, Jie, et al.
Veröffentlicht: (2025)
von: Ying, Jie, et al.
Veröffentlicht: (2025)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
von: Pires, Ramon, et al.
Veröffentlicht: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
FREB-TQA: A Fine-Grained Robustness Evaluation Benchmark for Table Question Answering
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios
von: Yang, Tingyue, et al.
Veröffentlicht: (2025)
von: Yang, Tingyue, et al.
Veröffentlicht: (2025)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
CCR-Bench: A Comprehensive Benchmark for Evaluating LLMs on Complex Constraints, Control Flows, and Real-World Cases
von: Xue, Xiaona, et al.
Veröffentlicht: (2026)
von: Xue, Xiaona, et al.
Veröffentlicht: (2026)
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
Med-R$^2$: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine
von: Lu, Keer, et al.
Veröffentlicht: (2025)
von: Lu, Keer, et al.
Veröffentlicht: (2025)
GeoBenchX: Benchmarking LLMs in Agent Solving Multistep Geospatial Tasks
von: Krechetova, Varvara, et al.
Veröffentlicht: (2025)
von: Krechetova, Varvara, et al.
Veröffentlicht: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in Diplomacy
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
von: Lu, Keer, et al.
Veröffentlicht: (2024) -
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
von: Chen, Mingyang, et al.
Veröffentlicht: (2024) -
Data Proportion Detection for Optimized Data Management for Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2024) -
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
von: Zhang, Tao, et al.
Veröffentlicht: (2024) -
SAS-Bench: A Fine-Grained Benchmark for Evaluating Short Answer Scoring with Large Language Models
von: Lai, Peichao, et al.
Veröffentlicht: (2025)