FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
Fuente:
arXiv
Salvato in:
| Autori principali: | Xia, Congying, Xing, Chen, Du, Jiangshu, Yang, Xinyi, Feng, Yihao, Xu, Ran, Yin, Wenpeng, Xiong, Caiming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ScaleFormer: Span Representation Cumulation for Long-Context Transformer
di: Du, Jiangshu, et al.
Pubblicazione: (2025)
di: Du, Jiangshu, et al.
Pubblicazione: (2025)
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
Preference-grounded Token-level Guidance for Language Model Fine-tuning
di: Yang, Shentao, et al.
Pubblicazione: (2023)
di: Yang, Shentao, et al.
Pubblicazione: (2023)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
di: Wang, Yibo, et al.
Pubblicazione: (2025)
di: Wang, Yibo, et al.
Pubblicazione: (2025)
LLMs' Classification Performance is Overclaimed
di: Xu, Hanzi, et al.
Pubblicazione: (2024)
di: Xu, Hanzi, et al.
Pubblicazione: (2024)
Could AI Trace and Explain the Origins of AI-Generated Images and Text?
di: Fang, Hongchao, et al.
Pubblicazione: (2025)
di: Fang, Hongchao, et al.
Pubblicazione: (2025)
Toward Zero-Shot Instruction Following
di: Lou, Renze, et al.
Pubblicazione: (2023)
di: Lou, Renze, et al.
Pubblicazione: (2023)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Benchmarking LLMs for Political Science: A United Nations Perspective
di: Liang, Yueqing, et al.
Pubblicazione: (2025)
di: Liang, Yueqing, et al.
Pubblicazione: (2025)
Large Language Model Instruction Following: A Survey of Progresses and Challenges
di: Lou, Renze, et al.
Pubblicazione: (2023)
di: Lou, Renze, et al.
Pubblicazione: (2023)
UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG
di: Peng, Xiangyu, et al.
Pubblicazione: (2025)
di: Peng, Xiangyu, et al.
Pubblicazione: (2025)
AAAR-1.0: Assessing AI's Potential to Assist Research
di: Lou, Renze, et al.
Pubblicazione: (2024)
di: Lou, Renze, et al.
Pubblicazione: (2024)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
di: Liu, Hongtao, et al.
Pubblicazione: (2025)
di: Liu, Hongtao, et al.
Pubblicazione: (2025)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
di: Pakhomov, Egor, et al.
Pubblicazione: (2025)
di: Pakhomov, Egor, et al.
Pubblicazione: (2025)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
di: Wang, Lei, et al.
Pubblicazione: (2024)
di: Wang, Lei, et al.
Pubblicazione: (2024)
Shared Imagination: LLMs Hallucinate Alike
di: Zhou, Yilun, et al.
Pubblicazione: (2024)
di: Zhou, Yilun, et al.
Pubblicazione: (2024)
Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
di: wang, Jiaming, et al.
Pubblicazione: (2025)
di: wang, Jiaming, et al.
Pubblicazione: (2025)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
di: Zhang, Shu, et al.
Pubblicazione: (2023)
di: Zhang, Shu, et al.
Pubblicazione: (2023)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
Format-Adapter: Improving Reasoning Capability of LLMs by Adapting Suitable Format
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
Text2Data: Low-Resource Data Generation with Textual Control
di: Wang, Shiyu, et al.
Pubblicazione: (2024)
di: Wang, Shiyu, et al.
Pubblicazione: (2024)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
di: Zhang, Tao, et al.
Pubblicazione: (2024)
di: Zhang, Tao, et al.
Pubblicazione: (2024)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
di: Chen, Hailin, et al.
Pubblicazione: (2024)
di: Chen, Hailin, et al.
Pubblicazione: (2024)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
di: Li, Haoming, et al.
Pubblicazione: (2025)
di: Li, Haoming, et al.
Pubblicazione: (2025)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
di: Liu, Chuang, et al.
Pubblicazione: (2024)
di: Liu, Chuang, et al.
Pubblicazione: (2024)
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
di: Li, Ce, et al.
Pubblicazione: (2025)
di: Li, Ce, et al.
Pubblicazione: (2025)
Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis
di: Bhowmik, Shimanto, et al.
Pubblicazione: (2025)
di: Bhowmik, Shimanto, et al.
Pubblicazione: (2025)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
di: Son, Guijin, et al.
Pubblicazione: (2026)
di: Son, Guijin, et al.
Pubblicazione: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
di: Wang, Lei, et al.
Pubblicazione: (2024)
di: Wang, Lei, et al.
Pubblicazione: (2024)
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
di: Chen, Haolin, et al.
Pubblicazione: (2024)
di: Chen, Haolin, et al.
Pubblicazione: (2024)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
di: Laban, Philippe, et al.
Pubblicazione: (2023)
di: Laban, Philippe, et al.
Pubblicazione: (2023)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
di: Laban, Philippe, et al.
Pubblicazione: (2024)
di: Laban, Philippe, et al.
Pubblicazione: (2024)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
di: He, Yun, et al.
Pubblicazione: (2024)
di: He, Yun, et al.
Pubblicazione: (2024)
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
di: Oh, Hongseok, et al.
Pubblicazione: (2025)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
di: Singh, Harman, et al.
Pubblicazione: (2024)
di: Singh, Harman, et al.
Pubblicazione: (2024)
Unanswerability Evaluation for Retrieval Augmented Generation
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
di: Peng, Xiangyu, et al.
Pubblicazione: (2024)
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
di: Lin, Jingru, et al.
Pubblicazione: (2025)
di: Lin, Jingru, et al.
Pubblicazione: (2025)
Prompt-Reverse Inconsistency: LLM Self-Inconsistency Beyond Generative Randomness and Prompt Paraphrasing
di: Ahn, Jihyun Janice, et al.
Pubblicazione: (2025)
di: Ahn, Jihyun Janice, et al.
Pubblicazione: (2025)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
di: Xu, Hainiu, et al.
Pubblicazione: (2024)
di: Xu, Hainiu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ScaleFormer: Span Representation Cumulation for Long-Context Transformer
di: Du, Jiangshu, et al.
Pubblicazione: (2025) -
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
di: Peng, Xiangyu, et al.
Pubblicazione: (2024) -
Preference-grounded Token-level Guidance for Language Model Fine-tuning
di: Yang, Shentao, et al.
Pubblicazione: (2023) -
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
di: Wang, Yibo, et al.
Pubblicazione: (2025) -
LLMs' Classification Performance is Overclaimed
di: Xu, Hanzi, et al.
Pubblicazione: (2024)