CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Tao, Zhu, Chenglin, Shen, Yanjun, Luo, Wenjing, Zhang, Yan, Liang, Hao, Yang, Fan, Lin, Mingan, Qiao, Yujing, Chen, Weipeng, Cui, Bin, Zhang, Wentao, Zhou, Zenan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SysBench: Can Large Language Models Follow System Messages?
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024)
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
von: Zheng, Miao, et al.
Veröffentlicht: (2024)
von: Zheng, Miao, et al.
Veröffentlicht: (2024)
K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
von: Zhu, Chenglin, et al.
Veröffentlicht: (2025)
von: Zhu, Chenglin, et al.
Veröffentlicht: (2025)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)
von: Li, Youquan, et al.
Veröffentlicht: (2024)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
von: Chen, Mingyang, et al.
Veröffentlicht: (2024)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
von: Lu, Keer, et al.
Veröffentlicht: (2024)
von: Lu, Keer, et al.
Veröffentlicht: (2024)
Baichuan Alignment Technical Report
von: Lin, Mingan, et al.
Veröffentlicht: (2024)
von: Lin, Mingan, et al.
Veröffentlicht: (2024)
Data Proportion Detection for Optimized Data Management for Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Ocean-OCR: Towards General OCR Application via a Vision-Language Model
von: Chen, Song, et al.
Veröffentlicht: (2025)
von: Chen, Song, et al.
Veröffentlicht: (2025)
Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
von: Zhao, Yaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Yaqi, et al.
Veröffentlicht: (2024)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline
von: Dong, Guosheng, et al.
Veröffentlicht: (2024)
von: Dong, Guosheng, et al.
Veröffentlicht: (2024)
Fundamental Limits of Pulse Based UWB ISAC Systems: A Parameter Estimation Perspective
von: Liu, Fan, et al.
Veröffentlicht: (2024)
von: Liu, Fan, et al.
Veröffentlicht: (2024)
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
von: Feng, Xiang, et al.
Veröffentlicht: (2025)
von: Feng, Xiang, et al.
Veröffentlicht: (2025)
Towards Robust Sensor-Fusion Ground SLAM: A Comprehensive Benchmark and A Resilient Framework
von: Zhang, Deteng, et al.
Veröffentlicht: (2025)
von: Zhang, Deteng, et al.
Veröffentlicht: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
DARO: Difficulty-Aware Reweighting Policy Optimization
von: Zhou, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhou, Jingyu, et al.
Veröffentlicht: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
Baichuan-Omni Technical Report
von: Li, Yadong, et al.
Veröffentlicht: (2024)
von: Li, Yadong, et al.
Veröffentlicht: (2024)
GRADE: Probing Knowledge Gaps in LLMs through Gradient Subspace Dynamics
von: Wang, Yujing, et al.
Veröffentlicht: (2026)
von: Wang, Yujing, et al.
Veröffentlicht: (2026)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
CCTU: A Benchmark for Tool Use under Complex Constraints
von: Ye, Junjie, et al.
Veröffentlicht: (2026)
von: Ye, Junjie, et al.
Veröffentlicht: (2026)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
von: Guo, Tianyu, et al.
Veröffentlicht: (2025)
FusionBench: A Unified Library and Comprehensive Benchmark for Deep Model Fusion
von: Tang, Anke, et al.
Veröffentlicht: (2024)
von: Tang, Anke, et al.
Veröffentlicht: (2024)
Dissipated Correction Map Method with Trapezoidal Rule for the Simulations of Gravitational Waves from Spinning Compact Binary
von: Luo, Junjie, et al.
Veröffentlicht: (2024)
von: Luo, Junjie, et al.
Veröffentlicht: (2024)
LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
von: Wong, Zhen Hao, et al.
Veröffentlicht: (2025)
Optimization of Private Semantic Communication Performance: An Uncooperative Covert Communication Method
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2025)
VABench: A Comprehensive Benchmark for Audio-Video Generation
von: Hua, Daili, et al.
Veröffentlicht: (2025)
von: Hua, Daili, et al.
Veröffentlicht: (2025)
Intention Analysis Makes LLMs A Good Jailbreak Defender
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuqi, et al.
Veröffentlicht: (2024)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
von: Sun, Linzhuang, et al.
Veröffentlicht: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
EDU-MATRIX: A Society-Centric Generative Cognitive Digital Twin Architecture for Secondary Education
von: Zhai, Wenjing, et al.
Veröffentlicht: (2026)
von: Zhai, Wenjing, et al.
Veröffentlicht: (2026)
Comprehensive Analysis of Network Robustness Evaluation Based on Convolutional Neural Networks with Spatial Pyramid Pooling
von: Jiang, Wenjun, et al.
Veröffentlicht: (2023)
von: Jiang, Wenjun, et al.
Veröffentlicht: (2023)
BEATS: Optimizing LLM Mathematical Capabilities with BackVerify and Adaptive Disambiguate based Efficient Tree Search
von: Sun, Linzhuang, et al.
Veröffentlicht: (2024)
von: Sun, Linzhuang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SysBench: Can Large Language Models Follow System Messages?
von: Qin, Yanzhao, et al.
Veröffentlicht: (2024) -
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System
von: Zheng, Miao, et al.
Veröffentlicht: (2024) -
K12Vista: Exploring the Boundaries of MLLMs in K-12 Education
von: Li, Chong, et al.
Veröffentlicht: (2025) -
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
von: Zhu, Chenglin, et al.
Veröffentlicht: (2025) -
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)