CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Feng, Pang, Chengjie, Zhang, Yuehan, Luo, Chenyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMzSzŁ: a comprehensive LLM benchmark for Polish
by: Jassem, Krzysztof, et al.
Published: (2025)
by: Jassem, Krzysztof, et al.
Published: (2025)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
by: Liu, Shuyu, et al.
Published: (2025)
by: Liu, Shuyu, et al.
Published: (2025)
Harnessing Deep LLM Participation for Robust Entity Linking
by: Hou, Jiajun, et al.
Published: (2025)
by: Hou, Jiajun, et al.
Published: (2025)
Construction of a Syntactic Analysis Map for Yi Shui School through Text Mining and Natural Language Processing Research
by: Zhao, Hanqing, et al.
Published: (2024)
by: Zhao, Hanqing, et al.
Published: (2024)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
by: Ailem, Melissa, et al.
Published: (2024)
by: Ailem, Melissa, et al.
Published: (2024)
Applications of natural language processing in aviation safety: A review and qualitative analysis
by: Nanyonga, Aziida, et al.
Published: (2025)
by: Nanyonga, Aziida, et al.
Published: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
by: Nguyen, Nam V., et al.
Published: (2024)
by: Nguyen, Nam V., et al.
Published: (2024)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
by: Zhang, Yuehan, et al.
Published: (2024)
by: Zhang, Yuehan, et al.
Published: (2024)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
by: Wen, Zichen, et al.
Published: (2024)
by: Wen, Zichen, et al.
Published: (2024)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
by: Wu, Jian, et al.
Published: (2024)
by: Wu, Jian, et al.
Published: (2024)
Dynamic benchmarking framework for LLM-based conversational data capture
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
by: Aluffi, Pietro Alessandro, et al.
Published: (2025)
Pushing The Limit of LLM Capacity for Text Classification
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
by: Kankowski, Florian, et al.
Published: (2025)
by: Kankowski, Florian, et al.
Published: (2025)
Evaluating Zero-Shot Long-Context LLM Compression
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
by: Wei, Jiaqi, et al.
Published: (2026)
by: Wei, Jiaqi, et al.
Published: (2026)
Visualizing attention zones in machine reading comprehension models
by: Cui, Yiming, et al.
Published: (2024)
by: Cui, Yiming, et al.
Published: (2024)
DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
by: Wu, Junchao, et al.
Published: (2026)
by: Wu, Junchao, et al.
Published: (2026)
Linguini: A benchmark for language-agnostic linguistic reasoning
by: Sánchez, Eduardo, et al.
Published: (2024)
by: Sánchez, Eduardo, et al.
Published: (2024)
LLM4Decompile: Decompiling Binary Code with Large Language Models
by: Tan, Hanzhuo, et al.
Published: (2024)
by: Tan, Hanzhuo, et al.
Published: (2024)
A multilingual hallucination benchmark: MultiWikiQHalluA
by: Thoresen, Freja, et al.
Published: (2026)
by: Thoresen, Freja, et al.
Published: (2026)
HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?
by: Peng, Weihan, et al.
Published: (2026)
by: Peng, Weihan, et al.
Published: (2026)
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
by: Pacchiardi, Lorenzo, et al.
Published: (2024)
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
by: Zhang, Zekai, et al.
Published: (2024)
by: Zhang, Zekai, et al.
Published: (2024)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
SiLLM: Large Language Models for Simultaneous Machine Translation
by: Guo, Shoutao, et al.
Published: (2024)
by: Guo, Shoutao, et al.
Published: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
by: Wang, Chonghua, et al.
Published: (2024)
by: Wang, Chonghua, et al.
Published: (2024)
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
by: Yan, Jianhao, et al.
Published: (2025)
by: Yan, Jianhao, et al.
Published: (2025)
ElasticMem: Latent Memory as a Learnable Resource for LLM Agents
by: Feng, Tao, et al.
Published: (2026)
by: Feng, Tao, et al.
Published: (2026)
Counting Ability of Large Language Models and Impact of Tokenization
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning
by: Cui, Ziang, et al.
Published: (2026)
by: Cui, Ziang, et al.
Published: (2026)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
by: Min, Junghyun, et al.
Published: (2025)
by: Min, Junghyun, et al.
Published: (2025)
Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMs
by: Zhang, Xiang, et al.
Published: (2025)
by: Zhang, Xiang, et al.
Published: (2025)
Rethinking Verification for LLM Code Generation: From Generation to Testing
by: Ma, Zihan, et al.
Published: (2025)
by: Ma, Zihan, et al.
Published: (2025)
An Expert-grounded benchmark of General Purpose LLMs in LCA
by: Donaldson, Artur, et al.
Published: (2025)
by: Donaldson, Artur, et al.
Published: (2025)
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
by: Mou, Yutao, et al.
Published: (2025)
by: Mou, Yutao, et al.
Published: (2025)
Mathematical Reasoning Enhanced LLM for Formula Derivation: A Case Study on Fiber NLI Modellin
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
Annotation alignment: Comparing LLM and human annotations of conversational safety
by: Movva, Rajiv, et al.
Published: (2024)
by: Movva, Rajiv, et al.
Published: (2024)
Ensembling Large Language Models to Characterize Affective Dynamics in Student-AI Tutor Dialogues
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Anchor function: a type of benchmark functions for studying language models
by: Zhang, Zhongwang, et al.
Published: (2024)
by: Zhang, Zhongwang, et al.
Published: (2024)
Similar Items
-
LLMzSzŁ: a comprehensive LLM benchmark for Polish
by: Jassem, Krzysztof, et al.
Published: (2025) -
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
by: Liu, Shuyu, et al.
Published: (2025) -
Harnessing Deep LLM Participation for Robust Entity Linking
by: Hou, Jiajun, et al.
Published: (2025) -
Construction of a Syntactic Analysis Map for Yi Shui School through Text Mining and Natural Language Processing Research
by: Zhao, Hanqing, et al.
Published: (2024) -
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
by: Ailem, Melissa, et al.
Published: (2024)