CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Feng, Pang, Chengjie, Zhang, Yuehan, Luo, Chenyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLMzSzŁ: a comprehensive LLM benchmark for Polish
por: Jassem, Krzysztof, et al.
Publicado: (2025)
por: Jassem, Krzysztof, et al.
Publicado: (2025)
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
por: Liu, Shuyu, et al.
Publicado: (2025)
por: Liu, Shuyu, et al.
Publicado: (2025)
Harnessing Deep LLM Participation for Robust Entity Linking
por: Hou, Jiajun, et al.
Publicado: (2025)
por: Hou, Jiajun, et al.
Publicado: (2025)
Construction of a Syntactic Analysis Map for Yi Shui School through Text Mining and Natural Language Processing Research
por: Zhao, Hanqing, et al.
Publicado: (2024)
por: Zhao, Hanqing, et al.
Publicado: (2024)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
por: Ailem, Melissa, et al.
Publicado: (2024)
por: Ailem, Melissa, et al.
Publicado: (2024)
Applications of natural language processing in aviation safety: A review and qualitative analysis
por: Nanyonga, Aziida, et al.
Publicado: (2025)
por: Nanyonga, Aziida, et al.
Publicado: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
por: Nguyen, Nam V., et al.
Publicado: (2024)
por: Nguyen, Nam V., et al.
Publicado: (2024)
PersonaMark: Personalized LLM watermarking for model protection and user attribution
por: Zhang, Yuehan, et al.
Publicado: (2024)
por: Zhang, Yuehan, et al.
Publicado: (2024)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
por: Wen, Zichen, et al.
Publicado: (2024)
por: Wen, Zichen, et al.
Publicado: (2024)
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark
por: Wu, Jian, et al.
Publicado: (2024)
por: Wu, Jian, et al.
Publicado: (2024)
Dynamic benchmarking framework for LLM-based conversational data capture
por: Aluffi, Pietro Alessandro, et al.
Publicado: (2025)
por: Aluffi, Pietro Alessandro, et al.
Publicado: (2025)
Pushing The Limit of LLM Capacity for Text Classification
por: Zhang, Yazhou, et al.
Publicado: (2024)
por: Zhang, Yazhou, et al.
Publicado: (2024)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
por: Kankowski, Florian, et al.
Publicado: (2025)
por: Kankowski, Florian, et al.
Publicado: (2025)
Evaluating Zero-Shot Long-Context LLM Compression
por: Wang, Chenyu, et al.
Publicado: (2024)
por: Wang, Chenyu, et al.
Publicado: (2024)
When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning
por: Wei, Jiaqi, et al.
Publicado: (2026)
por: Wei, Jiaqi, et al.
Publicado: (2026)
Visualizing attention zones in machine reading comprehension models
por: Cui, Yiming, et al.
Publicado: (2024)
por: Cui, Yiming, et al.
Publicado: (2024)
DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
por: Wu, Junchao, et al.
Publicado: (2026)
por: Wu, Junchao, et al.
Publicado: (2026)
Linguini: A benchmark for language-agnostic linguistic reasoning
por: Sánchez, Eduardo, et al.
Publicado: (2024)
por: Sánchez, Eduardo, et al.
Publicado: (2024)
LLM4Decompile: Decompiling Binary Code with Large Language Models
por: Tan, Hanzhuo, et al.
Publicado: (2024)
por: Tan, Hanzhuo, et al.
Publicado: (2024)
A multilingual hallucination benchmark: MultiWikiQHalluA
por: Thoresen, Freja, et al.
Publicado: (2026)
por: Thoresen, Freja, et al.
Publicado: (2026)
HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?
por: Peng, Weihan, et al.
Publicado: (2026)
por: Peng, Weihan, et al.
Publicado: (2026)
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis
por: Zhang, Bohan, et al.
Publicado: (2025)
por: Zhang, Bohan, et al.
Publicado: (2025)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion
por: Zhang, Zekai, et al.
Publicado: (2024)
por: Zhang, Zekai, et al.
Publicado: (2024)
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
por: Wang, Jingxing, et al.
Publicado: (2026)
por: Wang, Jingxing, et al.
Publicado: (2026)
SiLLM: Large Language Models for Simultaneous Machine Translation
por: Guo, Shoutao, et al.
Publicado: (2024)
por: Guo, Shoutao, et al.
Publicado: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
por: Wang, Chonghua, et al.
Publicado: (2024)
por: Wang, Chonghua, et al.
Publicado: (2024)
RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction
por: Yan, Jianhao, et al.
Publicado: (2025)
por: Yan, Jianhao, et al.
Publicado: (2025)
ElasticMem: Latent Memory as a Learnable Resource for LLM Agents
por: Feng, Tao, et al.
Publicado: (2026)
por: Feng, Tao, et al.
Publicado: (2026)
Counting Ability of Large Language Models and Impact of Tokenization
por: Zhang, Xiang, et al.
Publicado: (2024)
por: Zhang, Xiang, et al.
Publicado: (2024)
HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning
por: Cui, Ziang, et al.
Publicado: (2026)
por: Cui, Ziang, et al.
Publicado: (2026)
\textsc{CantoNLU}: A benchmark for Cantonese natural language understanding
por: Min, Junghyun, et al.
Publicado: (2025)
por: Min, Junghyun, et al.
Publicado: (2025)
Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMs
por: Zhang, Xiang, et al.
Publicado: (2025)
por: Zhang, Xiang, et al.
Publicado: (2025)
Rethinking Verification for LLM Code Generation: From Generation to Testing
por: Ma, Zihan, et al.
Publicado: (2025)
por: Ma, Zihan, et al.
Publicado: (2025)
An Expert-grounded benchmark of General Purpose LLMs in LCA
por: Donaldson, Artur, et al.
Publicado: (2025)
por: Donaldson, Artur, et al.
Publicado: (2025)
SaRO: Enhancing LLM Safety through Reasoning-based Alignment
por: Mou, Yutao, et al.
Publicado: (2025)
por: Mou, Yutao, et al.
Publicado: (2025)
Mathematical Reasoning Enhanced LLM for Formula Derivation: A Case Study on Fiber NLI Modellin
por: Zhang, Yao, et al.
Publicado: (2026)
por: Zhang, Yao, et al.
Publicado: (2026)
Annotation alignment: Comparing LLM and human annotations of conversational safety
por: Movva, Rajiv, et al.
Publicado: (2024)
por: Movva, Rajiv, et al.
Publicado: (2024)
Ensembling Large Language Models to Characterize Affective Dynamics in Student-AI Tutor Dialogues
por: Zhang, Chenyu, et al.
Publicado: (2025)
por: Zhang, Chenyu, et al.
Publicado: (2025)
Anchor function: a type of benchmark functions for studying language models
por: Zhang, Zhongwang, et al.
Publicado: (2024)
por: Zhang, Zhongwang, et al.
Publicado: (2024)
Ejemplares similares
-
LLMzSzŁ: a comprehensive LLM benchmark for Polish
por: Jassem, Krzysztof, et al.
Publicado: (2025) -
PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice
por: Liu, Shuyu, et al.
Publicado: (2025) -
Harnessing Deep LLM Participation for Robust Entity Linking
por: Hou, Jiajun, et al.
Publicado: (2025) -
Construction of a Syntactic Analysis Map for Yi Shui School through Text Mining and Natural Language Processing Research
por: Zhao, Hanqing, et al.
Publicado: (2024) -
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
por: Ailem, Melissa, et al.
Publicado: (2024)