LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Beiming, Cui, Zhizhuo, Hu, Siteng, Li, Xiaohua, Lin, Haifeng, Zhang, Zhengxin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
by: Lei, Xiang, et al.
Published: (2025)
by: Lei, Xiang, et al.
Published: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026)
by: Haque, Md. Asraful, et al.
Published: (2026)
Towards Probabilistic Question Answering Over Tabular Data
by: Shen, Chen, et al.
Published: (2025)
by: Shen, Chen, et al.
Published: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
by: Lee, Wooin, et al.
Published: (2026)
by: Lee, Wooin, et al.
Published: (2026)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
DPDisc: From Factoid Questions to Data Product Requests for Open-World Data Product Discovery over Tables and Text
by: Zhang, Liangliang, et al.
Published: (2025)
by: Zhang, Liangliang, et al.
Published: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
by: Zong, Chang, et al.
Published: (2024)
by: Zong, Chang, et al.
Published: (2024)
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
by: Kunt, Tim, et al.
Published: (2026)
by: Kunt, Tim, et al.
Published: (2026)
SocialX: A Modular Platform for Multi-Source Big Data Research in Indonesia
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
by: Saputra, Muhammad Apriandito Arya, et al.
Published: (2026)
Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs
by: Pandya, Vedant
Published: (2026)
by: Pandya, Vedant
Published: (2026)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
by: Soman, Sumit, et al.
Published: (2025)
by: Soman, Sumit, et al.
Published: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
by: Tu, Songjun, et al.
Published: (2026)
by: Tu, Songjun, et al.
Published: (2026)
LLM-supported document separation for printed reviews from zbMATH Open
by: Pluzhnikov, Ivan, et al.
Published: (2026)
by: Pluzhnikov, Ivan, et al.
Published: (2026)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
by: Badshah, Sher, et al.
Published: (2025)
by: Badshah, Sher, et al.
Published: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
FARSIQA: Faithful and Advanced RAG System for Islamic Question Answering
by: Asl, Mohammad Aghajani, et al.
Published: (2025)
by: Asl, Mohammad Aghajani, et al.
Published: (2025)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
by: Pan, Leyi, et al.
Published: (2025)
by: Pan, Leyi, et al.
Published: (2025)
GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs
by: Da, Longchao, et al.
Published: (2025)
by: Da, Longchao, et al.
Published: (2025)
Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative Analysis
by: Amamou, Hazem, et al.
Published: (2026)
by: Amamou, Hazem, et al.
Published: (2026)
When Many-Shot Prompting Fails: An Empirical Study of LLM Code Translation
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification
by: Turaev, Sherzod, et al.
Published: (2026)
by: Turaev, Sherzod, et al.
Published: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
by: Nwokocha, Caleb Princewill
Published: (2022)
by: Nwokocha, Caleb Princewill
Published: (2022)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
by: Sun, Jingyi, et al.
Published: (2024)
by: Sun, Jingyi, et al.
Published: (2024)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2023)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
by: Nguyen, Minh Hoang, et al.
Published: (2025)
by: Nguyen, Minh Hoang, et al.
Published: (2025)
Pitfalls in Evaluating Interpretability Agents
by: Haklay, Tal, et al.
Published: (2026)
by: Haklay, Tal, et al.
Published: (2026)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
by: Johnson, Warren
Published: (2026)
by: Johnson, Warren
Published: (2026)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
by: Wang, Yongjie, et al.
Published: (2025)
by: Wang, Yongjie, et al.
Published: (2025)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
by: Zhang, Zhehao, et al.
Published: (2025)
by: Zhang, Zhehao, et al.
Published: (2025)
The Superalignment of Superhuman Intelligence with Large Language Models
by: Huang, Minlie, et al.
Published: (2024)
by: Huang, Minlie, et al.
Published: (2024)
Evaluating Pixel Language Models on Non-Standardized Languages
by: Muñoz-Ortiz, Alberto, et al.
Published: (2024)
by: Muñoz-Ortiz, Alberto, et al.
Published: (2024)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
by: Gupta, Aayush
Published: (2025)
by: Gupta, Aayush
Published: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
by: Kim, Heejun, et al.
Published: (2026)
by: Kim, Heejun, et al.
Published: (2026)
AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs
by: Perera, Manoj Madushanka, et al.
Published: (2026)
by: Perera, Manoj Madushanka, et al.
Published: (2026)
EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference
by: Kumar, Aayush
Published: (2025)
by: Kumar, Aayush
Published: (2025)
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
by: Hasan, Md Arid, et al.
Published: (2025)
by: Hasan, Md Arid, et al.
Published: (2025)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024)
by: Gómez-Rodríguez, Carlos, et al.
Published: (2024)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
by: Bayram, M. Ali, et al.
Published: (2024)
by: Bayram, M. Ali, et al.
Published: (2024)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Similar Items
-
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
by: Lei, Xiang, et al.
Published: (2025) -
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026) -
Towards Probabilistic Question Answering Over Tabular Data
by: Shen, Chen, et al.
Published: (2025) -
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
by: Lee, Wooin, et al.
Published: (2026) -
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
by: Yim, Wen-wai, et al.
Published: (2025)