MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Xuan, Weihao, Yang, Rui, Qi, Heli, Zeng, Qingcheng, Xiao, Yunze, Feng, Aosong, Liu, Dairui, Xing, Yun, Wang, Junjue, Gao, Fan, Lu, Jinghui, Jiang, Yuang, Li, Huitao, Li, Xin, Yu, Kunyu, Dong, Ruihai, Gu, Shangding, Li, Yuekang, Xie, Xiaofei, Juefei-Xu, Felix, Khomh, Foutse, Yoshie, Osamu, Chen, Qingyu, Teodoro, Douglas, Liu, Nan, Goebel, Randy, Ma, Lei, Marrese-Taylor, Edison, Lu, Shijian, Iwasawa, Yusuke, Matsuo, Yutaka, Li, Irene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
by: Xuan, Weihao, et al.
Published: (2026)
by: Xuan, Weihao, et al.
Published: (2026)
Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts
by: Gao, Fan, et al.
Published: (2023)
by: Gao, Fan, et al.
Published: (2023)
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
Topic-Centric Explanations for News Recommendation
by: Liu, Dairui, et al.
Published: (2023)
by: Liu, Dairui, et al.
Published: (2023)
On the Effectiveness of Log Representation for Log-based Anomaly Detection
by: Wu, Xingfang, et al.
Published: (2023)
by: Wu, Xingfang, et al.
Published: (2023)
Protecting Privacy in Software Logs: What Should Be Anonymized?
by: Aghili, Roozbeh, et al.
Published: (2024)
by: Aghili, Roozbeh, et al.
Published: (2024)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
by: Wu, Xingfang, et al.
Published: (2024)
by: Wu, Xingfang, et al.
Published: (2024)
Representation Improvement in Latent Space for Search-Based Testing of Autonomous Robotic Systems
by: Humeniuk, Dmytro, et al.
Published: (2025)
by: Humeniuk, Dmytro, et al.
Published: (2025)
An Efficient Model Maintenance Approach for MLOps
by: Majidi, Forough, et al.
Published: (2024)
by: Majidi, Forough, et al.
Published: (2024)
Understanding Web Application Workloads and Their Applications: Systematic Literature Review and Characterization
by: Aghili, Roozbeh, et al.
Published: (2024)
by: Aghili, Roozbeh, et al.
Published: (2024)
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Tracing Optimization for Performance Modeling and Regression Detection
by: Shahedi, Kaveh, et al.
Published: (2024)
by: Shahedi, Kaveh, et al.
Published: (2024)
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
by: Aghili, Roozbeh, et al.
Published: (2025)
by: Aghili, Roozbeh, et al.
Published: (2025)
Machine Learning Robustness: A Primer
by: Braiek, Houssem Ben, et al.
Published: (2024)
by: Braiek, Houssem Ben, et al.
Published: (2024)
Leveraging Large Language Models for Concept Graph Recovery and Question Answering in NLP Education
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
Graphusion: Leveraging Large Language Models for Scientific Knowledge Graph Fusion and Construction in NLP Education
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs
by: Chen, Yingjian, et al.
Published: (2026)
by: Chen, Yingjian, et al.
Published: (2026)
Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis
by: Song, Da, et al.
Published: (2026)
by: Song, Da, et al.
Published: (2026)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
GIST: Generated Inputs Sets Transferability in Deep Learning
by: Tambon, Florian, et al.
Published: (2023)
by: Tambon, Florian, et al.
Published: (2023)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Reinforcement Learning Informed Evolutionary Search for Autonomous Systems Testing
by: Humeniuk, Dmytro, et al.
Published: (2023)
by: Humeniuk, Dmytro, et al.
Published: (2023)
Evaluating and Enhancing Segmentation Model Robustness with Metamorphic Testing
by: Mzoughi, Seif, et al.
Published: (2025)
by: Mzoughi, Seif, et al.
Published: (2025)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025)
by: Majdinasab, Vahid, et al.
Published: (2025)
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4
by: Abukhalaf, Seif, et al.
Published: (2024)
by: Abukhalaf, Seif, et al.
Published: (2024)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
by: Oueslati, Khouloud, et al.
Published: (2025)
by: Oueslati, Khouloud, et al.
Published: (2025)
Graphusion: A RAG Framework for Knowledge Graph Construction with a Global Perspective
by: Yang, Rui, et al.
Published: (2024)
by: Yang, Rui, et al.
Published: (2024)
An Empirical Study on Method-Level Performance Evolution in Open-Source Java Projects
by: Shahedi, Kaveh, et al.
Published: (2025)
by: Shahedi, Kaveh, et al.
Published: (2025)
From Technical Excellence to Practical Adoption: Lessons Learned Building an ML-Enhanced Trace Analysis Tool
by: Shahedi, Kaveh, et al.
Published: (2025)
by: Shahedi, Kaveh, et al.
Published: (2025)
Improving the Robustness of Large Language Models for Code Tasks via Fine-tuning with Perturbed Data
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Structural Anchors and Reasoning Fragility:Understanding CoT Robustness in LLM4Code
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
QMon: Monitoring the Execution of Quantum Circuits with Mid-Circuit Measurement and Reset
by: Ma, Ning, et al.
Published: (2025)
by: Ma, Ning, et al.
Published: (2025)
Refining GPT-3 Embeddings with a Siamese Structure for Technical Post Duplicate Detection
by: Wu, Xingfang, et al.
Published: (2023)
by: Wu, Xingfang, et al.
Published: (2023)
RecPrompt: A Self-tuning Prompting Framework for News Recommendation Using Large Language Models
by: Liu, Dairui, et al.
Published: (2023)
by: Liu, Dairui, et al.
Published: (2023)
Transformers4NewsRec: A Transformer-based News Recommendation Framework
by: Liu, Dairui, et al.
Published: (2024)
by: Liu, Dairui, et al.
Published: (2024)
Direction-aware 3D Large Multimodal Models
by: Liu, Quan, et al.
Published: (2026)
by: Liu, Quan, et al.
Published: (2026)
An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges
by: Foalem, Patrick Loic, et al.
Published: (2026)
by: Foalem, Patrick Loic, et al.
Published: (2026)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
by: Da Silva, Leuson, et al.
Published: (2024)
by: Da Silva, Leuson, et al.
Published: (2024)
One Size Does Not Fit All: Architecture-Aware Adaptive Batch Scheduling with DEBA
by: Belias, François, et al.
Published: (2025)
by: Belias, François, et al.
Published: (2025)
Fault Localization in Deep Learning-based Software: A System-level Approach
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
by: Morovati, Mohammad Mehdi, et al.
Published: (2024)
Similar Items
-
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
by: Xuan, Weihao, et al.
Published: (2026) -
Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts
by: Gao, Fan, et al.
Published: (2023) -
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
by: Xuan, Weihao, et al.
Published: (2025) -
Topic-Centric Explanations for News Recommendation
by: Liu, Dairui, et al.
Published: (2023) -
On the Effectiveness of Log Representation for Log-based Anomaly Detection
by: Wu, Xingfang, et al.
Published: (2023)