Gespeichert in:
| Hauptverfasser: | Sokol, Anna, Daly, Elizabeth, Hind, Michael, Piorkowski, David, Zhang, Xiangliang, Moniz, Nuno, Chawla, Nitesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.12974 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
von: Hofmann, Aris, et al.
Veröffentlicht: (2025)
von: Hofmann, Aris, et al.
Veröffentlicht: (2025)
Conformalized Selective Regression
von: Sokol, Anna, et al.
Veröffentlicht: (2024)
von: Sokol, Anna, et al.
Veröffentlicht: (2024)
Breaking Language Barriers: Equitable Performance in Multilingual Language Models
von: Nagar, Tanay, et al.
Veröffentlicht: (2025)
von: Nagar, Tanay, et al.
Veröffentlicht: (2025)
Class-Aware Contrastive Optimization for Imbalanced Text Classification
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2024)
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
von: Guo, Taicheng, et al.
Veröffentlicht: (2026)
von: Guo, Taicheng, et al.
Veröffentlicht: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Capability-Oriented Training Induced Alignment Risk
von: Zhou, Yujun, et al.
Veröffentlicht: (2026)
von: Zhou, Yujun, et al.
Veröffentlicht: (2026)
Large Language Model based Multi-Agents: A Survey of Progress and Challenges
von: Guo, Taicheng, et al.
Veröffentlicht: (2024)
von: Guo, Taicheng, et al.
Veröffentlicht: (2024)
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2024)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
von: Ye, Jiayi, et al.
Veröffentlicht: (2024)
Fast Explanations via Policy Gradient-Optimized Explainer
von: Pan, Deng, et al.
Veröffentlicht: (2024)
von: Pan, Deng, et al.
Veröffentlicht: (2024)
SMOTExT: SMOTE meets Large Language Models
von: Bystroński, Mateusz, et al.
Veröffentlicht: (2025)
von: Bystroński, Mateusz, et al.
Veröffentlicht: (2025)
PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
von: Bao, Han, et al.
Veröffentlicht: (2026)
von: Bao, Han, et al.
Veröffentlicht: (2026)
Cards Against LLMs: Benchmarking Humor Alignment in Large Language Models
von: Fettach, Yousra, et al.
Veröffentlicht: (2026)
von: Fettach, Yousra, et al.
Veröffentlicht: (2026)
AnyLoss: Transforming Classification Metrics into Loss Functions
von: Han, Doheon, et al.
Veröffentlicht: (2024)
von: Han, Doheon, et al.
Veröffentlicht: (2024)
Intersectional Divergence: Measuring Fairness in Regression
von: Germino, Joe, et al.
Veröffentlicht: (2025)
von: Germino, Joe, et al.
Veröffentlicht: (2025)
Emergent Social Intelligence Risks in Generative Multi-Agent Systems
von: Huang, Yue, et al.
Veröffentlicht: (2026)
von: Huang, Yue, et al.
Veröffentlicht: (2026)
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
von: Shi, Kaiwen, et al.
Veröffentlicht: (2026)
von: Shi, Kaiwen, et al.
Veröffentlicht: (2026)
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
von: Chen, Si, et al.
Veröffentlicht: (2026)
von: Chen, Si, et al.
Veröffentlicht: (2026)
Do Multimodal Large Language Models Understand Welding?
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
von: Li, Peiyu, et al.
Veröffentlicht: (2025)
von: Li, Peiyu, et al.
Veröffentlicht: (2025)
Evaluating a Methodology for Increasing AI Transparency: A Case Study
von: Piorkowski, David, et al.
Veröffentlicht: (2022)
von: Piorkowski, David, et al.
Veröffentlicht: (2022)
Differentially-Private Data Synthetisation for Efficient Re-Identification Risk Control
von: Carvalho, Tânia, et al.
Veröffentlicht: (2022)
von: Carvalho, Tânia, et al.
Veröffentlicht: (2022)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
Are we making much progress? Revisiting chemical reaction yield prediction from an imbalanced regression perspective
von: Ma, Yihong, et al.
Veröffentlicht: (2024)
von: Ma, Yihong, et al.
Veröffentlicht: (2024)
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models
von: Tirupathi, Seshu, et al.
Veröffentlicht: (2025)
von: Tirupathi, Seshu, et al.
Veröffentlicht: (2025)
Redefining Simplicity: Benchmarking Large Language Models from Lexical to Document Simplification
von: Qiang, Jipeng, et al.
Veröffentlicht: (2025)
von: Qiang, Jipeng, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
von: Wu, Shiwei, et al.
Veröffentlicht: (2024)
Graph Neural Prompting with Large Language Models
von: Tian, Yijun, et al.
Veröffentlicht: (2023)
von: Tian, Yijun, et al.
Veröffentlicht: (2023)
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
von: Lin, Teng, et al.
Veröffentlicht: (2025)
von: Lin, Teng, et al.
Veröffentlicht: (2025)
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models
von: Jang, Haeun, et al.
Veröffentlicht: (2026)
von: Jang, Haeun, et al.
Veröffentlicht: (2026)
Benchmarking Benchmark Leakage in Large Language Models
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
von: Xu, Ruijie, et al.
Veröffentlicht: (2024)
Reliable Control-Point Selection for Steering Reasoning in Large Language Models
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
von: Zhuang, Haomin, et al.
Veröffentlicht: (2026)
Can Large Language Models Master Complex Card Games?
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
A Survey on Large Language Model Benchmarks
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
von: Ni, Shiwen, et al.
Veröffentlicht: (2025)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
von: Gao, Lang, et al.
Veröffentlicht: (2024)
von: Gao, Lang, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models with Psychometrics
von: Li, Yuan, et al.
Veröffentlicht: (2024)
von: Li, Yuan, et al.
Veröffentlicht: (2024)
LLMs4All: A Review of Large Language Models Across Academic Disciplines
von: Ye, Yanfang, et al.
Veröffentlicht: (2025)
von: Ye, Yanfang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
von: Hofmann, Aris, et al.
Veröffentlicht: (2025) -
Conformalized Selective Regression
von: Sokol, Anna, et al.
Veröffentlicht: (2024) -
Breaking Language Barriers: Equitable Performance in Multilingual Language Models
von: Nagar, Tanay, et al.
Veröffentlicht: (2025) -
Class-Aware Contrastive Optimization for Imbalanced Text Classification
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2024) -
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)