HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Haokun, Huang, Sicong, Hu, Jingyu, Zhou, Yangqiaoyu, Tan, Chenhao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hypothesis Generation with Large Language Models
by: Zhou, Yangqiaoyu, et al.
Published: (2024)
by: Zhou, Yangqiaoyu, et al.
Published: (2024)
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2024)
by: Liu, Haokun, et al.
Published: (2024)
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
by: Li, Mingxuan, et al.
Published: (2025)
by: Li, Mingxuan, et al.
Published: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
by: Bai, Xiaoyan, et al.
Published: (2025)
by: Bai, Xiaoyan, et al.
Published: (2025)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions
by: Heddaya, Mourad, et al.
Published: (2024)
by: Heddaya, Mourad, et al.
Published: (2024)
Sparse Autoencoders for Hypothesis Generation
by: Movva, Rajiv, et al.
Published: (2025)
by: Movva, Rajiv, et al.
Published: (2025)
GPT-4V Cannot Generate Radiology Reports Yet
by: Jiang, Yuyang, et al.
Published: (2024)
by: Jiang, Yuyang, et al.
Published: (2024)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models
by: Song, Kefan, et al.
Published: (2025)
by: Song, Kefan, et al.
Published: (2025)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
by: Ghaboura, Sara, et al.
Published: (2024)
by: Ghaboura, Sara, et al.
Published: (2024)
DarkBench: Benchmarking Dark Patterns in Large Language Models
by: Kran, Esben, et al.
Published: (2025)
by: Kran, Esben, et al.
Published: (2025)
MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023)
by: Kıcıman, Emre, et al.
Published: (2023)
Towards Enriched Controllability for Educational Question Generation
by: Leite, Bernardo, et al.
Published: (2023)
by: Leite, Bernardo, et al.
Published: (2023)
The Lock-in Hypothesis: Stagnation by Algorithm
by: Qiu, Tianyi Alex, et al.
Published: (2025)
by: Qiu, Tianyi Alex, et al.
Published: (2025)
Superficial Safety Alignment Hypothesis
by: Li, Jianwei, et al.
Published: (2024)
by: Li, Jianwei, et al.
Published: (2024)
EigenBench: A Comparative Behavioral Measure of Value Alignment
by: Chang, Jonathn, et al.
Published: (2025)
by: Chang, Jonathn, et al.
Published: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Language of Bargaining
by: Heddaya, Mourad, et al.
Published: (2023)
by: Heddaya, Mourad, et al.
Published: (2023)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
by: Sosto, Mae, et al.
Published: (2024)
by: Sosto, Mae, et al.
Published: (2024)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
by: Kim, Kyuhee, et al.
Published: (2025)
by: Kim, Kyuhee, et al.
Published: (2025)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Embedding Enhancement via Fine-Tuned Language Models for Learner-Item Cognitive Modeling
by: Liu, Yuanhao, et al.
Published: (2026)
by: Liu, Yuanhao, et al.
Published: (2026)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
Towards Autonomous Mathematics Research
by: Feng, Tony, et al.
Published: (2026)
by: Feng, Tony, et al.
Published: (2026)
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
by: Li, Nathaniel, et al.
Published: (2024)
by: Li, Nathaniel, et al.
Published: (2024)
HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets
by: Yuan, Ying, et al.
Published: (2025)
by: Yuan, Ying, et al.
Published: (2025)
PakBBQ: A Culturally Adapted Bias Benchmark for QA
by: Hashmat, Abdullah, et al.
Published: (2025)
by: Hashmat, Abdullah, et al.
Published: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Similar Items
-
Hypothesis Generation with Large Language Models
by: Zhou, Yangqiaoyu, et al.
Published: (2024) -
Literature Meets Data: A Synergistic Approach to Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2024) -
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation
by: Li, Mingxuan, et al.
Published: (2025) -
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
by: Nguyen, Dang, et al.
Published: (2025) -
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)