BenchScope: How Many Independent Signals Does Your Benchmark Provide?
Fuente:
arXiv
Salvato in:
| Autori principali: | Sha, Tommy, Zhao, Stella |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
di: Reda, Waleed, et al.
Pubblicazione: (2025)
di: Reda, Waleed, et al.
Pubblicazione: (2025)
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
di: Arimbur, Johin Johny
Pubblicazione: (2026)
di: Arimbur, Johin Johny
Pubblicazione: (2026)
ManiBench: A Benchmark for Testing Visual-Logic Drift and Syntactic Hallucinations in Manim Code Generation
di: Oli, Nabin
Pubblicazione: (2026)
di: Oli, Nabin
Pubblicazione: (2026)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
di: Jiang, Yukun, et al.
Pubblicazione: (2026)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
di: Yan, Kai, et al.
Pubblicazione: (2025)
di: Yan, Kai, et al.
Pubblicazione: (2025)
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
di: Li, Xiangyi, et al.
Pubblicazione: (2026)
di: Li, Xiangyi, et al.
Pubblicazione: (2026)
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
di: Abouzeid, Abdelrahman
Pubblicazione: (2026)
di: Abouzeid, Abdelrahman
Pubblicazione: (2026)
Game of Trust: How Trustworthy Does Your Blockchain Think You Are?
di: Drineas, Petros, et al.
Pubblicazione: (2025)
di: Drineas, Petros, et al.
Pubblicazione: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
di: Pozzobon, Luiza, et al.
Pubblicazione: (2024)
di: Pozzobon, Luiza, et al.
Pubblicazione: (2024)
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
di: Sui, Xingyu, et al.
Pubblicazione: (2026)
di: Sui, Xingyu, et al.
Pubblicazione: (2026)
ElecBench: a Power Dispatch Evaluation Benchmark for Large Language Models
di: Zhou, Xiyuan, et al.
Pubblicazione: (2024)
di: Zhou, Xiyuan, et al.
Pubblicazione: (2024)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
di: Wang, Zeyu, et al.
Pubblicazione: (2026)
di: Wang, Zeyu, et al.
Pubblicazione: (2026)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
di: Huang, Zixuan, et al.
Pubblicazione: (2026)
How Many Instructions Can LLMs Follow at Once?
di: Jaroslawicz, Daniel, et al.
Pubblicazione: (2025)
di: Jaroslawicz, Daniel, et al.
Pubblicazione: (2025)
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
di: Shen, Hui, et al.
Pubblicazione: (2025)
di: Shen, Hui, et al.
Pubblicazione: (2025)
Bench-CoE: a Framework for Collaboration of Experts from Benchmark
di: Wang, Yuanshuai, et al.
Pubblicazione: (2024)
di: Wang, Yuanshuai, et al.
Pubblicazione: (2024)
Riemann-Bench: A Benchmark for Moonshot Mathematics
di: Garre, Suhaas, et al.
Pubblicazione: (2026)
di: Garre, Suhaas, et al.
Pubblicazione: (2026)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
di: Li, Yifei, et al.
Pubblicazione: (2025)
di: Li, Yifei, et al.
Pubblicazione: (2025)
CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data
di: Cheng, Zhao, et al.
Pubblicazione: (2024)
di: Cheng, Zhao, et al.
Pubblicazione: (2024)
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
di: Peng, Zixiang, et al.
Pubblicazione: (2026)
di: Peng, Zixiang, et al.
Pubblicazione: (2026)
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy
di: Xiao, Peiyao, et al.
Pubblicazione: (2026)
di: Xiao, Peiyao, et al.
Pubblicazione: (2026)
AMS-IO-Bench and AMS-IO-Agent: Benchmarking and Structured Reasoning for Analog and Mixed-Signal Integrated Circuit Input/Output Design
di: Zhang, Zhishuai, et al.
Pubblicazione: (2025)
di: Zhang, Zhishuai, et al.
Pubblicazione: (2025)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
di: Shi, Wentao, et al.
Pubblicazione: (2026)
di: Shi, Wentao, et al.
Pubblicazione: (2026)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
di: Shen, Xu, et al.
Pubblicazione: (2025)
di: Shen, Xu, et al.
Pubblicazione: (2025)
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
di: Zhou, Xiyuan, et al.
Pubblicazione: (2025)
di: Zhou, Xiyuan, et al.
Pubblicazione: (2025)
TSI-Bench: Benchmarking Time Series Imputation
di: Du, Wenjie, et al.
Pubblicazione: (2024)
di: Du, Wenjie, et al.
Pubblicazione: (2024)
GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization
di: Nimase, Ojas, et al.
Pubblicazione: (2026)
di: Nimase, Ojas, et al.
Pubblicazione: (2026)
How Persuasive is Your Context?
di: Nguyen, Tu, et al.
Pubblicazione: (2025)
di: Nguyen, Tu, et al.
Pubblicazione: (2025)
AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting
di: Xu, Xing, et al.
Pubblicazione: (2026)
di: Xu, Xing, et al.
Pubblicazione: (2026)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
di: Liu, Junqi, et al.
Pubblicazione: (2025)
di: Liu, Junqi, et al.
Pubblicazione: (2025)
DCA-Bench: A Benchmark for Dataset Curation Agents
di: Huang, Benhao, et al.
Pubblicazione: (2024)
di: Huang, Benhao, et al.
Pubblicazione: (2024)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
di: Cheng, Zihao, et al.
Pubblicazione: (2026)
di: Cheng, Zihao, et al.
Pubblicazione: (2026)
PushupBench: Your VLM is not good at counting pushups
di: Li, Shengzhi, et al.
Pubblicazione: (2026)
di: Li, Shengzhi, et al.
Pubblicazione: (2026)
RedacBench: Can AI Erase Your Secrets?
di: Jeon, Hyunjun, et al.
Pubblicazione: (2026)
di: Jeon, Hyunjun, et al.
Pubblicazione: (2026)
lmgame-Bench: How Good are LLMs at Playing Games?
di: Hu, Lanxiang, et al.
Pubblicazione: (2025)
di: Hu, Lanxiang, et al.
Pubblicazione: (2025)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
di: Hu, Yulin, et al.
Pubblicazione: (2026)
di: Hu, Yulin, et al.
Pubblicazione: (2026)
PSPA-Bench: A Personalized Benchmark for Smartphone GUI Agent
di: Nie, Hongyi, et al.
Pubblicazione: (2026)
di: Nie, Hongyi, et al.
Pubblicazione: (2026)
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
di: Li, Ningyuan, et al.
Pubblicazione: (2026)
di: Li, Ningyuan, et al.
Pubblicazione: (2026)
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
di: Ma, Haoxuan, et al.
Pubblicazione: (2026)
di: Ma, Haoxuan, et al.
Pubblicazione: (2026)
ConstraintBench: Benchmarking LLM Constraint Reasoning on Direct Optimization
di: Tso, Joseph, et al.
Pubblicazione: (2026)
di: Tso, Joseph, et al.
Pubblicazione: (2026)
Documenti analoghi
-
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
di: Reda, Waleed, et al.
Pubblicazione: (2025) -
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
di: Arimbur, Johin Johny
Pubblicazione: (2026) -
ManiBench: A Benchmark for Testing Visual-Logic Drift and Syntactic Hallucinations in Manim Code Generation
di: Oli, Nabin
Pubblicazione: (2026) -
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
di: Jiang, Yukun, et al.
Pubblicazione: (2026) -
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
di: Yan, Kai, et al.
Pubblicazione: (2025)