Gespeichert in:
| Hauptverfasser: | Sha, Tommy, Zhao, Stella |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.29357 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
von: Arimbur, Johin Johny
Veröffentlicht: (2026)
von: Arimbur, Johin Johny
Veröffentlicht: (2026)
ManiBench: A Benchmark for Testing Visual-Logic Drift and Syntactic Hallucinations in Manim Code Generation
von: Oli, Nabin
Veröffentlicht: (2026)
von: Oli, Nabin
Veröffentlicht: (2026)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
von: Yan, Kai, et al.
Veröffentlicht: (2025)
von: Yan, Kai, et al.
Veröffentlicht: (2025)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
von: Li, Xiangyi, et al.
Veröffentlicht: (2026)
von: Li, Xiangyi, et al.
Veröffentlicht: (2026)
Game of Trust: How Trustworthy Does Your Blockchain Think You Are?
von: Drineas, Petros, et al.
Veröffentlicht: (2025)
von: Drineas, Petros, et al.
Veröffentlicht: (2025)
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
von: Abouzeid, Abdelrahman
Veröffentlicht: (2026)
von: Abouzeid, Abdelrahman
Veröffentlicht: (2026)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024)
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
von: Sui, Xingyu, et al.
Veröffentlicht: (2026)
von: Sui, Xingyu, et al.
Veröffentlicht: (2026)
ElecBench: a Power Dispatch Evaluation Benchmark for Large Language Models
von: Zhou, Xiyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Xiyuan, et al.
Veröffentlicht: (2024)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
Does Your Reasoning Model Implicitly Know When to Stop Thinking?
von: Huang, Zixuan, et al.
Veröffentlicht: (2026)
von: Huang, Zixuan, et al.
Veröffentlicht: (2026)
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
von: Shen, Hui, et al.
Veröffentlicht: (2025)
von: Shen, Hui, et al.
Veröffentlicht: (2025)
How Many Instructions Can LLMs Follow at Once?
von: Jaroslawicz, Daniel, et al.
Veröffentlicht: (2025)
von: Jaroslawicz, Daniel, et al.
Veröffentlicht: (2025)
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
von: Peng, Zixiang, et al.
Veröffentlicht: (2026)
von: Peng, Zixiang, et al.
Veröffentlicht: (2026)
Bench-CoE: a Framework for Collaboration of Experts from Benchmark
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024)
von: Wang, Yuanshuai, et al.
Veröffentlicht: (2024)
How Persuasive is Your Context?
von: Nguyen, Tu, et al.
Veröffentlicht: (2025)
von: Nguyen, Tu, et al.
Veröffentlicht: (2025)
Riemann-Bench: A Benchmark for Moonshot Mathematics
von: Garre, Suhaas, et al.
Veröffentlicht: (2026)
von: Garre, Suhaas, et al.
Veröffentlicht: (2026)
TSI-Bench: Benchmarking Time Series Imputation
von: Du, Wenjie, et al.
Veröffentlicht: (2024)
von: Du, Wenjie, et al.
Veröffentlicht: (2024)
GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization
von: Nimase, Ojas, et al.
Veröffentlicht: (2026)
von: Nimase, Ojas, et al.
Veröffentlicht: (2026)
CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data
von: Cheng, Zhao, et al.
Veröffentlicht: (2024)
von: Cheng, Zhao, et al.
Veröffentlicht: (2024)
SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy
von: Xiao, Peiyao, et al.
Veröffentlicht: (2026)
von: Xiao, Peiyao, et al.
Veröffentlicht: (2026)
PushupBench: Your VLM is not good at counting pushups
von: Li, Shengzhi, et al.
Veröffentlicht: (2026)
von: Li, Shengzhi, et al.
Veröffentlicht: (2026)
RedacBench: Can AI Erase Your Secrets?
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
von: Jeon, Hyunjun, et al.
Veröffentlicht: (2026)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
AMS-IO-Bench and AMS-IO-Agent: Benchmarking and Structured Reasoning for Analog and Mixed-Signal Integrated Circuit Input/Output Design
von: Zhang, Zhishuai, et al.
Veröffentlicht: (2025)
von: Zhang, Zhishuai, et al.
Veröffentlicht: (2025)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
von: Shen, Xu, et al.
Veröffentlicht: (2025)
von: Shen, Xu, et al.
Veröffentlicht: (2025)
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
von: Zhou, Xiyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Xiyuan, et al.
Veröffentlicht: (2025)
AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting
von: Xu, Xing, et al.
Veröffentlicht: (2026)
von: Xu, Xing, et al.
Veröffentlicht: (2026)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
von: Hu, Yulin, et al.
Veröffentlicht: (2026)
von: Hu, Yulin, et al.
Veröffentlicht: (2026)
YourBench: Easy Custom Evaluation Sets for Everyone
von: Shashidhar, Sumuk, et al.
Veröffentlicht: (2025)
von: Shashidhar, Sumuk, et al.
Veröffentlicht: (2025)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
von: Liu, Junqi, et al.
Veröffentlicht: (2025)
von: Liu, Junqi, et al.
Veröffentlicht: (2025)
DCA-Bench: A Benchmark for Dataset Curation Agents
von: Huang, Benhao, et al.
Veröffentlicht: (2024)
von: Huang, Benhao, et al.
Veröffentlicht: (2024)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
von: Cheng, Zihao, et al.
Veröffentlicht: (2026)
AI incidents and 'networked trouble': The case for a research agenda
von: Shane, Tommy Shaffer
Veröffentlicht: (2024)
von: Shane, Tommy Shaffer
Veröffentlicht: (2024)
C-SEO Bench: Does Conversational SEO Work?
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
lmgame-Bench: How Good are LLMs at Playing Games?
von: Hu, Lanxiang, et al.
Veröffentlicht: (2025)
von: Hu, Lanxiang, et al.
Veröffentlicht: (2025)
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
von: Zhao, Kangjia, et al.
Veröffentlicht: (2024)
von: Zhao, Kangjia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
von: Reda, Waleed, et al.
Veröffentlicht: (2025) -
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
von: Arimbur, Johin Johny
Veröffentlicht: (2026) -
ManiBench: A Benchmark for Testing Visual-Logic Drift and Syntactic Hallucinations in Manim Code Generation
von: Oli, Nabin
Veröffentlicht: (2026) -
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
von: Yan, Kai, et al.
Veröffentlicht: (2025) -
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)