INTEGRALBENCH: Benchmarking LLMs with Definite Integral Problems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tang, Bintao, Yang, Xin, Wang, Yuhao, Qiu, Zixuan, Ji, Zimo, Jiang, Wenyuan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards Provable (In)Secure Model Weight Release Schemes
par: Yang, Xin, et autres
Publié: (2025)
par: Yang, Xin, et autres
Publié: (2025)
Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges
par: Ji, Zimo, et autres
Publié: (2025)
par: Ji, Zimo, et autres
Publié: (2025)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
par: Ji, Zimo, et autres
Publié: (2026)
par: Ji, Zimo, et autres
Publié: (2026)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
par: Yang, Chang, et autres
Publié: (2025)
par: Yang, Chang, et autres
Publié: (2025)
Event Stream-based Visual Object Tracking: HDETrack V2 and A High-Definition Benchmark
par: Wang, Shiao, et autres
Publié: (2025)
par: Wang, Shiao, et autres
Publié: (2025)
VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
par: Gu, Zixuan, et autres
Publié: (2025)
par: Gu, Zixuan, et autres
Publié: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
par: Jiang, Botian, et autres
Publié: (2024)
par: Jiang, Botian, et autres
Publié: (2024)
Feasible Pairings for Decentralized Integral Controllability of Non-Square Systems
par: Tong, Yuhao, et autres
Publié: (2026)
par: Tong, Yuhao, et autres
Publié: (2026)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
par: Yang, Xin, et autres
Publié: (2026)
par: Yang, Xin, et autres
Publié: (2026)
A Definition of Open-Ended Learning Problems for Goal-Conditioned Agents
par: Sigaud, Olivier, et autres
Publié: (2023)
par: Sigaud, Olivier, et autres
Publié: (2023)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
par: Mohammadi, Seyedali, et autres
Publié: (2025)
par: Mohammadi, Seyedali, et autres
Publié: (2025)
NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs
par: Qi, Yunjin, et autres
Publié: (2026)
par: Qi, Yunjin, et autres
Publié: (2026)
T-DuMpRa: Teacher-guided Dual-path Multi-prototype Retrieval Augmented framework for fine-grained medical image classification
par: Tang, Zixuan, et autres
Publié: (2026)
par: Tang, Zixuan, et autres
Publié: (2026)
Why Does New Knowledge Create Messy Ripple Effects in LLMs?
par: Qin, Jiaxin, et autres
Publié: (2024)
par: Qin, Jiaxin, et autres
Publié: (2024)
Improving LLMs' Generalized Reasoning Abilities by Graph Problems
par: Zhang, Qifan, et autres
Publié: (2025)
par: Zhang, Qifan, et autres
Publié: (2025)
Flames: Benchmarking Value Alignment of LLMs in Chinese
par: Huang, Kexin, et autres
Publié: (2023)
par: Huang, Kexin, et autres
Publié: (2023)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
par: Hu, Ruida, et autres
Publié: (2025)
par: Hu, Ruida, et autres
Publié: (2025)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
par: Ning, Xuefei, et autres
Publié: (2023)
par: Ning, Xuefei, et autres
Publié: (2023)
Can LLMs Solve ASP Problems? Insights from a Benchmarking Study (Extended Version)
par: Ren, Lin, et autres
Publié: (2025)
par: Ren, Lin, et autres
Publié: (2025)
Exploring Adversarial Robustness of LiDAR-Camera Fusion Model in Autonomous Driving
par: Yang, Bo, et autres
Publié: (2023)
par: Yang, Bo, et autres
Publié: (2023)
DHP Benchmark: Are LLMs Good NLG Evaluators?
par: Wang, Yicheng, et autres
Publié: (2024)
par: Wang, Yicheng, et autres
Publié: (2024)
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
par: Yang, Kai, et autres
Publié: (2025)
par: Yang, Kai, et autres
Publié: (2025)
Trapezoidal Gradient Descent for Effective Reinforcement Learning in Spiking Networks
par: Pan, Yuhao, et autres
Publié: (2024)
par: Pan, Yuhao, et autres
Publié: (2024)
Advancing and Benchmarking Personalized Tool Invocation for LLMs
par: Huang, Xu, et autres
Publié: (2025)
par: Huang, Xu, et autres
Publié: (2025)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
par: Xu, Wanghan, et autres
Publié: (2025)
par: Xu, Wanghan, et autres
Publié: (2025)
Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks
par: Tian, Aaron Xuxiang, et autres
Publié: (2025)
par: Tian, Aaron Xuxiang, et autres
Publié: (2025)
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
par: Zhang, Terry Jingchen, et autres
Publié: (2025)
par: Zhang, Terry Jingchen, et autres
Publié: (2025)
FourierCSP: Differentiable Constraint Satisfaction Problem Solving by Walsh-Fourier Expansion
par: Cen, Yunuo, et autres
Publié: (2025)
par: Cen, Yunuo, et autres
Publié: (2025)
On the Definition of Intelligence
par: Ng, Kei-Sing
Publié: (2025)
par: Ng, Kei-Sing
Publié: (2025)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
par: Zhang, Yiming, et autres
Publié: (2025)
par: Zhang, Yiming, et autres
Publié: (2025)
A Memory-Augmented LLM-Driven Method for Autonomous Merging of 3D Printing Work Orders
par: Liu, Yuhao, et autres
Publié: (2025)
par: Liu, Yuhao, et autres
Publié: (2025)
AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization
par: Jiang, Zixuan, et autres
Publié: (2025)
par: Jiang, Zixuan, et autres
Publié: (2025)
CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
par: Li, Zhihao, et autres
Publié: (2025)
par: Li, Zhihao, et autres
Publié: (2025)
When Does Hierarchy Help? Benchmarking Agent Coordination in Event-Driven Industrial Scheduling
par: Wang, Ziqi, et autres
Publié: (2026)
par: Wang, Ziqi, et autres
Publié: (2026)
ProBench: Benchmarking GUI Agents with Accurate Process Information
par: Yang, Leyang, et autres
Publié: (2025)
par: Yang, Leyang, et autres
Publié: (2025)
SmartPlay: A Benchmark for LLMs as Intelligent Agents
par: Wu, Yue, et autres
Publié: (2023)
par: Wu, Yue, et autres
Publié: (2023)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
par: Ji, Yuhao, et autres
Publié: (2024)
par: Ji, Yuhao, et autres
Publié: (2024)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
par: Xu, Ruiling, et autres
Publié: (2025)
par: Xu, Ruiling, et autres
Publié: (2025)
Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs
par: Zhu, Longyuan, et autres
Publié: (2026)
par: Zhu, Longyuan, et autres
Publié: (2026)
Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms
par: Wang, Xiao, et autres
Publié: (2024)
par: Wang, Xiao, et autres
Publié: (2024)
Documents similaires
-
Towards Provable (In)Secure Model Weight Release Schemes
par: Yang, Xin, et autres
Publié: (2025) -
Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges
par: Ji, Zimo, et autres
Publié: (2025) -
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
par: Ji, Zimo, et autres
Publié: (2026) -
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
par: Yang, Chang, et autres
Publié: (2025) -
Event Stream-based Visual Object Tracking: HDETrack V2 and A High-Definition Benchmark
par: Wang, Shiao, et autres
Publié: (2025)