Enregistré dans:
| Auteurs principaux: | Jin, Hexi, Liu, Stephen, Li, Yuheng, Malik, Simran, Zhang, Yiying |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.16942 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ConvexBench: Can LLMs Recognize Convex Functions?
par: Liu, Yepeng, et autres
Publié: (2026)
par: Liu, Yepeng, et autres
Publié: (2026)
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
par: Li, Yubo, et autres
Publié: (2026)
par: Li, Yubo, et autres
Publié: (2026)
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
par: Xie, Sixiong, et autres
Publié: (2026)
par: Xie, Sixiong, et autres
Publié: (2026)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
par: Zhang, Zhe, et autres
Publié: (2025)
par: Zhang, Zhe, et autres
Publié: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
par: Ouyang, Anne, et autres
Publié: (2025)
par: Ouyang, Anne, et autres
Publié: (2025)
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
par: Cheng, Zihao, et autres
Publié: (2026)
par: Cheng, Zihao, et autres
Publié: (2026)
SC-Bench: A Large-Scale Dataset for Smart Contract Auditing
par: Xia, Shihao, et autres
Publié: (2024)
par: Xia, Shihao, et autres
Publié: (2024)
EXP-Bench: Can AI Conduct AI Research Experiments?
par: Kon, Patrick Tser Jern, et autres
Publié: (2025)
par: Kon, Patrick Tser Jern, et autres
Publié: (2025)
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
par: Li, Bingyu, et autres
Publié: (2026)
par: Li, Bingyu, et autres
Publié: (2026)
Deep Research Bench: Evaluating AI Web Research Agents
par: FutureSearch, et autres
Publié: (2025)
par: FutureSearch, et autres
Publié: (2025)
Multimodal Multihop Source Retrieval for Web Question Answering
par: Yarrabelly, Navya, et autres
Publié: (2025)
par: Yarrabelly, Navya, et autres
Publié: (2025)
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
par: Zanoli, Christopher, et autres
Publié: (2026)
par: Zanoli, Christopher, et autres
Publié: (2026)
LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
par: Zhang, Danqing, et autres
Publié: (2025)
par: Zhang, Danqing, et autres
Publié: (2025)
Can AI Assistance Aid in the Grading of Handwritten Answer Sheets?
par: Sil, Pritam, et autres
Publié: (2024)
par: Sil, Pritam, et autres
Publié: (2024)
InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents
par: Du, Yaxin, et autres
Publié: (2025)
par: Du, Yaxin, et autres
Publié: (2025)
Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact
par: Xu, Haofei, et autres
Publié: (2026)
par: Xu, Haofei, et autres
Publié: (2026)
Dynamic Multi-Agent Orchestration and Retrieval for Multi-Source Question-Answer Systems using Large Language Models
par: Seabra, Antony, et autres
Publié: (2024)
par: Seabra, Antony, et autres
Publié: (2024)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
par: Liu, Chenxu, et autres
Publié: (2026)
par: Liu, Chenxu, et autres
Publié: (2026)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
par: Li, Jidong, et autres
Publié: (2025)
par: Li, Jidong, et autres
Publié: (2025)
Generating High-Quality Datasets for Code Editing via Open-Source Language Models
par: Zhang, Zekai, et autres
Publié: (2025)
par: Zhang, Zekai, et autres
Publié: (2025)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
par: Zhang, Zehua, et autres
Publié: (2025)
par: Zhang, Zehua, et autres
Publié: (2025)
CocoaBench: Evaluating Unified Digital Agents in the Wild
par: CocoaBench Team, et autres
Publié: (2026)
par: CocoaBench Team, et autres
Publié: (2026)
ClawBench: Can AI Agents Complete Everyday Online Tasks?
par: Zhang, Yuxuan, et autres
Publié: (2026)
par: Zhang, Yuxuan, et autres
Publié: (2026)
TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing
par: Wang, Jiaming, et autres
Publié: (2025)
par: Wang, Jiaming, et autres
Publié: (2025)
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
par: Yan, Yuchen, et autres
Publié: (2025)
par: Yan, Yuchen, et autres
Publié: (2025)
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
par: Wang, Qiyao, et autres
Publié: (2026)
par: Wang, Qiyao, et autres
Publié: (2026)
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
par: Zhang, Zhenxing, et autres
Publié: (2025)
par: Zhang, Zhenxing, et autres
Publié: (2025)
LJ-Spoof: A Generatively Varied Corpus for Audio Anti-Spoofing and Synthesis Source Tracing
par: Subramani, Surya, et autres
Publié: (2026)
par: Subramani, Surya, et autres
Publié: (2026)
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
par: Yao, Yang, et autres
Publié: (2025)
par: Yao, Yang, et autres
Publié: (2025)
RedacBench: Can AI Erase Your Secrets?
par: Jeon, Hyunjun, et autres
Publié: (2026)
par: Jeon, Hyunjun, et autres
Publié: (2026)
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
par: Liu, Xuan, et autres
Publié: (2026)
par: Liu, Xuan, et autres
Publié: (2026)
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
par: Yang, Haoyue, et autres
Publié: (2026)
par: Yang, Haoyue, et autres
Publié: (2026)
Cognify: Supercharging Gen-AI Workflows With Hierarchical Autotuning
par: He, Zijian, et autres
Publié: (2025)
par: He, Zijian, et autres
Publié: (2025)
Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
par: Xu, Kai, et autres
Publié: (2025)
par: Xu, Kai, et autres
Publié: (2025)
Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?
par: Nan, Yang, et autres
Publié: (2025)
par: Nan, Yang, et autres
Publié: (2025)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
par: Liu, Junpeng, et autres
Publié: (2024)
par: Liu, Junpeng, et autres
Publié: (2024)
Prompt Sensitivity and Answer Consistency of Small Open-Source Language Models for Clinical Question Answering in Low-Resource Healthcare
par: Hariprasad, Shravani
Publié: (2026)
par: Hariprasad, Shravani
Publié: (2026)
Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
par: Zhang, Jingran, et autres
Publié: (2025)
par: Zhang, Jingran, et autres
Publié: (2025)
WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning
par: Lai, Peichao, et autres
Publié: (2025)
par: Lai, Peichao, et autres
Publié: (2025)
Evaluating AI for Law: Bridging the Gap with Open-Source Solutions
par: Bhambhoria, Rohan, et autres
Publié: (2024)
par: Bhambhoria, Rohan, et autres
Publié: (2024)
Documents similaires
-
ConvexBench: Can LLMs Recognize Convex Functions?
par: Liu, Yepeng, et autres
Publié: (2026) -
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
par: Li, Yubo, et autres
Publié: (2026) -
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
par: Xie, Sixiong, et autres
Publié: (2026) -
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
par: Zhang, Zhe, et autres
Publié: (2025) -
KernelBench: Can LLMs Write Efficient GPU Kernels?
par: Ouyang, Anne, et autres
Publié: (2025)