DeepEval/DeepEval: Artifacts Associated with the Paper Under Review in TOSEM
Fuente:
Zenodo
Salvato in:
| Autori principali: | Xiangyue, Ma, Xiaoting, Du, Chenglong, Li, Xiaoke, Fang, Wenjie, Ding, Zheng, Zheng, Jiangtao, Meng |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Artifact for TOSEM paper: Exploring Development Methods for Reactive Synthesis Specifications
di: Ma'ayan, Dor, et al.
Pubblicazione: (2025)
di: Ma'ayan, Dor, et al.
Pubblicazione: (2025)
Deep Researcher Agent: An Autonomous Framework for 24/7 Deep Learning Experimentation with Zero-Cost Monitoring
di: Zhang, Xiangyue
Pubblicazione: (2026)
di: Zhang, Xiangyue
Pubblicazione: (2026)
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
di: Zhang, Hongzhi, et al.
Pubblicazione: (2026)
di: Zhang, Hongzhi, et al.
Pubblicazione: (2026)
DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
di: Xie, Qianqian, et al.
Pubblicazione: (2026)
di: Xie, Qianqian, et al.
Pubblicazione: (2026)
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
di: Wang, Yibo, et al.
Pubblicazione: (2026)
di: Wang, Yibo, et al.
Pubblicazione: (2026)
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
di: Ye, Fangda, et al.
Pubblicazione: (2026)
di: Ye, Fangda, et al.
Pubblicazione: (2026)
DevEval: Evaluating Code Generation in Practical Software Projects
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
DeepSlide: From Artifacts to Presentation Delivery
di: Yang, Ming, et al.
Pubblicazione: (2026)
di: Yang, Ming, et al.
Pubblicazione: (2026)
ReviewEval: An Evaluation Framework for AI-Generated Reviews
di: Garg, Madhav Krishan, et al.
Pubblicazione: (2025)
di: Garg, Madhav Krishan, et al.
Pubblicazione: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
di: Ho, Xanh, et al.
Pubblicazione: (2026)
di: Ho, Xanh, et al.
Pubblicazione: (2026)
ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts
di: Liu, Shuang, et al.
Pubblicazione: (2025)
di: Liu, Shuang, et al.
Pubblicazione: (2025)
On Randomness in Agentic Evals
di: Bjarnason, Bjarni Haukur, et al.
Pubblicazione: (2026)
di: Bjarnason, Bjarni Haukur, et al.
Pubblicazione: (2026)
AlphaEval: Evaluating Agents in Production
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
di: Du, Junjia, et al.
Pubblicazione: (2025)
di: Du, Junjia, et al.
Pubblicazione: (2025)
MarineEval: Assessing the Marine Intelligence of Vision-Language Models
di: Wong, YuK-Kwan, et al.
Pubblicazione: (2025)
di: Wong, YuK-Kwan, et al.
Pubblicazione: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
di: Shu, Lei, et al.
Pubblicazione: (2023)
di: Shu, Lei, et al.
Pubblicazione: (2023)
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles
di: Huang, Shulin, et al.
Pubblicazione: (2023)
di: Huang, Shulin, et al.
Pubblicazione: (2023)
AmbiGraph-Eval: Can LLMs Effectively Handle Ambiguous Graph Queries?
di: Tian, Yuchen, et al.
Pubblicazione: (2025)
di: Tian, Yuchen, et al.
Pubblicazione: (2025)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
di: Chen, Zehui, et al.
Pubblicazione: (2023)
di: Chen, Zehui, et al.
Pubblicazione: (2023)
EvolMathEval: Towards Evolvable Benchmarks for Mathematical Reasoning via Evolutionary Testing
di: Wang, Shengbo, et al.
Pubblicazione: (2025)
di: Wang, Shengbo, et al.
Pubblicazione: (2025)
QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models
di: Kang, Zhaolu, et al.
Pubblicazione: (2026)
di: Kang, Zhaolu, et al.
Pubblicazione: (2026)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
HumanEval on Latest GPT Models -- 2024
di: Li, Daniel, et al.
Pubblicazione: (2024)
di: Li, Daniel, et al.
Pubblicazione: (2024)
SciSafeEval: A Comprehensive Benchmark for Safety Alignment of Large Language Models in Scientific Tasks
di: Li, Tianhao, et al.
Pubblicazione: (2024)
di: Li, Tianhao, et al.
Pubblicazione: (2024)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
di: Yang, Yuhang, et al.
Pubblicazione: (2025)
di: Yang, Yuhang, et al.
Pubblicazione: (2025)
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
di: He, Zheqi, et al.
Pubblicazione: (2025)
di: He, Zheqi, et al.
Pubblicazione: (2025)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
di: Tang, Guowei, et al.
Pubblicazione: (2026)
di: Tang, Guowei, et al.
Pubblicazione: (2026)
RepEval: Effective Text Evaluation with LLM Representation
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
di: Sheng, Shuqian, et al.
Pubblicazione: (2024)
Sentiment Analysis in SemEval: A Review of Sentiment Identification Approaches
di: Haddaoui, Bousselham El, et al.
Pubblicazione: (2025)
di: Haddaoui, Bousselham El, et al.
Pubblicazione: (2025)
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
di: Jiang, Lai, et al.
Pubblicazione: (2025)
di: Jiang, Lai, et al.
Pubblicazione: (2025)
CloudEval-YAML: A Practical Benchmark for Cloud Configuration Generation
di: Xu, Yifei, et al.
Pubblicazione: (2023)
di: Xu, Yifei, et al.
Pubblicazione: (2023)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning
di: Zhang, Qin, et al.
Pubblicazione: (2026)
di: Zhang, Qin, et al.
Pubblicazione: (2026)
SuiteEval: Simplifying Retrieval Benchmarks
di: Parry, Andrew, et al.
Pubblicazione: (2026)
di: Parry, Andrew, et al.
Pubblicazione: (2026)
PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
di: Xiao, Yujia, et al.
Pubblicazione: (2025)
di: Xiao, Yujia, et al.
Pubblicazione: (2025)
RadEval: A framework for radiology text evaluation
di: Xu, Justin, et al.
Pubblicazione: (2025)
di: Xu, Justin, et al.
Pubblicazione: (2025)
mattwilliamson13/MonEval: Update
di: tylercreech, et al.
Pubblicazione: (2025)
di: tylercreech, et al.
Pubblicazione: (2025)
SycEval: Evaluating LLM Sycophancy
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
di: Fanous, Aaron, et al.
Pubblicazione: (2025)
Measuring all the noises of LLM Evals
di: Wang, Sida
Pubblicazione: (2025)
di: Wang, Sida
Pubblicazione: (2025)
McEval: Massively Multilingual Code Evaluation
di: Chai, Linzheng, et al.
Pubblicazione: (2024)
di: Chai, Linzheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Artifact for TOSEM paper: Exploring Development Methods for Reactive Synthesis Specifications
di: Ma'ayan, Dor, et al.
Pubblicazione: (2025) -
Deep Researcher Agent: An Autonomous Framework for 24/7 Deep Learning Experimentation with Zero-Cost Monitoring
di: Zhang, Xiangyue
Pubblicazione: (2026) -
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
di: Zhang, Hongzhi, et al.
Pubblicazione: (2026) -
DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
di: Xie, Qianqian, et al.
Pubblicazione: (2026) -
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
di: Wang, Yibo, et al.
Pubblicazione: (2026)