SysTradeBench: An Iterative Build-Test-Patch Benchmark for Strategy-to-Code Trading Systems with Drift-Aware Diagnostics
Fuente:
arXiv
Salvato in:
| Autori principali: | Cao, Yuchen, Zhang, Hanlin, Keung, Jacky Wai, Chen, Yang, Song, Linqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing
di: Wang, Yifei, et al.
Pubblicazione: (2025)
di: Wang, Yifei, et al.
Pubblicazione: (2025)
Fight Fire with Fire: How Much Can We Trust ChatGPT on Source Code-Related Tasks?
di: Yu, Xiao, et al.
Pubblicazione: (2024)
di: Yu, Xiao, et al.
Pubblicazione: (2024)
Data Preparation for Deep Learning based Code Smell Detection: A Systematic Literature Review
di: Zhang, Fengji, et al.
Pubblicazione: (2024)
di: Zhang, Fengji, et al.
Pubblicazione: (2024)
R2Code: A Self-Reflective LLM Framework for Requirements-to-Code Traceability
di: Wang, Yifei, et al.
Pubblicazione: (2026)
di: Wang, Yifei, et al.
Pubblicazione: (2026)
Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
di: Yang, Zhen, et al.
Pubblicazione: (2025)
di: Yang, Zhen, et al.
Pubblicazione: (2025)
Towards Understanding Bugs in Distributed Training and Inference Frameworks for Large Language Models
di: Yu, Xiao, et al.
Pubblicazione: (2025)
di: Yu, Xiao, et al.
Pubblicazione: (2025)
Delving into Parameter-Efficient Fine-Tuning in Code Change Learning: An Empirical Study
di: Liu, Shuo, et al.
Pubblicazione: (2024)
di: Liu, Shuo, et al.
Pubblicazione: (2024)
UniAda: Universal Adaptive Multi-objective Adversarial Attack for End-to-End Autonomous Driving Systems
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
di: Aleithan, Reem, et al.
Pubblicazione: (2024)
di: Aleithan, Reem, et al.
Pubblicazione: (2024)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
di: Muna, Rabeya Khatun, et al.
Pubblicazione: (2026)
di: Muna, Rabeya Khatun, et al.
Pubblicazione: (2026)
Hybrid Privacy Policy-Code Consistency Check using Knowledge Graphs and LLMs
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
di: Wang, Sizhe, et al.
Pubblicazione: (2025)
di: Wang, Sizhe, et al.
Pubblicazione: (2025)
A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
di: Ma, Xiaoxue, et al.
Pubblicazione: (2025)
di: Ma, Xiaoxue, et al.
Pubblicazione: (2025)
InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution
di: Li, KeFan, et al.
Pubblicazione: (2025)
di: Li, KeFan, et al.
Pubblicazione: (2025)
Unlocking LLM Repair Capabilities Through Cross-Language Translation and Multi-Agent Refinement
di: Luo, Wenqiang, et al.
Pubblicazione: (2025)
di: Luo, Wenqiang, et al.
Pubblicazione: (2025)
ClassEval-T: Evaluating Large Language Models in Class-Level Code Translation
di: Xue, Pengyu, et al.
Pubblicazione: (2024)
di: Xue, Pengyu, et al.
Pubblicazione: (2024)
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
di: Li, Xuefeng, et al.
Pubblicazione: (2026)
di: Li, Xuefeng, et al.
Pubblicazione: (2026)
iCoRe: An Iterative Correlation-Aware Retriever for Bug Reproduction Test Generation
di: Wang, Junyi, et al.
Pubblicazione: (2026)
di: Wang, Junyi, et al.
Pubblicazione: (2026)
Exploring and Unleashing the Power of Large Language Models in Automated Code Translation
di: Yang, Zhen, et al.
Pubblicazione: (2024)
di: Yang, Zhen, et al.
Pubblicazione: (2024)
An Iterative Test-and-Repair Framework for Competitive Code Generation
di: Tang, Lingxiao, et al.
Pubblicazione: (2026)
di: Tang, Lingxiao, et al.
Pubblicazione: (2026)
When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair
di: Luo, Wenqiang, et al.
Pubblicazione: (2024)
di: Luo, Wenqiang, et al.
Pubblicazione: (2024)
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
di: Xu, Qiao, et al.
Pubblicazione: (2026)
di: Xu, Qiao, et al.
Pubblicazione: (2026)
Maximizing Patch Coverage for Testing of Highly-Configurable Software without Exploding Build Times
di: Yıldıran, Necip Fazıl, et al.
Pubblicazione: (2024)
di: Yıldıran, Necip Fazıl, et al.
Pubblicazione: (2024)
Towards Engineering Multi-Agent LLMs: A Protocol-Driven Approach
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face
di: Liu, Yujian, et al.
Pubblicazione: (2026)
di: Liu, Yujian, et al.
Pubblicazione: (2026)
Towards Requirements Engineering for GenAI-Enabled Software: Bridging Responsibility Gaps through Human Oversight Requirements
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
RubberDuckBench: A Benchmark for AI Coding Assistants
di: Mohammed, Ferida, et al.
Pubblicazione: (2026)
di: Mohammed, Ferida, et al.
Pubblicazione: (2026)
Lares: LLM-driven Code Slice Semantic Search for Patch Presence Testing
di: Li, Siyuan, et al.
Pubblicazione: (2025)
di: Li, Siyuan, et al.
Pubblicazione: (2025)
SysPro: Reproducing System-level Concurrency Bugs from Bug Reports
di: Zaman, Tarannum Shaila, et al.
Pubblicazione: (2026)
di: Zaman, Tarannum Shaila, et al.
Pubblicazione: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
di: Orlanski, Gabriel, et al.
Pubblicazione: (2026)
di: Orlanski, Gabriel, et al.
Pubblicazione: (2026)
SysLLMatic: Large Language Models are Software System Optimizers
di: Peng, Huiyun, et al.
Pubblicazione: (2025)
di: Peng, Huiyun, et al.
Pubblicazione: (2025)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
di: Guo, Xiaoyu, et al.
Pubblicazione: (2025)
di: Guo, Xiaoyu, et al.
Pubblicazione: (2025)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
di: Ouyang, Shuyin, et al.
Pubblicazione: (2025)
di: Ouyang, Shuyin, et al.
Pubblicazione: (2025)
Uncertainty Modeling for SysML v2
di: Zhang, Man, et al.
Pubblicazione: (2026)
di: Zhang, Man, et al.
Pubblicazione: (2026)
Lightweight Model Editing for LLMs to Correct Deprecated API Recommendations
di: Lin, Guancheng, et al.
Pubblicazione: (2025)
di: Lin, Guancheng, et al.
Pubblicazione: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
di: Jing, Huihao, et al.
Pubblicazione: (2026)
di: Jing, Huihao, et al.
Pubblicazione: (2026)
FedLAD: A Modular and Adaptive Testbed for Federated Log Anomaly Detection
di: Liao, Yihan, et al.
Pubblicazione: (2025)
di: Liao, Yihan, et al.
Pubblicazione: (2025)
Exposing and Defending Membership Leakage in Vulnerability Prediction Models
di: Liao, Yihan, et al.
Pubblicazione: (2025)
di: Liao, Yihan, et al.
Pubblicazione: (2025)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
di: Liu, Linbo, et al.
Pubblicazione: (2025)
di: Liu, Linbo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Chart2Code-MoLA: Efficient Multi-Modal Code Generation via Adaptive Expert Routing
di: Wang, Yifei, et al.
Pubblicazione: (2025) -
Fight Fire with Fire: How Much Can We Trust ChatGPT on Source Code-Related Tasks?
di: Yu, Xiao, et al.
Pubblicazione: (2024) -
Data Preparation for Deep Learning based Code Smell Detection: A Systematic Literature Review
di: Zhang, Fengji, et al.
Pubblicazione: (2024) -
R2Code: A Self-Reflective LLM Framework for Requirements-to-Code Traceability
di: Wang, Yifei, et al.
Pubblicazione: (2026) -
Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)