A Comprehensive Evaluation of Four End-to-End AI Autopilots Using CCTest and the Carla Leaderboard
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Changwen, Sifakis, Joseph, Yan, Rongjie, Zhang, Jian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rigorous Simulation-based Testing for Autonomous Driving Systems -- Targeting the Achilles' Heel of Four Open Autopilots
di: Li, Changwen, et al.
Pubblicazione: (2024)
di: Li, Changwen, et al.
Pubblicazione: (2024)
Testing Autonomous Driving Systems -- What Really Matters and What Doesn't
di: Li, Changwen, et al.
Pubblicazione: (2025)
di: Li, Changwen, et al.
Pubblicazione: (2025)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
di: Guo, Hanyang, et al.
Pubblicazione: (2025)
di: Guo, Hanyang, et al.
Pubblicazione: (2025)
On Autopilot? An Empirical Study of Human-AI Teaming and Review Practices in Open Source
di: Gao, Haoyu, et al.
Pubblicazione: (2026)
di: Gao, Haoyu, et al.
Pubblicazione: (2026)
A Viable Paradigm of Software Automation: Iterative End-to-End Automated Software Development
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
End-user Comprehension of Transfer Risks in Smart Contracts
di: Panicker, Yustynn, et al.
Pubblicazione: (2024)
di: Panicker, Yustynn, et al.
Pubblicazione: (2024)
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
di: Zeng, Zhengran, et al.
Pubblicazione: (2025)
GenIA-E2ETest: A Generative AI-Based Approach for End-to-End Test Automation
di: Júnior, Elvis, et al.
Pubblicazione: (2025)
di: Júnior, Elvis, et al.
Pubblicazione: (2025)
End-to-End Automated Logging via Multi-Agent Framework
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
Feature-Driven End-To-End Test Generation
di: Alian, Parsa, et al.
Pubblicazione: (2024)
di: Alian, Parsa, et al.
Pubblicazione: (2024)
Do Comments and Expertise Still Matter? An Experiment on Programmers' Adoption of AI-Generated JavaScript Code
di: Li, Changwen, et al.
Pubblicazione: (2025)
di: Li, Changwen, et al.
Pubblicazione: (2025)
FeaGPT: an End-to-End agentic-AI for Finite Element Analysis
di: Qi, Yupeng, et al.
Pubblicazione: (2025)
di: Qi, Yupeng, et al.
Pubblicazione: (2025)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
di: Lu, Pengrui, et al.
Pubblicazione: (2026)
VISCA: Inferring Component Abstractions for Automated End-to-End Testing
di: Alian, Parsa, et al.
Pubblicazione: (2025)
di: Alian, Parsa, et al.
Pubblicazione: (2025)
CCISolver: End-to-End Detection and Repair of Method-Level Code-Comment Inconsistency
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
di: Zhong, Renyi, et al.
Pubblicazione: (2025)
RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
di: Peng, Zhiyuan, et al.
Pubblicazione: (2026)
di: Peng, Zhiyuan, et al.
Pubblicazione: (2026)
An End-to-End Approach for Fixing Concurrency Bugs via SHB-Based Context Extractor
di: Li, Zhuang, et al.
Pubblicazione: (2026)
di: Li, Zhuang, et al.
Pubblicazione: (2026)
UniSTPA: A Safety Analysis Framework for End-to-End Autonomous Driving
di: Kou, Hongrui, et al.
Pubblicazione: (2025)
di: Kou, Hongrui, et al.
Pubblicazione: (2025)
Multi-Agent End-to-End Vulnerability Management for Mitigating Recurring Vulnerabilities
di: Zheng, Zelong, et al.
Pubblicazione: (2026)
di: Zheng, Zelong, et al.
Pubblicazione: (2026)
UniAda: Universal Adaptive Multi-objective Adversarial Attack for End-to-End Autonomous Driving Systems
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
Hallucination to Consensus: Multi-Agent LLMs for End-to-End JUnit Test Generation
di: Xu, Qinghua, et al.
Pubblicazione: (2025)
di: Xu, Qinghua, et al.
Pubblicazione: (2025)
FastLog: An End-to-End Method to Efficiently Generate and Insert Logging Statements
di: Xie, Xiaoyuan, et al.
Pubblicazione: (2023)
di: Xie, Xiaoyuan, et al.
Pubblicazione: (2023)
United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task Learning
di: He, Minghua, et al.
Pubblicazione: (2025)
di: He, Minghua, et al.
Pubblicazione: (2025)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
di: Hu, Ruida, et al.
Pubblicazione: (2026)
di: Hu, Ruida, et al.
Pubblicazione: (2026)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
di: Zhao, Zhimin, et al.
Pubblicazione: (2024)
di: Zhao, Zhimin, et al.
Pubblicazione: (2024)
WEFix: Intelligent Automatic Generation of Explicit Waits for Efficient Web End-to-End Flaky Tests
di: Liu, Xinyue, et al.
Pubblicazione: (2024)
di: Liu, Xinyue, et al.
Pubblicazione: (2024)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
WebTestPilot: Agentic End-to-End Web Testing against Natural Language Specification by Inferring Oracles with Symbolized GUI Elements
di: Teoh, Xiwen, et al.
Pubblicazione: (2026)
di: Teoh, Xiwen, et al.
Pubblicazione: (2026)
MatClaw: An Autonomous Code-First LLM Agent for End-to-End Materials Exploration
di: Zhang, Chenmu, et al.
Pubblicazione: (2026)
di: Zhang, Chenmu, et al.
Pubblicazione: (2026)
Boosting End-to-End Database Isolation Checking via Mini-Transactions (Extended Version)
di: Wei, Hengfeng, et al.
Pubblicazione: (2025)
di: Wei, Hengfeng, et al.
Pubblicazione: (2025)
Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
di: Yapağcı, Eray, et al.
Pubblicazione: (2025)
di: Yapağcı, Eray, et al.
Pubblicazione: (2025)
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
di: Chen, Zhilong, et al.
Pubblicazione: (2025)
di: Chen, Zhilong, et al.
Pubblicazione: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
di: Tran, Hung, et al.
Pubblicazione: (2026)
di: Tran, Hung, et al.
Pubblicazione: (2026)
Scalable Back-End for an AI-Based Diabetes Prediction Application
di: Radityo, Henry Anand Septian, et al.
Pubblicazione: (2025)
di: Radityo, Henry Anand Septian, et al.
Pubblicazione: (2025)
Signal-First Architectures: Rethinking Front-End Reactivity
di: Balasubramanian, Shrinivass Arunachalam
Pubblicazione: (2025)
di: Balasubramanian, Shrinivass Arunachalam
Pubblicazione: (2025)
OneLog: Towards End-to-End Training in Software Log Anomaly Detection
di: Hashemi, Shayan, et al.
Pubblicazione: (2021)
di: Hashemi, Shayan, et al.
Pubblicazione: (2021)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
di: Ni, Ziyi, et al.
Pubblicazione: (2024)
di: Ni, Ziyi, et al.
Pubblicazione: (2024)
When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning
di: Morabito, Roberto, et al.
Pubblicazione: (2025)
di: Morabito, Roberto, et al.
Pubblicazione: (2025)
E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
di: He, Keyu, et al.
Pubblicazione: (2026)
di: He, Keyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Rigorous Simulation-based Testing for Autonomous Driving Systems -- Targeting the Achilles' Heel of Four Open Autopilots
di: Li, Changwen, et al.
Pubblicazione: (2024) -
Testing Autonomous Driving Systems -- What Really Matters and What Doesn't
di: Li, Changwen, et al.
Pubblicazione: (2025) -
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
di: Guo, Hanyang, et al.
Pubblicazione: (2025) -
On Autopilot? An Empirical Study of Human-AI Teaming and Review Practices in Open Source
di: Gao, Haoyu, et al.
Pubblicazione: (2026) -
A Viable Paradigm of Software Automation: Iterative End-to-End Automated Software Development
di: Li, Jia, et al.
Pubblicazione: (2025)