AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Ferrag, Mohamed Amine, Lakas, Abderrahmane, Debbah, Merouane |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
6G Needs Agents: Toward Agentic AI-Native Networks for Autonomous Intelligence
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
How Small Can 6G Reason? Scaling Tiny Language Models for AI-Native Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
$α^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
LIDSA: Cognitive Arbitration for Signal-Free Autonomous Intersection Management
by: Lakas, Abderrahmane, et al.
Published: (2026)
by: Lakas, Abderrahmane, et al.
Published: (2026)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
VARS-FL: Validation-Aligned Client Selection for Non-IID Federated Learning in IoT Systems
by: Lakas, Mohamed, et al.
Published: (2026)
by: Lakas, Mohamed, et al.
Published: (2026)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
A Tutorial on Cognitive Biases in Agentic AI-Driven 6G Autonomous Networks
by: Chergui, Hatim, et al.
Published: (2025)
by: Chergui, Hatim, et al.
Published: (2025)
The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification
by: Tihanyi, Norbert, et al.
Published: (2023)
by: Tihanyi, Norbert, et al.
Published: (2023)
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
MX-AI: Agentic Observability and Control Platform for Open and AI-RAN
by: Chatzistefanidis, Ilias, et al.
Published: (2025)
by: Chatzistefanidis, Ilias, et al.
Published: (2025)
Advanced Architectures Integrated with Agentic AI for Next-Generation Wireless Networks
by: Dev, Kapal, et al.
Published: (2025)
by: Dev, Kapal, et al.
Published: (2025)
GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models
by: Loumachi, Fatma Yasmine, et al.
Published: (2024)
by: Loumachi, Fatma Yasmine, et al.
Published: (2024)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
by: Dubniczky, Richard A., et al.
Published: (2025)
by: Dubniczky, Richard A., et al.
Published: (2025)
LLM-Based Agentic Negotiation for 6G: Addressing Uncertainty Neglect and Tail-Event Risk
by: Chergui, Hatim, et al.
Published: (2025)
by: Chergui, Hatim, et al.
Published: (2025)
Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics
by: Liu, Junqi, et al.
Published: (2026)
by: Liu, Junqi, et al.
Published: (2026)
Dynamic Intelligence Assessment: Benchmarking LLMs on the Road to AGI with a Focus on Model Confidence
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Integration of Agentic AI with 6G Networks for Mission-Critical Applications: Use-case and Challenges
by: Khowaja, Sunder Ali, et al.
Published: (2025)
by: Khowaja, Sunder Ali, et al.
Published: (2025)
Agentic AI for Self-Driving Laboratories in Soft Matter: Taxonomy, Benchmarks,and Open Challenges
by: Chen, Xuanzhou, et al.
Published: (2026)
by: Chen, Xuanzhou, et al.
Published: (2026)
Safety2Drive: Safety-Critical Scenario Benchmark for the Evaluation of Autonomous Driving
by: Li, Jingzheng, et al.
Published: (2025)
by: Li, Jingzheng, et al.
Published: (2025)
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
by: Ma, Yunsheng, et al.
Published: (2023)
by: Ma, Yunsheng, et al.
Published: (2023)
Non-Identical Diffusion Models in MIMO-OFDM Channel Generation
by: Yang, Yuzhi, et al.
Published: (2025)
by: Yang, Yuzhi, et al.
Published: (2025)
Diffusion-Based Generative Priors for Efficient Beam Alignment in Directional Networks
by: Othman, Esraa Fahmy, et al.
Published: (2026)
by: Othman, Esraa Fahmy, et al.
Published: (2026)
From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
by: Jiang, Feibo, et al.
Published: (2025)
by: Jiang, Feibo, et al.
Published: (2025)
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
by: Colle, Vincenzo, et al.
Published: (2025)
by: Colle, Vincenzo, et al.
Published: (2025)
Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms
by: Yu, Kejin, et al.
Published: (2026)
by: Yu, Kejin, et al.
Published: (2026)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
by: Paglieri, Davide, et al.
Published: (2024)
by: Paglieri, Davide, et al.
Published: (2024)
Towards Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-critical Scenarios
by: Li, Jingzheng, et al.
Published: (2025)
by: Li, Jingzheng, et al.
Published: (2025)
AgentRAN: An Agentic AI Architecture for Autonomous Control of Open 6G Networks
by: Elkael, Maxime, et al.
Published: (2025)
by: Elkael, Maxime, et al.
Published: (2025)
ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving
by: Ruan, Qiyu, et al.
Published: (2026)
by: Ruan, Qiyu, et al.
Published: (2026)
Similar Items
-
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
by: Ferrag, Mohamed Amine, et al.
Published: (2025) -
6G Needs Agents: Toward Agentic AI-Native Networks for Autonomous Intelligence
by: Ferrag, Mohamed Amine, et al.
Published: (2026) -
6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026) -
$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026) -
How Small Can 6G Reason? Scaling Tiny Language Models for AI-Native Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)