Dynamic Intelligence Assessment: Benchmarking LLMs on the Road to AGI with a Focus on Model Confidence
Fuente:
arXiv
Guardado en:
| Autores principales: | Tihanyi, Norbert, Bisztray, Tamas, Dubniczky, Richard A., Toth, Rebeka, Borsos, Bertalan, Cherif, Bilel, Ferrag, Mohamed Amine, Muzsai, Lajos, Jain, Ridhi, Marinelli, Ryan, Cordeiro, Lucas C., Debbah, Merouane, Mavroeidis, Vasileios, Josang, Audun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
por: Bisztray, Tamas, et al.
Publicado: (2025)
por: Bisztray, Tamas, et al.
Publicado: (2025)
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
por: Tihanyi, Norbert, et al.
Publicado: (2025)
por: Tihanyi, Norbert, et al.
Publicado: (2025)
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
por: Tihanyi, Norbert, et al.
Publicado: (2025)
por: Tihanyi, Norbert, et al.
Publicado: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
por: Tihanyi, Norbert, et al.
Publicado: (2024)
por: Tihanyi, Norbert, et al.
Publicado: (2024)
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
por: Dubniczky, Richard A., et al.
Publicado: (2025)
por: Dubniczky, Richard A., et al.
Publicado: (2025)
The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification
por: Tihanyi, Norbert, et al.
Publicado: (2023)
por: Tihanyi, Norbert, et al.
Publicado: (2023)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
por: Ferrag, Mohamed Amine, et al.
Publicado: (2024)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2024)
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
por: Tihanyi, Norbert, et al.
Publicado: (2024)
por: Tihanyi, Norbert, et al.
Publicado: (2024)
DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response
por: Cherif, Bilel, et al.
Publicado: (2025)
por: Cherif, Bilel, et al.
Publicado: (2025)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection
por: Dubniczky, Richard A., et al.
Publicado: (2025)
por: Dubniczky, Richard A., et al.
Publicado: (2025)
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
por: Jain, Ridhi, et al.
Publicado: (2024)
por: Jain, Ridhi, et al.
Publicado: (2024)
Sustaining Cyber Awareness: The Long-Term Impact of Continuous Phishing Training and Emotional Triggers
por: Toth, Rebeka, et al.
Publicado: (2025)
por: Toth, Rebeka, et al.
Publicado: (2025)
The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs
por: Toth, Rebeka, et al.
Publicado: (2025)
por: Toth, Rebeka, et al.
Publicado: (2025)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
por: Tóth, Rebeka, et al.
Publicado: (2024)
por: Tóth, Rebeka, et al.
Publicado: (2024)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
por: Tihanyi, Norbert, et al.
Publicado: (2023)
por: Tihanyi, Norbert, et al.
Publicado: (2023)
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2025)
$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
6G Needs Agents: Toward Agentic AI-Native Networks for Autonomous Intelligence
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
LIDSA: Cognitive Arbitration for Signal-Free Autonomous Intersection Management
por: Lakas, Abderrahmane, et al.
Publicado: (2026)
por: Lakas, Abderrahmane, et al.
Publicado: (2026)
How Small Can 6G Reason? Scaling Tiny Language Models for AI-Native Networks
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
$α^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2026)
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices
por: Ferrag, Mohamed Amine, et al.
Publicado: (2023)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2023)
Guest Editors' Introduction. Trust and Trust Management
por: Audun Jøsang
Publicado: (2010)
por: Audun Jøsang
Publicado: (2010)
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs?
por: Ferrag, Mohamed Amine, et al.
Publicado: (2023)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2023)
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection
por: Marinelli, Ryan, et al.
Publicado: (2025)
por: Marinelli, Ryan, et al.
Publicado: (2025)
Edge Learning for 6G-enabled Internet of Things: A Comprehensive Survey of Vulnerabilities, Datasets, and Defenses
por: Ferrag, Mohamed Amine, et al.
Publicado: (2023)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2023)
TIPS: Threat Sharing Information Platform for Enhanced Security
por: Pasumarthy, Lakshmi Rama Kiran, et al.
Publicado: (2024)
por: Pasumarthy, Lakshmi Rama Kiran, et al.
Publicado: (2024)
Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
por: Muzsai, Lajos, et al.
Publicado: (2025)
por: Muzsai, Lajos, et al.
Publicado: (2025)
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
por: Muzsai, Lajos, et al.
Publicado: (2024)
por: Muzsai, Lajos, et al.
Publicado: (2024)
Adversarial Attacks and Defenses in 6G Network-Assisted IoT Systems
por: Son, Bui Duc, et al.
Publicado: (2024)
por: Son, Bui Duc, et al.
Publicado: (2024)
VARS-FL: Validation-Aligned Client Selection for Non-IID Federated Learning in IoT Systems
por: Lakas, Mohamed, et al.
Publicado: (2026)
por: Lakas, Mohamed, et al.
Publicado: (2026)
LlamBERT: Large-scale low-cost data annotation in NLP
por: Csanády, Bálint, et al.
Publicado: (2024)
por: Csanády, Bálint, et al.
Publicado: (2024)
Towards Incident Response Orchestration and Automation for the Advanced Metering Infrastructure
por: Lekidis, Alexios, et al.
Publicado: (2024)
por: Lekidis, Alexios, et al.
Publicado: (2024)
LLM-Powered Intent-Based Categorization of Phishing Emails
por: Eilertsen, Even, et al.
Publicado: (2025)
por: Eilertsen, Even, et al.
Publicado: (2025)
Towards Agentic Investigation of Security Alerts
por: Eilertsen, Even, et al.
Publicado: (2026)
por: Eilertsen, Even, et al.
Publicado: (2026)
Ejemplares similares
-
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
por: Bisztray, Tamas, et al.
Publicado: (2025) -
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
por: Tihanyi, Norbert, et al.
Publicado: (2025) -
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
por: Tihanyi, Norbert, et al.
Publicado: (2025) -
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
por: Tihanyi, Norbert, et al.
Publicado: (2024) -
You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models
por: Dubniczky, Richard A., et al.
Publicado: (2025)