Reasoning Beyond Limits: Advances and Open Problems for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ferrag, Mohamed Amine, Tihanyi, Norbert, Debbah, Merouane |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
by: Tihanyi, Norbert, et al.
Published: (2023)
by: Tihanyi, Norbert, et al.
Published: (2023)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
by: Seddik, Mohamed El Amine, et al.
Published: (2024)
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution
by: Tihanyi, Norbert, et al.
Published: (2025)
by: Tihanyi, Norbert, et al.
Published: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
LIDSA: Cognitive Arbitration for Signal-Free Autonomous Intersection Management
by: Lakas, Abderrahmane, et al.
Published: (2026)
by: Lakas, Abderrahmane, et al.
Published: (2026)
Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward
by: Chavan, Arnav, et al.
Published: (2024)
by: Chavan, Arnav, et al.
Published: (2024)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
by: Zhang, Chao, et al.
Published: (2025)
by: Zhang, Chao, et al.
Published: (2025)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
How Small Can 6G Reason? Scaling Tiny Language Models for AI-Native Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
Securing Tomorrow's Smart Cities: Investigating Software Security in Internet of Vehicles and Deep Learning Technologies
by: Jain, Ridhi, et al.
Published: (2024)
by: Jain, Ridhi, et al.
Published: (2024)
VARS-FL: Validation-Aligned Client Selection for Non-IID Federated Learning in IoT Systems
by: Lakas, Mohamed, et al.
Published: (2026)
by: Lakas, Mohamed, et al.
Published: (2026)
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
6G Needs Agents: Toward Agentic AI-Native Networks for Autonomous Intelligence
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
$α^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
Fusion Matters: Length-Aware Analysis of Positional-Encoding Fusion in Transformers
by: Hallam, Mohamed Amine, et al.
Published: (2026)
by: Hallam, Mohamed Amine, et al.
Published: (2026)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
by: Younsi, Adam, et al.
Published: (2025)
by: Younsi, Adam, et al.
Published: (2025)
TeleOracle: Fine-Tuned Retrieval-Augmented Generation with Long-Context Support for Network
by: Alabbasi, Nouf, et al.
Published: (2024)
by: Alabbasi, Nouf, et al.
Published: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
by: Lin, Bill Yuchen, et al.
Published: (2025)
by: Lin, Bill Yuchen, et al.
Published: (2025)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
by: Qu, Yuxiao, et al.
Published: (2025)
by: Qu, Yuxiao, et al.
Published: (2025)
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
by: Pei, Qizhi, et al.
Published: (2025)
by: Pei, Qizhi, et al.
Published: (2025)
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
by: Ferrag, Mohamed Amine, et al.
Published: (2023)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
by: Shao, Zhihong, et al.
Published: (2024)
by: Shao, Zhihong, et al.
Published: (2024)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
by: Ramezanali, Mohammad, et al.
Published: (2025)
by: Ramezanali, Mohammad, et al.
Published: (2025)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
by: Duan, Jinhao, et al.
Published: (2024)
by: Duan, Jinhao, et al.
Published: (2024)
GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models
by: Loumachi, Fatma Yasmine, et al.
Published: (2024)
by: Loumachi, Fatma Yasmine, et al.
Published: (2024)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025)
by: Shojaee, Parshin, et al.
Published: (2025)
Learning to Reason in LLMs by Expectation Maximization
by: Lee, Junghyun, et al.
Published: (2025)
by: Lee, Junghyun, et al.
Published: (2025)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
MillStone: How Open-Minded Are LLMs?
by: Triedman, Harold, et al.
Published: (2025)
by: Triedman, Harold, et al.
Published: (2025)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
by: Bisztray, Tamas, et al.
Published: (2025)
by: Bisztray, Tamas, et al.
Published: (2025)
Similar Items
-
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
by: Ferrag, Mohamed Amine, et al.
Published: (2025) -
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024) -
A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification
by: Tihanyi, Norbert, et al.
Published: (2023) -
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
by: Seddik, Mohamed El Amine, et al.
Published: (2024) -
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
by: Ferrag, Mohamed Amine, et al.
Published: (2026)