TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bariah, Lina, Mefgouda, Brahim, Tavakkoli, Farbod, Molero, Enrique, Powell, Louis, Debbah, Merouane |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RF-Analyzer: Can Vision-Language Models Learn RF Understanding from Synthetic Data?
von: Bara, Anis, et al.
Veröffentlicht: (2026)
von: Bara, Anis, et al.
Veröffentlicht: (2026)
Generative AI for Immersive Communication: The Next Frontier in Internet-of-Senses Through 6G
von: Sehad, Nassim, et al.
Veröffentlicht: (2024)
von: Sehad, Nassim, et al.
Veröffentlicht: (2024)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025)
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025)
TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models
von: Zou, Hang, et al.
Veröffentlicht: (2024)
von: Zou, Hang, et al.
Veröffentlicht: (2024)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Immersive Media and Massive Twinning: Advancing Towards the Metaverse
von: Hamidouche, Wassim, et al.
Veröffentlicht: (2023)
von: Hamidouche, Wassim, et al.
Veröffentlicht: (2023)
Large Language Models for Power Scheduling: A User-Centric Approach
von: Mongaillard, Thomas, et al.
Veröffentlicht: (2024)
von: Mongaillard, Thomas, et al.
Veröffentlicht: (2024)
Diffusion-Based Generative Priors for Efficient Beam Alignment in Directional Networks
von: Othman, Esraa Fahmy, et al.
Veröffentlicht: (2026)
von: Othman, Esraa Fahmy, et al.
Veröffentlicht: (2026)
Reading Radio from Camera: Visually-Grounded, Lightweight, and Interpretable RSSI Prediction
von: Yan, Sen, et al.
Veröffentlicht: (2025)
von: Yan, Sen, et al.
Veröffentlicht: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
von: Almeida, Thales Sales, et al.
Veröffentlicht: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
von: Wang, Peng, et al.
Veröffentlicht: (2025)
von: Wang, Peng, et al.
Veröffentlicht: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
von: Bragg, Jonathan, et al.
Veröffentlicht: (2025)
Diffusion Models for Wireless Transceivers: From Pilot-Efficient Channel Estimation to AI-Native 6G Receivers
von: Yang, Yuzhi, et al.
Veröffentlicht: (2025)
von: Yang, Yuzhi, et al.
Veröffentlicht: (2025)
MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications
von: Kumar, Anshul, et al.
Veröffentlicht: (2025)
von: Kumar, Anshul, et al.
Veröffentlicht: (2025)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
von: Fang, Shicheng, et al.
Veröffentlicht: (2026)
MAPS: A Multilingual Benchmark for Agent Performance and Security
von: Hofman, Omer, et al.
Veröffentlicht: (2025)
von: Hofman, Omer, et al.
Veröffentlicht: (2025)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2024)
von: Wu, Cheng-Kuang, et al.
Veröffentlicht: (2024)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ferrag, Mohamed Amine, et al.
Veröffentlicht: (2025)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
von: He, Muyu, et al.
Veröffentlicht: (2026)
von: He, Muyu, et al.
Veröffentlicht: (2026)
Telecom World Models: Unifying Digital Twins, Foundation Models, and Predictive Planning for 6G
von: Zou, Hang, et al.
Veröffentlicht: (2026)
von: Zou, Hang, et al.
Veröffentlicht: (2026)
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
von: Wang, Wei, et al.
Veröffentlicht: (2024)
von: Wang, Wei, et al.
Veröffentlicht: (2024)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
von: Jeon, YoungHoon, et al.
Veröffentlicht: (2026)
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
von: FutureSearch, et al.
Veröffentlicht: (2025)
von: FutureSearch, et al.
Veröffentlicht: (2025)
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
von: Fu, Xing, et al.
Veröffentlicht: (2026)
von: Fu, Xing, et al.
Veröffentlicht: (2026)
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
von: Dong, Xuan, et al.
Veröffentlicht: (2026)
von: Dong, Xuan, et al.
Veröffentlicht: (2026)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
von: Diao, Lingxiao, et al.
Veröffentlicht: (2025)
von: Diao, Lingxiao, et al.
Veröffentlicht: (2025)
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
von: Tang, Xiangru, et al.
Veröffentlicht: (2025)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
von: Lee, Gyubok, et al.
Veröffentlicht: (2025)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
von: Zhang, Chao, et al.
Veröffentlicht: (2025)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
von: Ki, Dayeon, et al.
Veröffentlicht: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
von: Du, Mingxuan, et al.
Veröffentlicht: (2025)
von: Du, Mingxuan, et al.
Veröffentlicht: (2025)
SEM-RAG: Structure-Preserving Multimodal Graph Compilation and Entropy-Guided Retrieval for Telecommunication Standards
von: Yang, Yuzhi, et al.
Veröffentlicht: (2026)
von: Yang, Yuzhi, et al.
Veröffentlicht: (2026)
Goal-Oriented State Information Compression for Linear Dynamical System Control
von: Wang, Li, et al.
Veröffentlicht: (2024)
von: Wang, Li, et al.
Veröffentlicht: (2024)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
von: Deng, Shihan, et al.
Veröffentlicht: (2024)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
von: Yang, Jingbo, et al.
Veröffentlicht: (2026)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
von: Cheng, Xiang, et al.
Veröffentlicht: (2026)
von: Cheng, Xiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RF-Analyzer: Can Vision-Language Models Learn RF Understanding from Synthetic Data?
von: Bara, Anis, et al.
Veröffentlicht: (2026) -
Generative AI for Immersive Communication: The Next Frontier in Internet-of-Senses Through 6G
von: Sehad, Nassim, et al.
Veröffentlicht: (2024) -
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025) -
TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models
von: Zou, Hang, et al.
Veröffentlicht: (2024) -
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)