TelcoAgent-Bench: A Multilingual Benchmark for Telecom AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Bariah, Lina, Mefgouda, Brahim, Tavakkoli, Farbod, Molero, Enrique, Powell, Louis, Debbah, Merouane |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RF-Analyzer: Can Vision-Language Models Learn RF Understanding from Synthetic Data?
by: Bara, Anis, et al.
Published: (2026)
by: Bara, Anis, et al.
Published: (2026)
Generative AI for Immersive Communication: The Next Frontier in Internet-of-Senses Through 6G
by: Sehad, Nassim, et al.
Published: (2024)
by: Sehad, Nassim, et al.
Published: (2024)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
by: Colle, Vincenzo, et al.
Published: (2025)
by: Colle, Vincenzo, et al.
Published: (2025)
TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models
by: Zou, Hang, et al.
Published: (2024)
by: Zou, Hang, et al.
Published: (2024)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
Immersive Media and Massive Twinning: Advancing Towards the Metaverse
by: Hamidouche, Wassim, et al.
Published: (2023)
by: Hamidouche, Wassim, et al.
Published: (2023)
Large Language Models for Power Scheduling: A User-Centric Approach
by: Mongaillard, Thomas, et al.
Published: (2024)
by: Mongaillard, Thomas, et al.
Published: (2024)
Diffusion-Based Generative Priors for Efficient Beam Alignment in Directional Networks
by: Othman, Esraa Fahmy, et al.
Published: (2026)
by: Othman, Esraa Fahmy, et al.
Published: (2026)
Reading Radio from Camera: Visually-Grounded, Lightweight, and Interpretable RSSI Prediction
by: Yan, Sen, et al.
Published: (2025)
by: Yan, Sen, et al.
Published: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
by: Bragg, Jonathan, et al.
Published: (2025)
by: Bragg, Jonathan, et al.
Published: (2025)
Diffusion Models for Wireless Transceivers: From Pilot-Efficient Channel Estimation to AI-Native 6G Receivers
by: Yang, Yuzhi, et al.
Published: (2025)
by: Yang, Yuzhi, et al.
Published: (2025)
MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications
by: Kumar, Anshul, et al.
Published: (2025)
by: Kumar, Anshul, et al.
Published: (2025)
AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts
by: Fang, Shicheng, et al.
Published: (2026)
by: Fang, Shicheng, et al.
Published: (2026)
MAPS: A Multilingual Benchmark for Agent Performance and Security
by: Hofman, Omer, et al.
Published: (2025)
by: Hofman, Omer, et al.
Published: (2025)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
by: Wu, Cheng-Kuang, et al.
Published: (2024)
by: Wu, Cheng-Kuang, et al.
Published: (2024)
Reasoning Beyond Limits: Advances and Open Problems for LLMs
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
$\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
by: He, Muyu, et al.
Published: (2026)
by: He, Muyu, et al.
Published: (2026)
Telecom World Models: Unifying Digital Twins, Foundation Models, and Predictive Planning for 6G
by: Zou, Hang, et al.
Published: (2026)
by: Zou, Hang, et al.
Published: (2026)
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
by: Jeon, YoungHoon, et al.
Published: (2026)
by: Jeon, YoungHoon, et al.
Published: (2026)
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
by: Fu, Xing, et al.
Published: (2026)
by: Fu, Xing, et al.
Published: (2026)
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
by: Dong, Xuan, et al.
Published: (2026)
by: Dong, Xuan, et al.
Published: (2026)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
by: Diao, Lingxiao, et al.
Published: (2025)
by: Diao, Lingxiao, et al.
Published: (2025)
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
by: Tang, Xiangru, et al.
Published: (2025)
by: Tang, Xiangru, et al.
Published: (2025)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
by: Lee, Gyubok, et al.
Published: (2025)
by: Lee, Gyubok, et al.
Published: (2025)
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
by: Zhang, Chao, et al.
Published: (2025)
by: Zhang, Chao, et al.
Published: (2025)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
by: Du, Mingxuan, et al.
Published: (2025)
by: Du, Mingxuan, et al.
Published: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
by: Ding, Shuangrui, et al.
Published: (2026)
by: Ding, Shuangrui, et al.
Published: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
SEM-RAG: Structure-Preserving Multimodal Graph Compilation and Entropy-Guided Retrieval for Telecommunication Standards
by: Yang, Yuzhi, et al.
Published: (2026)
by: Yang, Yuzhi, et al.
Published: (2026)
Goal-Oriented State Information Compression for Linear Dynamical System Control
by: Wang, Li, et al.
Published: (2024)
by: Wang, Li, et al.
Published: (2024)
GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations
by: Yang, Jingbo, et al.
Published: (2026)
by: Yang, Jingbo, et al.
Published: (2026)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
by: Cheng, Xiang, et al.
Published: (2026)
by: Cheng, Xiang, et al.
Published: (2026)
Similar Items
-
RF-Analyzer: Can Vision-Language Models Learn RF Understanding from Synthetic Data?
by: Bara, Anis, et al.
Published: (2026) -
Generative AI for Immersive Communication: The Next Frontier in Internet-of-Senses Through 6G
by: Sehad, Nassim, et al.
Published: (2024) -
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
by: Colle, Vincenzo, et al.
Published: (2025) -
TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models
by: Zou, Hang, et al.
Published: (2024) -
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
by: Li, Xin, et al.
Published: (2025)