Latency and Cost of Multi-Agent Intelligent Tutoring at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Elhaimeur, Iizalaarab, Chrisochoides, Nikos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ITAS: A Multi-Agent Architecture for LLM-Based Intelligent Tutoring
von: Elhaimeur, Iizalaarab, et al.
Veröffentlicht: (2026)
von: Elhaimeur, Iizalaarab, et al.
Veröffentlicht: (2026)
From Prototype to Classroom: An Intelligent Tutoring System for Quantum Education
von: Elhaimeur, Iizalaarab, et al.
Veröffentlicht: (2026)
von: Elhaimeur, Iizalaarab, et al.
Veröffentlicht: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
Efficacy of a Computer Tutor that Models Expert Human Tutors
von: Olney, Andrew M., et al.
Veröffentlicht: (2025)
von: Olney, Andrew M., et al.
Veröffentlicht: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
von: Luo, Zizhang, et al.
Veröffentlicht: (2026)
von: Luo, Zizhang, et al.
Veröffentlicht: (2026)
From 50% to Mastery in 3 Days: A Low-Resource SOP for Localizing Graduate-Level AI Tutors via Shadow-RAG
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains
von: Shintani, Seine A.
Veröffentlicht: (2026)
von: Shintani, Seine A.
Veröffentlicht: (2026)
Methodological Foundations for AI-Driven Survey Question Generation
von: Mburu, Ted K., et al.
Veröffentlicht: (2025)
von: Mburu, Ted K., et al.
Veröffentlicht: (2025)
A Fast Parallel Median Filtering Algorithm Using Hierarchical Tiling
von: Sugy, Louis
Veröffentlicht: (2025)
von: Sugy, Louis
Veröffentlicht: (2025)
From Untamed Black Box to Interpretable Pedagogical Orchestration: The Ensemble of Specialized LLMs Architecture for Adaptive Tutoring
von: Kadir, Nizam
Veröffentlicht: (2026)
von: Kadir, Nizam
Veröffentlicht: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
Assessing Crime Disclosure Patterns in a Large-Scale Cybercrime Forum
von: Hoheisel, Raphael, et al.
Veröffentlicht: (2026)
von: Hoheisel, Raphael, et al.
Veröffentlicht: (2026)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
von: Patherya, Kausar, et al.
Veröffentlicht: (2025)
von: Patherya, Kausar, et al.
Veröffentlicht: (2025)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
von: Brant, Thiago, et al.
Veröffentlicht: (2026)
von: Brant, Thiago, et al.
Veröffentlicht: (2026)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
von: Ganjihal, Sanjeev Rao
Veröffentlicht: (2026)
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
von: Iannelli, Michael, et al.
Veröffentlicht: (2024)
von: Iannelli, Michael, et al.
Veröffentlicht: (2024)
How Large Language Models Are Changing MOOC Essay Answers: A Comparison of Pre- and Post-LLM Responses
von: Leppänen, Leo, et al.
Veröffentlicht: (2025)
von: Leppänen, Leo, et al.
Veröffentlicht: (2025)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
ChatGPT and Gemini participated in the Korean College Scholastic Ability Test -- Earth Science I
von: Ga, Seok-Hyun, et al.
Veröffentlicht: (2025)
von: Ga, Seok-Hyun, et al.
Veröffentlicht: (2025)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
von: Sarker, Yeahia, et al.
Veröffentlicht: (2026)
von: Sarker, Yeahia, et al.
Veröffentlicht: (2026)
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
von: Zeng, Lingling, et al.
Veröffentlicht: (2025)
von: Zeng, Lingling, et al.
Veröffentlicht: (2025)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
von: Nouri, Azam
Veröffentlicht: (2026)
von: Nouri, Azam
Veröffentlicht: (2026)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
von: Coimbra, Bruno Moreira, et al.
Veröffentlicht: (2025)
von: Coimbra, Bruno Moreira, et al.
Veröffentlicht: (2025)
Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation
von: Hoyl, Matias
Veröffentlicht: (2026)
von: Hoyl, Matias
Veröffentlicht: (2026)
De-DSI: Decentralised Differentiable Search Index
von: Neague, Petru, et al.
Veröffentlicht: (2024)
von: Neague, Petru, et al.
Veröffentlicht: (2024)
The Reliance Negotiation Framework: A Dynamic Process Model of Student LLM Engagement in Academic Writing
von: Hossain, Shahin
Veröffentlicht: (2026)
von: Hossain, Shahin
Veröffentlicht: (2026)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
von: Zehra, Sehar, et al.
Veröffentlicht: (2025)
von: Zehra, Sehar, et al.
Veröffentlicht: (2025)
Social Dynamics of DAOs: Power, Onboarding, and Inclusivity
von: Kozlova, Victoria, et al.
Veröffentlicht: (2025)
von: Kozlova, Victoria, et al.
Veröffentlicht: (2025)
Neural Router: Semantic Content Matching for Agentic AI
von: Lovén, Lauri, et al.
Veröffentlicht: (2026)
von: Lovén, Lauri, et al.
Veröffentlicht: (2026)
Balancing Innovation and Integrity: AI Integration in Liberal Arts College Administration
von: Read, Ian Olivo
Veröffentlicht: (2025)
von: Read, Ian Olivo
Veröffentlicht: (2025)
CS-Guide: Leveraging LLMs and Student Reflections to Provide Frequent, Scalable Academic Monitoring Feedback to Computer Science Students
von: Chacko, Samuel Jacob, et al.
Veröffentlicht: (2025)
von: Chacko, Samuel Jacob, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ITAS: A Multi-Agent Architecture for LLM-Based Intelligent Tutoring
von: Elhaimeur, Iizalaarab, et al.
Veröffentlicht: (2026) -
From Prototype to Classroom: An Intelligent Tutoring System for Quantum Education
von: Elhaimeur, Iizalaarab, et al.
Veröffentlicht: (2026) -
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
von: Penke, Carolin, et al.
Veröffentlicht: (2025) -
Efficacy of a Computer Tutor that Models Expert Human Tutors
von: Olney, Andrew M., et al.
Veröffentlicht: (2025) -
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)