NinjaLLM: Fast, Scalable and Cost-effective RAG using Amazon SageMaker and AWS Trainium and Inferentia2
Fuente:
arXiv
Guardado en:
| Autores principales: | Xue, Tengfei, Li, Xuefeng, Smirnov, Roman, Azim, Tahir, Sadrieh, Arash, Pahlavan, Babak |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Programming Language Ensemble for Code Generation in Large Language Model
por: Xue, Tengfei, et al.
Publicado: (2024)
por: Xue, Tengfei, et al.
Publicado: (2024)
YOLO based Ocean Eddy Localization with AWS SageMaker
por: Mostafa, Seraj Al Mahmud, et al.
Publicado: (2024)
por: Mostafa, Seraj Al Mahmud, et al.
Publicado: (2024)
GPU Programming for AI Workflow Development on AWS SageMaker: An Instructional Approach
por: Srinivasan, Sriram, et al.
Publicado: (2025)
por: Srinivasan, Sriram, et al.
Publicado: (2025)
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
por: Lee, Myeonghwa, et al.
Publicado: (2024)
por: Lee, Myeonghwa, et al.
Publicado: (2024)
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
por: Fan, Haozheng, et al.
Publicado: (2024)
por: Fan, Haozheng, et al.
Publicado: (2024)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
por: Song, Dinghong, et al.
Publicado: (2025)
por: Song, Dinghong, et al.
Publicado: (2025)
Fine-Tuning Open-Weight Language Models to Deliver Cognitive Behavioral Therapy for Depression: A Feasibility Study
por: Tahir, Talha
Publicado: (2024)
por: Tahir, Talha
Publicado: (2024)
Partial order on involutive permutations and double Schubert cells
por: Smirnov, Evgeny
Publicado: (2024)
por: Smirnov, Evgeny
Publicado: (2024)
FlexStructRAG: Flexible Structure-Aware Multi-Granular Relational Retrieval for RAG
por: Chen, Mengzhu, et al.
Publicado: (2026)
por: Chen, Mengzhu, et al.
Publicado: (2026)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
por: Hu, Junyi, et al.
Publicado: (2024)
por: Hu, Junyi, et al.
Publicado: (2024)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
por: Qi, Jinhu, et al.
Publicado: (2024)
por: Qi, Jinhu, et al.
Publicado: (2024)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
por: Li, Xu, et al.
Publicado: (2026)
por: Li, Xu, et al.
Publicado: (2026)
ARAGOG: Advanced RAG Output Grading
por: Eibich, Matouš, et al.
Publicado: (2024)
por: Eibich, Matouš, et al.
Publicado: (2024)
Correctness is not Faithfulness in RAG Attributions
por: Wallat, Jonas, et al.
Publicado: (2024)
por: Wallat, Jonas, et al.
Publicado: (2024)
Persona Inconstancy in Multi-Agent LLM Collaboration: Conformity, Confabulation, and Impersonation
por: Baltaji, Razan, et al.
Publicado: (2024)
por: Baltaji, Razan, et al.
Publicado: (2024)
Physical oceanography during Walther Herwig I cruise WH27
por: Anonymous
Publicado: (2011)
por: Anonymous
Publicado: (2011)
Primary production within the euphotic layer of the Norwegian Sea in July 1997: measurements by different methods
por: Sapozhnikov, Victor V, et al.
Publicado: (2000)
por: Sapozhnikov, Victor V, et al.
Publicado: (2000)
(Table 2) Composition of polycyclic aromatics in bottom sediments from the Weddell Sea
por: Danyushevskaya, A I, et al.
Publicado: (1989)
por: Danyushevskaya, A I, et al.
Publicado: (1989)
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
por: Fiaz, Layba, et al.
Publicado: (2025)
por: Fiaz, Layba, et al.
Publicado: (2025)
Hydrochemistry measured on water bottle samples during Professor Kolesnikov cruise PK27
por: Piontkovski, Sergey, et al.
Publicado: (2011)
por: Piontkovski, Sergey, et al.
Publicado: (2011)
Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
por: Sorstkins, Andrejs
Publicado: (2025)
por: Sorstkins, Andrejs
Publicado: (2025)
Observations on Building RAG Systems for Technical Documents
por: Soman, Sumit, et al.
Publicado: (2024)
por: Soman, Sumit, et al.
Publicado: (2024)
KGiRAG: An Iterative GraphRAG Approach for Responding Sensemaking Queries
por: Iacob, Isabela, et al.
Publicado: (2026)
por: Iacob, Isabela, et al.
Publicado: (2026)
RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
por: Wang, Yongjie, et al.
Publicado: (2025)
por: Wang, Yongjie, et al.
Publicado: (2025)
UniC-RAG: Universal Knowledge Corruption Attacks to Retrieval-Augmented Generation
por: Geng, Runpeng, et al.
Publicado: (2025)
por: Geng, Runpeng, et al.
Publicado: (2025)
FACTUM: Mechanistic Detection of Citation Hallucination in Long-Form RAG
por: Dassen, Maxime, et al.
Publicado: (2026)
por: Dassen, Maxime, et al.
Publicado: (2026)
Physical oceanography during Maria S. Merian cruise MSM27
por: Schneider, Linn, et al.
Publicado: (2015)
por: Schneider, Linn, et al.
Publicado: (2015)
HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents
por: Jiang, Yilin, et al.
Publicado: (2026)
por: Jiang, Yilin, et al.
Publicado: (2026)
Synthia: Scalable Grounded Persona Generation from Social Media Data
por: Rahimzadeh, Vahid, et al.
Publicado: (2025)
por: Rahimzadeh, Vahid, et al.
Publicado: (2025)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
por: Tahir, Munief Hassan, et al.
Publicado: (2024)
por: Tahir, Munief Hassan, et al.
Publicado: (2024)
Composition of bottom sediments from the Mediterranean Sea off Syria
por: Shimkus, Kazimir M, et al.
Publicado: (2000)
por: Shimkus, Kazimir M, et al.
Publicado: (2000)
Physical oceanography, CFC-11 and CFC-12 measured on water bottle samples during Maria S. Merian cruise MSM27
por: Kieke, Dagmar, et al.
Publicado: (2020)
por: Kieke, Dagmar, et al.
Publicado: (2020)
Physical oceanography during POLARSTERN cruise ARK-IX/4
por: Schauer, Ursula, et al.
Publicado: (2010)
por: Schauer, Ursula, et al.
Publicado: (2010)
Contextually Aware E-Commerce Product Question Answering using RAG
por: Tangarajan, Praveen, et al.
Publicado: (2025)
por: Tangarajan, Praveen, et al.
Publicado: (2025)
Chlorophyll a measured on water bottle samples during POLARSTERN cruise ARK-IX/4
por: Nöthig, Eva-Maria, et al.
Publicado: (2015)
por: Nöthig, Eva-Maria, et al.
Publicado: (2015)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
por: Lai, Junyu, et al.
Publicado: (2025)
por: Lai, Junyu, et al.
Publicado: (2025)
Chemical oceanography during POLARSTERN cruise ARK-IX/4
por: Luchetta, Anna, et al.
Publicado: (2021)
por: Luchetta, Anna, et al.
Publicado: (2021)
Bidirectional RAG: Safe Self-Improving Retrieval-Augmented Generation Through Multi-Stage Validation
por: Chinthala, Teja
Publicado: (2025)
por: Chinthala, Teja
Publicado: (2025)
Legal RAG Bench: an end-to-end benchmark for legal RAG
por: Butler, Abdur-Rahman, et al.
Publicado: (2026)
por: Butler, Abdur-Rahman, et al.
Publicado: (2026)
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
por: Fang, Bowen, et al.
Publicado: (2026)
por: Fang, Bowen, et al.
Publicado: (2026)
Ejemplares similares
-
Multi-Programming Language Ensemble for Code Generation in Large Language Model
por: Xue, Tengfei, et al.
Publicado: (2024) -
YOLO based Ocean Eddy Localization with AWS SageMaker
por: Mostafa, Seraj Al Mahmud, et al.
Publicado: (2024) -
GPU Programming for AI Workflow Development on AWS SageMaker: An Instructional Approach
por: Srinivasan, Sriram, et al.
Publicado: (2025) -
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers
por: Lee, Myeonghwa, et al.
Publicado: (2024) -
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
por: Fan, Haozheng, et al.
Publicado: (2024)