MixLLM: Dynamic Routing in Mixed Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Xinyuan, Liu, Yanchi, Cheng, Wei, Zhao, Xujiang, Chen, Zhengzhang, Yu, Wenchao, Fu, Yanjie, Chen, Haifeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DBCopilot: Natural Language Querying over Massive Databases via Schema Routing
di: Wang, Tianshu, et al.
Pubblicazione: (2023)
di: Wang, Tianshu, et al.
Pubblicazione: (2023)
Mixed-Precision Embeddings for Large-Scale Recommendation Models
di: Li, Shiwei, et al.
Pubblicazione: (2024)
di: Li, Shiwei, et al.
Pubblicazione: (2024)
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
di: Yuan, Bo, et al.
Pubblicazione: (2026)
di: Yuan, Bo, et al.
Pubblicazione: (2026)
Taxonomy Inference for Tabular Data Using Large Language Models
di: Wu, Zhenyu, et al.
Pubblicazione: (2025)
di: Wu, Zhenyu, et al.
Pubblicazione: (2025)
Assessing SPARQL capabilities of Large Language Models
di: Meyer, Lars-Peter, et al.
Pubblicazione: (2024)
di: Meyer, Lars-Peter, et al.
Pubblicazione: (2024)
Towards Universal Dense Blocking for Entity Resolution
di: Wang, Tianshu, et al.
Pubblicazione: (2024)
di: Wang, Tianshu, et al.
Pubblicazione: (2024)
CARROT: A Learned Cost-Constrained Retrieval Optimization System for RAG
di: Wang, Ziting, et al.
Pubblicazione: (2024)
di: Wang, Ziting, et al.
Pubblicazione: (2024)
GaussMaster: An LLM-based Database Copilot System
di: Zhou, Wei, et al.
Pubblicazione: (2025)
di: Zhou, Wei, et al.
Pubblicazione: (2025)
An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models
di: Wang, Mengzhao, et al.
Pubblicazione: (2024)
di: Wang, Mengzhao, et al.
Pubblicazione: (2024)
A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection
di: Reis, Philipp, et al.
Pubblicazione: (2026)
di: Reis, Philipp, et al.
Pubblicazione: (2026)
Editing Conceptual Knowledge for Large Language Models
di: Wang, Xiaohan, et al.
Pubblicazione: (2024)
di: Wang, Xiaohan, et al.
Pubblicazione: (2024)
A Survey of LLM $\times$ DATA
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2025)
QPAD: Quantile-Preserving Approximate Dimension Reduction for Nearest Neighbors Preservation in High-Dimensional Vector Search
di: Fu, Jiuzhou, et al.
Pubblicazione: (2025)
di: Fu, Jiuzhou, et al.
Pubblicazione: (2025)
PentaRAG: Large-Scale Intelligent Knowledge Retrieval for Enterprise LLM Applications
di: Syarubany, Abu Hanif Muhammad, et al.
Pubblicazione: (2025)
di: Syarubany, Abu Hanif Muhammad, et al.
Pubblicazione: (2025)
TabSQLify: Enhancing Reasoning Capabilities of LLMs Through Table Decomposition
di: Nahid, Md Mahadi Hasan, et al.
Pubblicazione: (2024)
di: Nahid, Md Mahadi Hasan, et al.
Pubblicazione: (2024)
GRASP: Generic Reasoning And SPARQL Generation across Knowledge Graphs
di: Walter, Sebastian, et al.
Pubblicazione: (2025)
di: Walter, Sebastian, et al.
Pubblicazione: (2025)
Optimization of embeddings storage for RAG systems using quantization and dimensionality reduction techniques
di: Huerga-Pérez, Naamán, et al.
Pubblicazione: (2025)
di: Huerga-Pérez, Naamán, et al.
Pubblicazione: (2025)
Knowledge Graph-based Retrieval-Augmented Generation for Schema Matching
di: Ma, Chuangtao, et al.
Pubblicazione: (2025)
di: Ma, Chuangtao, et al.
Pubblicazione: (2025)
Domain Specific Question to SQL Conversion with Embedded Data Balancing Technique
di: Jyothi, et al.
Pubblicazione: (2025)
di: Jyothi, et al.
Pubblicazione: (2025)
Dr Web: a modern, query-based web data retrieval engine
di: Prifti, Ylli, et al.
Pubblicazione: (2025)
di: Prifti, Ylli, et al.
Pubblicazione: (2025)
In-depth Analysis of Graph-based RAG in a Unified Framework
di: Zhou, Yingli, et al.
Pubblicazione: (2025)
di: Zhou, Yingli, et al.
Pubblicazione: (2025)
Retrieval Augmented Generation using Engineering Design Knowledge
di: Siddharth, L., et al.
Pubblicazione: (2023)
di: Siddharth, L., et al.
Pubblicazione: (2023)
CrackSQL: A Hybrid SQL Dialect Translation System Powered by Large Language Models
di: Zhou, Wei, et al.
Pubblicazione: (2025)
di: Zhou, Wei, et al.
Pubblicazione: (2025)
LIRA: A Learning-based Query-aware Partition Framework for Large-scale ANN Search
di: Zeng, Ximu, et al.
Pubblicazione: (2025)
di: Zeng, Ximu, et al.
Pubblicazione: (2025)
A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives
di: Li, Pengyue, et al.
Pubblicazione: (2025)
di: Li, Pengyue, et al.
Pubblicazione: (2025)
Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems
di: Pham, Hoang, et al.
Pubblicazione: (2025)
di: Pham, Hoang, et al.
Pubblicazione: (2025)
Path-Constrained Retrieval: A Structural Approach to Reliable LLM Agent Reasoning Through Graph-Scoped Semantic Search
di: Oladokun, Joseph
Pubblicazione: (2025)
di: Oladokun, Joseph
Pubblicazione: (2025)
Access Paths for Efficient Ordering with Large Language Models
di: Zhao, Fuheng, et al.
Pubblicazione: (2025)
di: Zhao, Fuheng, et al.
Pubblicazione: (2025)
MURAD: A Large-Scale Multi-Domain Unified Reverse Arabic Dictionary Dataset
di: Sibaee, Serry, et al.
Pubblicazione: (2026)
di: Sibaee, Serry, et al.
Pubblicazione: (2026)
KeyInst: Keyword Instruction for Improving SQL Formulation in Text-to-SQL
di: Liu, Xiping, et al.
Pubblicazione: (2024)
di: Liu, Xiping, et al.
Pubblicazione: (2024)
Stitching Inner Product and Euclidean Metrics for Topology-aware Maximum Inner Product Search
di: Chen, Tingyang, et al.
Pubblicazione: (2025)
di: Chen, Tingyang, et al.
Pubblicazione: (2025)
OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System
di: Luo, Yujie, et al.
Pubblicazione: (2024)
di: Luo, Yujie, et al.
Pubblicazione: (2024)
MMAG: Mixed Memory-Augmented Generation for Large Language Models Applications
di: Zeppieri, Stefano
Pubblicazione: (2025)
di: Zeppieri, Stefano
Pubblicazione: (2025)
AlayaLaser: Efficient Index Layout and Search Strategy for Large-scale High-dimensional Vector Similarity Search
di: Chen, Weijian, et al.
Pubblicazione: (2026)
di: Chen, Weijian, et al.
Pubblicazione: (2026)
LR-SQL: A Supervised Fine-Tuning Method for Text2SQL Tasks under Low-Resource Scenarios
di: Wuzhenghong, Wen, et al.
Pubblicazione: (2024)
di: Wuzhenghong, Wen, et al.
Pubblicazione: (2024)
SIEVE: Effective Filtered Vector Search with Collection of Indexes
di: Li, Zhaoheng, et al.
Pubblicazione: (2025)
di: Li, Zhaoheng, et al.
Pubblicazione: (2025)
IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus
di: Gui, Honghao, et al.
Pubblicazione: (2024)
di: Gui, Honghao, et al.
Pubblicazione: (2024)
Efficient Data-aware Distance Comparison Operations for High-Dimensional Approximate Nearest Neighbor Search
di: Deng, Liwei, et al.
Pubblicazione: (2024)
di: Deng, Liwei, et al.
Pubblicazione: (2024)
Factual Inconsistencies in Multilingual Wikipedia Tables
di: Cappa, Silvia, et al.
Pubblicazione: (2025)
di: Cappa, Silvia, et al.
Pubblicazione: (2025)
Reveal Hidden Pitfalls and Navigate Next Generation of Vector Similarity Search from Task-Centric Views
di: Chen, Tingyang, et al.
Pubblicazione: (2025)
di: Chen, Tingyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DBCopilot: Natural Language Querying over Massive Databases via Schema Routing
di: Wang, Tianshu, et al.
Pubblicazione: (2023) -
Mixed-Precision Embeddings for Large-Scale Recommendation Models
di: Li, Shiwei, et al.
Pubblicazione: (2024) -
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
di: Yuan, Bo, et al.
Pubblicazione: (2026) -
Taxonomy Inference for Tabular Data Using Large Language Models
di: Wu, Zhenyu, et al.
Pubblicazione: (2025) -
Assessing SPARQL capabilities of Large Language Models
di: Meyer, Lars-Peter, et al.
Pubblicazione: (2024)