Cost-Aware Contrastive Routing for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Shirkavand, Reza, Gao, Shangqian, Yu, Peiran, Huang, Heng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
di: Shirkavand, Reza, et al.
Pubblicazione: (2024)
di: Shirkavand, Reza, et al.
Pubblicazione: (2024)
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
di: Ganjdanesh, Alireza, et al.
Pubblicazione: (2024)
di: Ganjdanesh, Alireza, et al.
Pubblicazione: (2024)
Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-Training
di: Shirkavand, Reza, et al.
Pubblicazione: (2025)
di: Shirkavand, Reza, et al.
Pubblicazione: (2025)
Privacy-Preserving LLMs Routing
di: Wu, Xidong, et al.
Pubblicazione: (2026)
di: Wu, Xidong, et al.
Pubblicazione: (2026)
Capability Self-Assessment: Teaching LLMs to Know Their Limits
di: Yang, Haoyan, et al.
Pubblicazione: (2026)
di: Yang, Haoyan, et al.
Pubblicazione: (2026)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
di: Ma, Shaocong, et al.
Pubblicazione: (2026)
di: Ma, Shaocong, et al.
Pubblicazione: (2026)
Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
di: Shirkavand, Reza, et al.
Pubblicazione: (2025)
di: Shirkavand, Reza, et al.
Pubblicazione: (2025)
Zeroth-Order Methods for Stochastic Nonconvex Nonsmooth Composite Optimization
di: Chen, Ziyi, et al.
Pubblicazione: (2025)
di: Chen, Ziyi, et al.
Pubblicazione: (2025)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
di: Gao, Shangqian, et al.
Pubblicazione: (2025)
di: Gao, Shangqian, et al.
Pubblicazione: (2025)
Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness
di: He, Qi, et al.
Pubblicazione: (2025)
di: He, Qi, et al.
Pubblicazione: (2025)
Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
di: Chen, Ziyi, et al.
Pubblicazione: (2025)
di: Chen, Ziyi, et al.
Pubblicazione: (2025)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
di: Xu, Zhiyuan, et al.
Pubblicazione: (2026)
di: Xu, Zhiyuan, et al.
Pubblicazione: (2026)
Dynamic Bayesian Optimization Framework for Instruction Tuning in Partial Differential Equation Discovery
di: Qu, Junqi, et al.
Pubblicazione: (2025)
di: Qu, Junqi, et al.
Pubblicazione: (2025)
Cost-Aware Routing for Efficient Text-To-Image Generation
di: Li, Qinchan, et al.
Pubblicazione: (2025)
di: Li, Qinchan, et al.
Pubblicazione: (2025)
DAK-UCB: Diversity-Aware Prompt Routing for LLMs and Generative Models
di: Jafari, Donya, et al.
Pubblicazione: (2026)
di: Jafari, Donya, et al.
Pubblicazione: (2026)
One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
di: Pulishetty, Roshini, et al.
Pubblicazione: (2025)
di: Pulishetty, Roshini, et al.
Pubblicazione: (2025)
Transformation-Augmented GRPO for Enhancing Exploration in Reasoning of Large Language Models
di: Le, Khiem, et al.
Pubblicazione: (2026)
di: Le, Khiem, et al.
Pubblicazione: (2026)
HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks
di: Zeng, Yiming, et al.
Pubblicazione: (2025)
di: Zeng, Yiming, et al.
Pubblicazione: (2025)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
di: Ding, Dujian, et al.
Pubblicazione: (2024)
di: Ding, Dujian, et al.
Pubblicazione: (2024)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
di: Tan, Qitao, et al.
Pubblicazione: (2025)
di: Tan, Qitao, et al.
Pubblicazione: (2025)
Breaking the Correlation Plateau: On the Optimization and Capacity Limits of Attention-Based Regressors
di: Yan, Jingquan, et al.
Pubblicazione: (2026)
di: Yan, Jingquan, et al.
Pubblicazione: (2026)
Auto-Train-Once: Controller Network Guided Automatic Network Pruning from Scratch
di: Wu, Xidong, et al.
Pubblicazione: (2024)
di: Wu, Xidong, et al.
Pubblicazione: (2024)
PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs
di: Xue, Jiaqi, et al.
Pubblicazione: (2025)
di: Xue, Jiaqi, et al.
Pubblicazione: (2025)
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
di: Okamoto, Mika, et al.
Pubblicazione: (2026)
di: Okamoto, Mika, et al.
Pubblicazione: (2026)
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)
di: Singla, Vasu, et al.
Pubblicazione: (2024)
DISP-LLM: Dimension-Independent Structural Pruning for Large Language Models
di: Gao, Shangqian, et al.
Pubblicazione: (2024)
di: Gao, Shangqian, et al.
Pubblicazione: (2024)
xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
Curvature Dynamic Black-box Attack: revisiting adversarial robustness via dynamic curvature estimation
di: Sun, Peiran
Pubblicazione: (2025)
di: Sun, Peiran
Pubblicazione: (2025)
Multi-intent Aware Contrastive Learning for Sequential Recommendation
di: Huang, Junshu, et al.
Pubblicazione: (2024)
di: Huang, Junshu, et al.
Pubblicazione: (2024)
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
di: Rawal, Ruchit, et al.
Pubblicazione: (2025)
di: Rawal, Ruchit, et al.
Pubblicazione: (2025)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
di: Rastegar, Reza
Pubblicazione: (2026)
di: Rastegar, Reza
Pubblicazione: (2026)
Cost-Aware Learning
di: Mohri, Clara, et al.
Pubblicazione: (2026)
di: Mohri, Clara, et al.
Pubblicazione: (2026)
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
di: Cao, Yadi, et al.
Pubblicazione: (2026)
di: Cao, Yadi, et al.
Pubblicazione: (2026)
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
di: Smith, James Seale, et al.
Pubblicazione: (2025)
di: Smith, James Seale, et al.
Pubblicazione: (2025)
Learning to Route LLMs with Confidence Tokens
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
Rethinking Constraint Awareness for Efficient State Embedding of Neural Routing Solver
di: Yu, Canhong, et al.
Pubblicazione: (2026)
di: Yu, Canhong, et al.
Pubblicazione: (2026)
Cost-TrustFL: Cost-Aware Hierarchical Federated Learning with Lightweight Reputation Evaluation across Multi-Cloud
di: Yang, Jixiao, et al.
Pubblicazione: (2025)
di: Yang, Jixiao, et al.
Pubblicazione: (2025)
RouteLLM: Learning to Route LLMs with Preference Data
di: Ong, Isaac, et al.
Pubblicazione: (2024)
di: Ong, Isaac, et al.
Pubblicazione: (2024)
Multi-Label Contrastive Learning : A Comprehensive Study
di: Audibert, Alexandre, et al.
Pubblicazione: (2024)
di: Audibert, Alexandre, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
di: Shirkavand, Reza, et al.
Pubblicazione: (2024) -
Not All Prompts Are Made Equal: Prompt-based Pruning of Text-to-Image Diffusion Models
di: Ganjdanesh, Alireza, et al.
Pubblicazione: (2024) -
Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-Training
di: Shirkavand, Reza, et al.
Pubblicazione: (2025) -
Privacy-Preserving LLMs Routing
di: Wu, Xidong, et al.
Pubblicazione: (2026) -
Capability Self-Assessment: Teaching LLMs to Know Their Limits
di: Yang, Haoyan, et al.
Pubblicazione: (2026)