When Does Confidence-Based Cascade Deferral Suffice?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jitkrittum, Wittawat, Gupta, Neha, Menon, Aditya Krishna, Narasimhan, Harikrishna, Rawat, Ankit Singh, Kumar, Sanjiv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
Cascade-Aware Training of Language Models
von: Wang, Congchao, et al.
Veröffentlicht: (2024)
von: Wang, Congchao, et al.
Veröffentlicht: (2024)
Universal Model Routing for Efficient LLM Inference
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025)
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025)
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
von: Lukasik, Michal, et al.
Veröffentlicht: (2025)
von: Lukasik, Michal, et al.
Veröffentlicht: (2025)
Gatekeeper: Improving Model Cascades Through Confidence Tuning
von: Rabanser, Stephan, et al.
Veröffentlicht: (2025)
von: Rabanser, Stephan, et al.
Veröffentlicht: (2025)
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
von: Rawat, Ankit Singh, et al.
Veröffentlicht: (2024)
von: Rawat, Ankit Singh, et al.
Veröffentlicht: (2024)
Regression-aware Inference with LLMs
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
Efficient Document Ranking with Learnable Late Interactions
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
On student-teacher deviations in distillation: does it pay to disobey?
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2023)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2023)
When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents
von: Rawat, Mrinal, et al.
Veröffentlicht: (2025)
von: Rawat, Mrinal, et al.
Veröffentlicht: (2025)
Cost-Aware Routing for Efficient Text-To-Image Generation
von: Li, Qinchan, et al.
Veröffentlicht: (2025)
von: Li, Qinchan, et al.
Veröffentlicht: (2025)
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
von: Farinhas, António, et al.
Veröffentlicht: (2025)
von: Farinhas, António, et al.
Veröffentlicht: (2025)
Benchmarking with MIMIC-IV, an irregular, spare clinical time series dataset
von: Bui, Hung, et al.
Veröffentlicht: (2024)
von: Bui, Hung, et al.
Veröffentlicht: (2024)
Regression with Multi-Expert Deferral
von: Mao, Anqi, et al.
Veröffentlicht: (2024)
von: Mao, Anqi, et al.
Veröffentlicht: (2024)
Budgeted Multiple-Expert Deferral
von: DeSalvo, Giulia, et al.
Veröffentlicht: (2025)
von: DeSalvo, Giulia, et al.
Veröffentlicht: (2025)
Optimized Deferral for Imbalanced Settings
von: Cortes, Corinna, et al.
Veröffentlicht: (2026)
von: Cortes, Corinna, et al.
Veröffentlicht: (2026)
Logarithmic Width Suffices for Robust Memorization
von: Egosi, Amitsour, et al.
Veröffentlicht: (2025)
von: Egosi, Amitsour, et al.
Veröffentlicht: (2025)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
Sparse Robust Classification via the Kernel Mean
von: van Rooyen, Brendan, et al.
Veröffentlicht: (2015)
von: van Rooyen, Brendan, et al.
Veröffentlicht: (2015)
Immediate Derivatives Suffice for Online Recurrent Adaptation
von: Merin, Aur Shalev
Veröffentlicht: (2026)
von: Merin, Aur Shalev
Veröffentlicht: (2026)
Online Decision Deferral under Budget Constraints
von: Reid, Mirabel, et al.
Veröffentlicht: (2024)
von: Reid, Mirabel, et al.
Veröffentlicht: (2024)
NeuroMemFPP: A recurrent neural approach for memory-aware parameter estimation in fractional Poisson process
von: Gupta, Neha, et al.
Veröffentlicht: (2025)
von: Gupta, Neha, et al.
Veröffentlicht: (2025)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Robust Validation: Confident Predictions Even When Distributions Shift
von: Cauchois, Maxime, et al.
Veröffentlicht: (2020)
von: Cauchois, Maxime, et al.
Veröffentlicht: (2020)
When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal
von: Phalod, Aditya Ajay
Veröffentlicht: (2026)
von: Phalod, Aditya Ajay
Veröffentlicht: (2026)
Querying Kernel Methods Suffices for Reconstructing their Training Data
von: Barzilai, Daniel, et al.
Veröffentlicht: (2025)
von: Barzilai, Daniel, et al.
Veröffentlicht: (2025)
Dual-Encoders for Extreme Multi-Label Classification
von: Gupta, Nilesh, et al.
Veröffentlicht: (2023)
von: Gupta, Nilesh, et al.
Veröffentlicht: (2023)
Bi-directional Model Cascading with Proxy Confidence
von: Warren, David, et al.
Veröffentlicht: (2025)
von: Warren, David, et al.
Veröffentlicht: (2025)
Identity-Free Deferral For Unseen Experts
von: Strong, Joshua, et al.
Veröffentlicht: (2025)
von: Strong, Joshua, et al.
Veröffentlicht: (2025)
The importance of feature preprocessing for differentially private linear optimization
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
Adaptive Regret for Bandits Made Possible: Two Queries Suffice
von: Lu, Zhou, et al.
Veröffentlicht: (2024)
von: Lu, Zhou, et al.
Veröffentlicht: (2024)
Towards Trustworthy AI Software Development Assistance
von: Maninger, Daniel, et al.
Veröffentlicht: (2023)
von: Maninger, Daniel, et al.
Veröffentlicht: (2023)
Investigation of Compressor Cascade Flow Using Physics- Informed Neural Networks with Adaptive Learning Strategy
von: Li, Zhihui, et al.
Veröffentlicht: (2023)
von: Li, Zhihui, et al.
Veröffentlicht: (2023)
HybridServe: Efficient Serving of Large AI Models with Confidence-Based Cascade Routing
von: Xue, Leyang, et al.
Veröffentlicht: (2025)
von: Xue, Leyang, et al.
Veröffentlicht: (2025)
Fine Tuning LLM for Enterprise: Practical Guidelines and Recommendations
von: J, Mathav Raj, et al.
Veröffentlicht: (2024)
von: J, Mathav Raj, et al.
Veröffentlicht: (2024)
Subsampling Suffices for Adaptive Data Analysis
von: Blanc, Guy
Veröffentlicht: (2023)
von: Blanc, Guy
Veröffentlicht: (2023)
Feature Augmentation of GNNs for ILPs: Local Uniqueness Suffices
von: Han, Qingyu, et al.
Veröffentlicht: (2025)
von: Han, Qingyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024) -
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024) -
Cascade-Aware Training of Language Models
von: Wang, Congchao, et al.
Veröffentlicht: (2024) -
Universal Model Routing for Efficient LLM Inference
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025) -
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
von: Lukasik, Michal, et al.
Veröffentlicht: (2025)