Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
Fuente:
arXiv
Salvato in:
| Autori principali: | Mofakhami, Mehrnaz, Bayat, Reza, Mitliagkas, Ioannis, Monteiro, Joao, Zantedeschi, Valentina |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Performative Prediction with Neural Networks
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2023)
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2023)
Learning to Defer for Causal Discovery with Imperfect Experts
di: Clivio, Oscar, et al.
Pubblicazione: (2025)
di: Clivio, Oscar, et al.
Pubblicazione: (2025)
Tight Lower Bounds and Improved Convergence in Performative Prediction
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024)
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
di: Wynn, Andrea, et al.
Pubblicazione: (2025)
di: Wynn, Andrea, et al.
Pubblicazione: (2025)
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
di: Naganuma, Hiroki, et al.
Pubblicazione: (2023)
di: Naganuma, Hiroki, et al.
Pubblicazione: (2023)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
di: Dobre, David, et al.
Pubblicazione: (2025)
di: Dobre, David, et al.
Pubblicazione: (2025)
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
di: Mahajan, Divyat, et al.
Pubblicazione: (2022)
di: Mahajan, Divyat, et al.
Pubblicazione: (2022)
Fast yet Safe: Early-Exiting with Risk Control
di: Jazbec, Metod, et al.
Pubblicazione: (2024)
di: Jazbec, Metod, et al.
Pubblicazione: (2024)
Expecting The Unexpected: Towards Broad Out-Of-Distribution Detection
di: Guille-Escuret, Charles, et al.
Pubblicazione: (2023)
di: Guille-Escuret, Charles, et al.
Pubblicazione: (2023)
Towards efficient representation identification in supervised learning
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
di: Feizi, Aarash, et al.
Pubblicazione: (2025)
di: Feizi, Aarash, et al.
Pubblicazione: (2025)
Sustainable Edge Intelligence Through Energy-Aware Early Exiting
di: Bullo, Marcello, et al.
Pubblicazione: (2023)
di: Bullo, Marcello, et al.
Pubblicazione: (2023)
Improving Prediction Certainty Estimation for Reliable Early Exiting via Null Space Projection
di: He, Jianing, et al.
Pubblicazione: (2025)
di: He, Jianing, et al.
Pubblicazione: (2025)
Steering Large Language Model Activations in Sparse Spaces
di: Bayat, Reza, et al.
Pubblicazione: (2025)
di: Bayat, Reza, et al.
Pubblicazione: (2025)
COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting Mechanism
di: He, Jianing, et al.
Pubblicazione: (2024)
di: He, Jianing, et al.
Pubblicazione: (2024)
Compositional Risk Minimization
di: Mahajan, Divyat, et al.
Pubblicazione: (2024)
di: Mahajan, Divyat, et al.
Pubblicazione: (2024)
ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
di: Zeng, Ziqian, et al.
Pubblicazione: (2023)
di: Zeng, Ziqian, et al.
Pubblicazione: (2023)
Early Exiting Predictive Coding Neural Networks for Edge AI
di: Zniber, Alaa, et al.
Pubblicazione: (2023)
di: Zniber, Alaa, et al.
Pubblicazione: (2023)
TACTiS-2: Better, Faster, Simpler Attentional Copulas for Multivariate Time Series
di: Ashok, Arjun, et al.
Pubblicazione: (2023)
di: Ashok, Arjun, et al.
Pubblicazione: (2023)
Fully Decentralized Joint Learning of Personalized Models and Collaboration Graphs
di: Zantedeschi, Valentina, et al.
Pubblicazione: (2019)
di: Zantedeschi, Valentina, et al.
Pubblicazione: (2019)
Navigating Potholes with Geometry-Aware Sharpness Minimization
di: Dufort-Labbé, Simon, et al.
Pubblicazione: (2026)
di: Dufort-Labbé, Simon, et al.
Pubblicazione: (2026)
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
di: Pan, Guanzhong, et al.
Pubblicazione: (2025)
di: Pan, Guanzhong, et al.
Pubblicazione: (2025)
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
di: Yu, Hantao, et al.
Pubblicazione: (2025)
di: Yu, Hantao, et al.
Pubblicazione: (2025)
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
di: Mahajan, Divyat, et al.
Pubblicazione: (2025)
di: Mahajan, Divyat, et al.
Pubblicazione: (2025)
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
di: Naganuma, Hiroki, et al.
Pubblicazione: (2026)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
di: Ren, Yiming, et al.
Pubblicazione: (2026)
di: Ren, Yiming, et al.
Pubblicazione: (2026)
DyCE: Dynamically Configurable Exiting for Deep Learning Compression and Real-time Scaling
di: Wang, Qingyuan, et al.
Pubblicazione: (2024)
di: Wang, Qingyuan, et al.
Pubblicazione: (2024)
FLAME: Adaptive and Reactive Concept Drift Mitigation for Federated Learning Deployments
di: Mavromatis, Ioannis, et al.
Pubblicazione: (2024)
di: Mavromatis, Ioannis, et al.
Pubblicazione: (2024)
ML Compass: Navigating Capability, Cost, and Compliance Trade-offs in AI Model Deployment
di: Digalakis Jr, Vassilis, et al.
Pubblicazione: (2025)
di: Digalakis Jr, Vassilis, et al.
Pubblicazione: (2025)
CrossQuant: A Post-Training Quantization Method with Smaller Quantization Kernel for Precise Large Language Model Compression
di: Liu, Wenyuan, et al.
Pubblicazione: (2024)
di: Liu, Wenyuan, et al.
Pubblicazione: (2024)
First Hallucination Tokens Are Different from Conditional Ones
di: Snel, Jakob, et al.
Pubblicazione: (2025)
di: Snel, Jakob, et al.
Pubblicazione: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
di: Bayat, Reza, et al.
Pubblicazione: (2024)
di: Bayat, Reza, et al.
Pubblicazione: (2024)
Multimodal Survival Analysis with Locally Deployable Large Language Models
di: Gögl, Moritz, et al.
Pubblicazione: (2026)
di: Gögl, Moritz, et al.
Pubblicazione: (2026)
Efficiently Deploying LLMs with Controlled Risk
di: Zellinger, Michael J., et al.
Pubblicazione: (2024)
di: Zellinger, Michael J., et al.
Pubblicazione: (2024)
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
di: Yin, Maxwell J., et al.
Pubblicazione: (2025)
di: Yin, Maxwell J., et al.
Pubblicazione: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
di: Griffin, Charlie, et al.
Pubblicazione: (2024)
di: Griffin, Charlie, et al.
Pubblicazione: (2024)
Understanding Adam Requires Better Rotation Dependent Assumptions
di: Zhang, Tianyue H., et al.
Pubblicazione: (2024)
di: Zhang, Tianyue H., et al.
Pubblicazione: (2024)
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
di: Zollo, Thomas P., et al.
Pubblicazione: (2023)
di: Zollo, Thomas P., et al.
Pubblicazione: (2023)
Documenti analoghi
-
Performative Prediction with Neural Networks
di: Mofakhami, Mehrnaz, et al.
Pubblicazione: (2023) -
Learning to Defer for Causal Discovery with Imperfect Experts
di: Clivio, Oscar, et al.
Pubblicazione: (2025) -
Tight Lower Bounds and Improved Convergence in Performative Prediction
di: Khorsandi, Pedram, et al.
Pubblicazione: (2024) -
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
di: Wynn, Andrea, et al.
Pubblicazione: (2025) -
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
di: Beaglehole, Daniel, et al.
Pubblicazione: (2024)