Two-dimensional early exit optimisation of LLM inference
Fuente:
arXiv
Saved in:
| Main Authors: | Hůla, Jan, Adamczyk, David, Filip, Tomáš, Pavlíček, Martin, Sosík, Petr |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-tuning multilingual language models in Twitter/X sentiment analysis: a study on Eastern-European V4 languages
by: Filip, Tomáš, et al.
Published: (2024)
by: Filip, Tomáš, et al.
Published: (2024)
BIPOLAR: Polarization-based granular framework for LLM bias evaluation
by: Pavlíček, Martin, et al.
Published: (2025)
by: Pavlíček, Martin, et al.
Published: (2025)
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
by: Fajcik, Martin, et al.
Published: (2024)
by: Fajcik, Martin, et al.
Published: (2024)
FlashEVA: Accelerating LLM inference via Efficient Attention
by: Kostelec, Juan Gabriel, et al.
Published: (2025)
by: Kostelec, Juan Gabriel, et al.
Published: (2025)
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
by: Somasundaram, Shwetha, et al.
Published: (2024)
by: Somasundaram, Shwetha, et al.
Published: (2024)
An Efficient Inference Framework for Early-exit Large Language Models
by: Miao, Ruijie, et al.
Published: (2024)
by: Miao, Ruijie, et al.
Published: (2024)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
by: Kucia, Filip J., et al.
Published: (2026)
by: Kucia, Filip J., et al.
Published: (2026)
Investigating the Robustness of Deductive Reasoning with Large Language Models
by: Hoppe, Fabian, et al.
Published: (2025)
by: Hoppe, Fabian, et al.
Published: (2025)
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
Controllable and explainable personality sliders for LLMs at inference time
by: Hoppe, Florian, et al.
Published: (2026)
by: Hoppe, Florian, et al.
Published: (2026)
Enhancing Structural Mapping with LLM-derived Abstractions for Analogical Reasoning in Narratives
by: Khojasteh, Mohammadhossein, et al.
Published: (2026)
by: Khojasteh, Mohammadhossein, et al.
Published: (2026)
The Knowledge-Behaviour Disconnect in LLM-based Chatbots
by: Broersen, Jan
Published: (2025)
by: Broersen, Jan
Published: (2025)
A comprehensive study of LLM-based argument classification: from Llama through DeepSeek to GPT-5.2
by: Pietroń, Marcin, et al.
Published: (2026)
by: Pietroń, Marcin, et al.
Published: (2026)
SemioLLM: Evaluating Large Language Models for Diagnostic Reasoning from Unstructured Clinical Narratives in Epilepsy
by: Dani, Meghal, et al.
Published: (2024)
by: Dani, Meghal, et al.
Published: (2024)
A comprehensive study of LLM-based argument classification: from LLAMA through GPT-4o to Deepseek-R1
by: Pietroń, Marcin, et al.
Published: (2025)
by: Pietroń, Marcin, et al.
Published: (2025)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
Benchmarking of LLM Detection: Comparing Two Competing Approaches
by: Pröhl, Thorsten, et al.
Published: (2024)
by: Pröhl, Thorsten, et al.
Published: (2024)
LLMzSzŁ: a comprehensive LLM benchmark for Polish
by: Jassem, Krzysztof, et al.
Published: (2025)
by: Jassem, Krzysztof, et al.
Published: (2025)
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning
by: Chadimová, Milena, et al.
Published: (2024)
by: Chadimová, Milena, et al.
Published: (2024)
PolicyBank: Evolving Policy Understanding for LLM Agents
by: Choi, Jihye, et al.
Published: (2026)
by: Choi, Jihye, et al.
Published: (2026)
Thinking Tokens for Language Modeling
by: Herel, David, et al.
Published: (2024)
by: Herel, David, et al.
Published: (2024)
Collapse of Self-trained Language Models
by: Herel, David, et al.
Published: (2024)
by: Herel, David, et al.
Published: (2024)
LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
by: Zolnour, Ali, et al.
Published: (2025)
by: Zolnour, Ali, et al.
Published: (2025)
PatentScore: Multi-dimensional Evaluation of LLM-Generated Patent Claims
by: Yoo, Yongmin, et al.
Published: (2025)
by: Yoo, Yongmin, et al.
Published: (2025)
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
by: Min, Hyangsuk, et al.
Published: (2025)
by: Min, Hyangsuk, et al.
Published: (2025)
HypeLoRA: Hyper-Network-Generated LoRA Adapters for Calibrated Language Model Fine-Tuning
by: Trojan, Bartosz, et al.
Published: (2026)
by: Trojan, Bartosz, et al.
Published: (2026)
Quantifying Generative Media Bias with a Corpus of Real-world and Generated News Articles
by: Trhlik, Filip, et al.
Published: (2024)
by: Trhlik, Filip, et al.
Published: (2024)
Towards Geo-Culturally Grounded LLM Generations
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
by: Lertvittayakumjorn, Piyawat, et al.
Published: (2025)
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
by: Sondej, Filip, et al.
Published: (2025)
by: Sondej, Filip, et al.
Published: (2025)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
LLM one-shot style transfer for Authorship Attribution and Verification
by: Miralles-González, Pablo, et al.
Published: (2025)
by: Miralles-González, Pablo, et al.
Published: (2025)
Efficient LLM Moderation with Multi-Layer Latent Prototypes
by: Chrabąszcz, Maciej, et al.
Published: (2025)
by: Chrabąszcz, Maciej, et al.
Published: (2025)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
by: Schmidt, Jan-Philipp
Published: (2026)
by: Schmidt, Jan-Philipp
Published: (2026)
Two-Stage Reasoning-Infused Learning: Improving Classification with LLM-Generated Reasoning
by: Henrichsen, Mads, et al.
Published: (2025)
by: Henrichsen, Mads, et al.
Published: (2025)
Sustainability via LLM Right-sizing
by: Haase, Jennifer, et al.
Published: (2025)
by: Haase, Jennifer, et al.
Published: (2025)
Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
by: Shin, Seungjun, et al.
Published: (2025)
by: Shin, Seungjun, et al.
Published: (2025)
Sound and Complete Neurosymbolic Reasoning with LLM-Grounded Interpretations
by: Allen, Bradley P., et al.
Published: (2025)
by: Allen, Bradley P., et al.
Published: (2025)
Can LLMs Evaluate What They Cannot Annotate? Revisiting LLM Reliability in Hate Speech Detection
by: Piot, Paloma, et al.
Published: (2025)
by: Piot, Paloma, et al.
Published: (2025)
Similar Items
-
Fine-tuning multilingual language models in Twitter/X sentiment analysis: a study on Eastern-European V4 languages
by: Filip, Tomáš, et al.
Published: (2024) -
BIPOLAR: Polarization-based granular framework for LLM bias evaluation
by: Pavlíček, Martin, et al.
Published: (2025) -
BenCzechMark : A Czech-centric Multitask and Multimetric Benchmark for Large Language Models with Duel Scoring Mechanism
by: Fajcik, Martin, et al.
Published: (2024) -
FlashEVA: Accelerating LLM inference via Efficient Attention
by: Kostelec, Juan Gabriel, et al.
Published: (2025) -
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
by: Somasundaram, Shwetha, et al.
Published: (2024)