Tandem Transformers for Inference Efficient LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | S, Aishwarya P, Nair, Pranav Ajit, Samaga, Yashas, Boyd, Toby, Kumar, Sanjiv, Jain, Prateek, Netrapalli, Praneeth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024)
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024)
A Faster Generalized Two-Stage Approximate Top-K
von: Samaga, Yashas, et al.
Veröffentlicht: (2025)
von: Samaga, Yashas, et al.
Veröffentlicht: (2025)
A model of errors in transformers
von: Raju, Suvrat, et al.
Veröffentlicht: (2026)
von: Raju, Suvrat, et al.
Veröffentlicht: (2026)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024)
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024)
Regression-aware Inference with LLMs
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
von: Simonds, Toby
Veröffentlicht: (2025)
von: Simonds, Toby
Veröffentlicht: (2025)
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
von: Addepalli, Sravanti, et al.
Veröffentlicht: (2024)
von: Addepalli, Sravanti, et al.
Veröffentlicht: (2024)
Time-Reversal Provides Unsupervised Feedback to LLMs
von: Varun, Yerram, et al.
Veröffentlicht: (2024)
von: Varun, Yerram, et al.
Veröffentlicht: (2024)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
von: Goru, Ritesh, et al.
Veröffentlicht: (2025)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
von: Ghadia, Ravi, et al.
Veröffentlicht: (2025)
von: Ghadia, Ravi, et al.
Veröffentlicht: (2025)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2025)
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2025)
Path-Consistency with Prefix Enhancement for Efficient Inference in LLMs
von: Zhu, Jiace, et al.
Veröffentlicht: (2024)
von: Zhu, Jiace, et al.
Veröffentlicht: (2024)
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
von: Vadlapati, Praneeth
Veröffentlicht: (2024)
von: Vadlapati, Praneeth
Veröffentlicht: (2024)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
von: Thareja, Rushil, et al.
Veröffentlicht: (2025)
von: Thareja, Rushil, et al.
Veröffentlicht: (2025)
Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models
von: Chhikara, Prateek
Veröffentlicht: (2025)
von: Chhikara, Prateek
Veröffentlicht: (2025)
VISTA: Visualization of Token Attribution via Efficient Analysis
von: Ahmed, Syed, et al.
Veröffentlicht: (2026)
von: Ahmed, Syed, et al.
Veröffentlicht: (2026)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2024)
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2024)
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2025)
von: Banayeeanzade, Amin, et al.
Veröffentlicht: (2025)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
von: Singhal, Raghav, et al.
Veröffentlicht: (2025)
von: Singhal, Raghav, et al.
Veröffentlicht: (2025)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
von: Machcha, Sravanthi, et al.
Veröffentlicht: (2026)
von: Machcha, Sravanthi, et al.
Veröffentlicht: (2026)
Deep sequence models tend to memorize geometrically; it is unclear why
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
GradPruner: Gradient-Guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2026)
von: Huang, Wei, et al.
Veröffentlicht: (2026)
Common Sense vs. Morality: The Curious Case of Narrative Focus Bias in LLMs
von: Purkayastha, Saugata, et al.
Veröffentlicht: (2026)
von: Purkayastha, Saugata, et al.
Veröffentlicht: (2026)
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction
von: Vadlapati, Praneeth
Veröffentlicht: (2024)
von: Vadlapati, Praneeth
Veröffentlicht: (2024)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
von: Ji, Tao, et al.
Veröffentlicht: (2025)
von: Ji, Tao, et al.
Veröffentlicht: (2025)
LLMs Are Prone to Fallacies in Causal Inference
von: Joshi, Nitish, et al.
Veröffentlicht: (2024)
von: Joshi, Nitish, et al.
Veröffentlicht: (2024)
Humans and LLMs Diverge on Probabilistic Inferences
von: Kamath, Gaurav, et al.
Veröffentlicht: (2026)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2026)
Draft-based Approximate Inference for LLMs
von: Galim, Kevin, et al.
Veröffentlicht: (2025)
von: Galim, Kevin, et al.
Veröffentlicht: (2025)
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
von: Anand, Avinash, et al.
Veröffentlicht: (2024)
ShishuLM : Achieving Optimal and Efficient Parameterization with Low Attention Transformer Models
von: Kumar, Shivanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Shivanshu, et al.
Veröffentlicht: (2025)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
von: Drinkall, Toby
Veröffentlicht: (2025)
von: Drinkall, Toby
Veröffentlicht: (2025)
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2023)
von: Wijesiriwardene, Thilini, et al.
Veröffentlicht: (2023)
Fairness Evaluation and Inference Level Mitigation in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Can Large Language Models Detect Misinformation in Scientific News Reporting?
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
von: Cao, Yupeng, et al.
Veröffentlicht: (2024)
Succeeding at Scale: Automated Dataset Construction and Query-Side Adaptation for Multi-Tenant Search
von: Jain, Prateek, et al.
Veröffentlicht: (2026)
von: Jain, Prateek, et al.
Veröffentlicht: (2026)
When Benchmarks Leak: Inference-Time Decontamination for LLMs
von: Chai, Jianzhe, et al.
Veröffentlicht: (2026)
von: Chai, Jianzhe, et al.
Veröffentlicht: (2026)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
von: Chehade, Mohamad, et al.
Veröffentlicht: (2025)
von: Chehade, Mohamad, et al.
Veröffentlicht: (2025)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
von: Cheng, Myra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024) -
A Faster Generalized Two-Stage Approximate Top-K
von: Samaga, Yashas, et al.
Veröffentlicht: (2025) -
A model of errors in transformers
von: Raju, Suvrat, et al.
Veröffentlicht: (2026) -
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
von: Nair, Pranav Ajit, et al.
Veröffentlicht: (2024) -
Regression-aware Inference with LLMs
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)