Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Jerry, Parthasarathi, Prasanna, Rezagholizadeh, Mehdi, Chandar, Sarath |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
Steering Large Language Model Activations in Sparse Spaces
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
von: Bayat, Reza, et al.
Veröffentlicht: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
Intelligent Switching for Reset-Free RL
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
Lookbehind-SAM: k steps back, 1 step forward
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025)
von: Jamialahmadi, Benyamin, et al.
Veröffentlicht: (2025)
Torque-Aware Momentum
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2024)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
von: Malviya, Pranshu, et al.
Veröffentlicht: (2023)
von: Malviya, Pranshu, et al.
Veröffentlicht: (2023)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
von: Nekoei, Hadi, et al.
Veröffentlicht: (2025)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2024)
Improving Context-Aware Preference Modeling for Language Models
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Learning to Select In-Context Demonstration Preferred by Large Language Model
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zheng, et al.
Veröffentlicht: (2025)
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
von: Chen, Sihan, et al.
Veröffentlicht: (2025)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
von: Li, Tenghui, et al.
Veröffentlicht: (2025)
von: Li, Tenghui, et al.
Veröffentlicht: (2025)
How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities
von: Huang, Jerry
Veröffentlicht: (2024)
von: Huang, Jerry
Veröffentlicht: (2024)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
von: Balestriero, Randall, et al.
Veröffentlicht: (2023)
von: Balestriero, Randall, et al.
Veröffentlicht: (2023)
Improving Large Language Models with Concept-Aware Fine-Tuning
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
PickLLM: Context-Aware RL-Assisted Large Language Model Routing
von: Sikeridis, Dimitrios, et al.
Veröffentlicht: (2024)
von: Sikeridis, Dimitrios, et al.
Veröffentlicht: (2024)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
von: Haridas, Akash, et al.
Veröffentlicht: (2026)
von: Haridas, Akash, et al.
Veröffentlicht: (2026)
LLaGA: Large Language and Graph Assistant
von: Chen, Runjin, et al.
Veröffentlicht: (2024)
von: Chen, Runjin, et al.
Veröffentlicht: (2024)
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
Context-Aware Membership Inference Attacks against Pre-trained Large Language Models
von: Chang, Hongyan, et al.
Veröffentlicht: (2024)
von: Chang, Hongyan, et al.
Veröffentlicht: (2024)
Small Encoders Can Rival Large Decoders in Detecting Groundedness
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
Calibrated Language Models and How to Find Them with Label Smoothing
von: Huang, Jerry, et al.
Veröffentlicht: (2025)
von: Huang, Jerry, et al.
Veröffentlicht: (2025)
Improving ARDS Diagnosis Through Context-Aware Concept Bottleneck Models
von: Narain, Anish, et al.
Veröffentlicht: (2025)
von: Narain, Anish, et al.
Veröffentlicht: (2025)
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy
von: Zhao, Yao, et al.
Veröffentlicht: (2023)
von: Zhao, Yao, et al.
Veröffentlicht: (2023)
Inference Optimization of Foundation Models on AI Accelerators
von: Park, Youngsuk, et al.
Veröffentlicht: (2024)
von: Park, Youngsuk, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024) -
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024) -
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025) -
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025) -
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)