Online Cascade Learning for Efficient Inference over Streams
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nie, Lunyiu, Ding, Zhimin, Hu, Erdong, Jermaine, Christopher, Chaudhuri, Swarat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Resource-efficient Inference with Foundation Model Programs
von: Nie, Lunyiu, et al.
Veröffentlicht: (2025)
von: Nie, Lunyiu, et al.
Veröffentlicht: (2025)
Batched Low-Rank Adaptation of Foundation Models
von: Wen, Yeming, et al.
Veröffentlicht: (2023)
von: Wen, Yeming, et al.
Veröffentlicht: (2023)
Learning Quantitative Automata Modulo Theories
von: Hsiung, Eric, et al.
Veröffentlicht: (2024)
von: Hsiung, Eric, et al.
Veröffentlicht: (2024)
When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding
von: Yang, Xu, et al.
Veröffentlicht: (2026)
von: Yang, Xu, et al.
Veröffentlicht: (2026)
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)
Efficient Tree-Structured Deep Research with Adaptive Resource Allocation
von: Nie, Lunyiu, et al.
Veröffentlicht: (2025)
von: Nie, Lunyiu, et al.
Veröffentlicht: (2025)
An In-Context Learning Agent for Formal Theorem-Proving
von: Thakur, Amitayush, et al.
Veröffentlicht: (2023)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2023)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
von: Tsoukalas, George, et al.
Veröffentlicht: (2024)
von: Tsoukalas, George, et al.
Veröffentlicht: (2024)
Are More Tokens Rational? Inference-Time Scaling in Language Models as Adaptive Resource Rationality
von: Hu, Zhimin, et al.
Veröffentlicht: (2026)
von: Hu, Zhimin, et al.
Veröffentlicht: (2026)
Why Code, Why Now: An Information-Theoretic Perspective on the Limits of Machine Learning
von: Zhao, Zhimin
Veröffentlicht: (2026)
von: Zhao, Zhimin
Veröffentlicht: (2026)
Cascade Speculative Drafting for Even Faster LLM Inference
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
von: Liu, Yongkang, et al.
Veröffentlicht: (2026)
von: Liu, Yongkang, et al.
Veröffentlicht: (2026)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
von: Hong, Yinrong, et al.
Veröffentlicht: (2025)
von: Hong, Yinrong, et al.
Veröffentlicht: (2025)
RAG-Modulo: Solving Sequential Tasks using Experience, Critics, and Language Models
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)
Learning-Time Encoding Shapes Unlearning in LLMs
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
von: Wu, Ruihan, et al.
Veröffentlicht: (2025)
Automata Learning from Preference and Equivalence Queries
von: Hsiung, Eric, et al.
Veröffentlicht: (2023)
von: Hsiung, Eric, et al.
Veröffentlicht: (2023)
Cascade Reward Sampling for Efficient Decoding-Time Alignment
von: Li, Bolian, et al.
Veröffentlicht: (2024)
von: Li, Bolian, et al.
Veröffentlicht: (2024)
Synthesize, Partition, then Adapt: Eliciting Diverse Samples from Foundation Models
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
GRASP: A Rehearsal Policy for Efficient Online Continual Learning
von: Harun, Md Yousuf, et al.
Veröffentlicht: (2023)
von: Harun, Md Yousuf, et al.
Veröffentlicht: (2023)
Star Attention: Efficient LLM Inference over Long Sequences
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
von: Acharya, Shantanu, et al.
Veröffentlicht: (2024)
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2024)
Grounding Data Science Code Generation with Input-Output Specifications
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
von: Deng, Hexuan, et al.
Veröffentlicht: (2025)
von: Deng, Hexuan, et al.
Veröffentlicht: (2025)
Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024)
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
OPTune: Efficient Online Preference Tuning
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
Efficient Learned Data Compression via Dual-Stream Feature Decoupling
von: Ma, Huidong, et al.
Veröffentlicht: (2026)
von: Ma, Huidong, et al.
Veröffentlicht: (2026)
PHONOS: PHOnetic Neutralization for Online Streaming Applications
von: Quamer, Waris, et al.
Veröffentlicht: (2026)
von: Quamer, Waris, et al.
Veröffentlicht: (2026)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
zip2zip: Inference-Time Adaptive Tokenization via Online Compression
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
von: Geng, Saibo, et al.
Veröffentlicht: (2025)
Universal Model Routing for Efficient LLM Inference
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025)
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
A Probabilistic Framework for Modular Continual Learning
von: Valkov, Lazar, et al.
Veröffentlicht: (2023)
von: Valkov, Lazar, et al.
Veröffentlicht: (2023)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
von: Agarwal, Saurabh, et al.
Veröffentlicht: (2024)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2024)
Symbolic Regression with a Learned Concept Library
von: Grayeli, Arya, et al.
Veröffentlicht: (2024)
von: Grayeli, Arya, et al.
Veröffentlicht: (2024)
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
von: Zhao, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Resource-efficient Inference with Foundation Model Programs
von: Nie, Lunyiu, et al.
Veröffentlicht: (2025) -
Batched Low-Rank Adaptation of Foundation Models
von: Wen, Yeming, et al.
Veröffentlicht: (2023) -
Learning Quantitative Automata Modulo Theories
von: Hsiung, Eric, et al.
Veröffentlicht: (2024) -
When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding
von: Yang, Xu, et al.
Veröffentlicht: (2026) -
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
von: Jain, Abhinav, et al.
Veröffentlicht: (2024)