SNLP: Layer-Parallel Inference via Structured Newton Corrections
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Ligong, Xu, Kai, Wang, Hao, Srivastava, Akash |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement
von: Dang, Meihua, et al.
Veröffentlicht: (2025)
von: Dang, Meihua, et al.
Veröffentlicht: (2025)
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods
von: Puri, Isha, et al.
Veröffentlicht: (2025)
von: Puri, Isha, et al.
Veröffentlicht: (2025)
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
von: Zhou, Enyu, et al.
Veröffentlicht: (2025)
Few-Step Diffusion Language Models via Trajectory Self-Distillation
von: Zhang, Tunyu, et al.
Veröffentlicht: (2026)
von: Zhang, Tunyu, et al.
Veröffentlicht: (2026)
Hopscotch: Discovering and Skipping Redundancies in Language Models
von: Eyceoz, Mustafa, et al.
Veröffentlicht: (2025)
von: Eyceoz, Mustafa, et al.
Veröffentlicht: (2025)
S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation
von: Han, Ligong, et al.
Veröffentlicht: (2026)
von: Han, Ligong, et al.
Veröffentlicht: (2026)
Spectrum-Aware Parameter Efficient Fine-Tuning for Diffusion Models
von: Zhang, Xinxi, et al.
Veröffentlicht: (2024)
von: Zhang, Xinxi, et al.
Veröffentlicht: (2024)
Mitigating Premature Exploitation in Particle-based Monte Carlo for Inference-Time Scaling
von: Giannone, Giorgio, et al.
Veröffentlicht: (2025)
von: Giannone, Giorgio, et al.
Veröffentlicht: (2025)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
von: Zhang, Qiuyang, et al.
Veröffentlicht: (2026)
von: Zhang, Qiuyang, et al.
Veröffentlicht: (2026)
Exploiting Student Parallelism for Efficient GPU Inference of BERT-like Models in Online Services
von: Wang, Weiyan, et al.
Veröffentlicht: (2024)
von: Wang, Weiyan, et al.
Veröffentlicht: (2024)
Enhancing Goal Inference via Correction Timing
von: Wang, Anjiabei, et al.
Veröffentlicht: (2026)
von: Wang, Anjiabei, et al.
Veröffentlicht: (2026)
LInK: Learning Joint Representations of Design and Performance Spaces through Contrastive Learning for Mechanism Synthesis
von: Nobari, Amin Heyrani, et al.
Veröffentlicht: (2024)
von: Nobari, Amin Heyrani, et al.
Veröffentlicht: (2024)
Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning
von: Nayak, Nikhil Shivakumar, et al.
Veröffentlicht: (2025)
von: Nayak, Nikhil Shivakumar, et al.
Veröffentlicht: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
von: Shi, Haizhou, et al.
Veröffentlicht: (2024)
von: Shi, Haizhou, et al.
Veröffentlicht: (2024)
BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language Models
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
Constrained Gaussian Process Motion Planning via Stein Variational Newton Inference
von: Li, Jiayun, et al.
Veröffentlicht: (2025)
von: Li, Jiayun, et al.
Veröffentlicht: (2025)
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
von: Pareja, Aldo, et al.
Veröffentlicht: (2024)
von: Pareja, Aldo, et al.
Veröffentlicht: (2024)
Differentially Private Synthetic Data Generation for Relational Databases
von: Alimohammadi, Kaveh, et al.
Veröffentlicht: (2024)
von: Alimohammadi, Kaveh, et al.
Veröffentlicht: (2024)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
Parallel Layer Normalization for Universal Approximation
von: Ni, Yunhao, et al.
Veröffentlicht: (2025)
von: Ni, Yunhao, et al.
Veröffentlicht: (2025)
LAB: Large-Scale Alignment for ChatBots
von: Sudalairaj, Shivchander, et al.
Veröffentlicht: (2024)
von: Sudalairaj, Shivchander, et al.
Veröffentlicht: (2024)
Hinge Regression Tree: A Newton Method for Oblique Regression Tree Splitting
von: Li, Hongyi, et al.
Veröffentlicht: (2026)
von: Li, Hongyi, et al.
Veröffentlicht: (2026)
QoS-Nets: Adaptive Approximate Neural Network Inference
von: Trommer, Elias, et al.
Veröffentlicht: (2024)
von: Trommer, Elias, et al.
Veröffentlicht: (2024)
GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric Feedback
von: Giannone, Giorgio, et al.
Veröffentlicht: (2026)
von: Giannone, Giorgio, et al.
Veröffentlicht: (2026)
A Probabilistic Framework for Modular Continual Learning
von: Valkov, Lazar, et al.
Veröffentlicht: (2023)
von: Valkov, Lazar, et al.
Veröffentlicht: (2023)
Beyond Statistical Similarity: Rethinking Metrics for Deep Generative Models in Engineering Design
von: Regenwetter, Lyle, et al.
Veröffentlicht: (2023)
von: Regenwetter, Lyle, et al.
Veröffentlicht: (2023)
Parallel Sequence Modeling via Generalized Spatial Propagation Network
von: Wang, Hongjun, et al.
Veröffentlicht: (2025)
von: Wang, Hongjun, et al.
Veröffentlicht: (2025)
MMD-Newton Method for Multi-objective Optimization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
Massively Parallel Exact Inference for Hawkes Processes
von: Raza, Ahmer, et al.
Veröffentlicht: (2026)
von: Raza, Ahmer, et al.
Veröffentlicht: (2026)
A Majorization-Minimization Gauss-Newton Method for 1-Bit Matrix Completion
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2023)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
Value Augmented Sampling for Language Model Alignment and Personalization
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
Parallelization of the K-Means Algorithm with Applications to Big Data Clustering
von: Srivastava, Ashish, et al.
Veröffentlicht: (2024)
von: Srivastava, Ashish, et al.
Veröffentlicht: (2024)
Not All Layers of LLMs Are Necessary During Inference
von: Fan, Siqi, et al.
Veröffentlicht: (2024)
von: Fan, Siqi, et al.
Veröffentlicht: (2024)
Urban context and delivery performance: Modelling service time for cargo bikes and vans across diverse urban environments
von: Schrader, Maxwell, et al.
Veröffentlicht: (2024)
von: Schrader, Maxwell, et al.
Veröffentlicht: (2024)
Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism
von: Wang, Kunyun, et al.
Veröffentlicht: (2025)
von: Wang, Kunyun, et al.
Veröffentlicht: (2025)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
von: Liu, Dong, et al.
Veröffentlicht: (2024)
von: Liu, Dong, et al.
Veröffentlicht: (2024)
Inference-Time Attribute Distribution Alignment for Unconditional Diffusion
von: Luan, Hao, et al.
Veröffentlicht: (2026)
von: Luan, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025) -
Inference-Time Scaling of Diffusion Language Models via Trajectory Refinement
von: Dang, Meihua, et al.
Veröffentlicht: (2025) -
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods
von: Puri, Isha, et al.
Veröffentlicht: (2025) -
CARD: A Cache-Assisted Parallel Speculative Decoding Framework via Query-and-Correct Paradigm for Accelerating LLM Inference
von: Zhou, Enyu, et al.
Veröffentlicht: (2025) -
Few-Step Diffusion Language Models via Trajectory Self-Distillation
von: Zhang, Tunyu, et al.
Veröffentlicht: (2026)