Recursive Inference Scaling: A Winning Path to Scalable Inference in Language and Multimodal Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Alabdulmohsin, Ibrahim, Zhai, Xiaohua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2024)
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2024)
Recursive Inference Machines for Neural Reasoning
von: Komisarczyk, Mieszko, et al.
Veröffentlicht: (2026)
von: Komisarczyk, Mieszko, et al.
Veröffentlicht: (2026)
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
von: Neth, Ashe, et al.
Veröffentlicht: (2025)
von: Neth, Ashe, et al.
Veröffentlicht: (2025)
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2023)
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2023)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
Scaling Inference-Efficient Language Models
von: Bian, Song, et al.
Veröffentlicht: (2025)
von: Bian, Song, et al.
Veröffentlicht: (2025)
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
von: Brown, Bradley, et al.
Veröffentlicht: (2024)
von: Brown, Bradley, et al.
Veröffentlicht: (2024)
Generative Score Inference for Multimodal Data
von: Tian, Xinyu, et al.
Veröffentlicht: (2026)
von: Tian, Xinyu, et al.
Veröffentlicht: (2026)
Learning to Condition: A Neural Heuristic for Scalable MPE Inference
von: Malhotra, Brij, et al.
Veröffentlicht: (2025)
von: Malhotra, Brij, et al.
Veröffentlicht: (2025)
When Dynamics Shift, Robust Task Inference Wins: Offline Imitation Learning with Behavior Foundation Models Revisited
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2026)
von: Agrawal, Rishabh, et al.
Veröffentlicht: (2026)
Consistency Models for Scalable and Fast Simulation-Based Inference
von: Schmitt, Marvin, et al.
Veröffentlicht: (2023)
von: Schmitt, Marvin, et al.
Veröffentlicht: (2023)
Context Parallelism for Scalable Million-Token Inference
von: Yang, Amy, et al.
Veröffentlicht: (2024)
von: Yang, Amy, et al.
Veröffentlicht: (2024)
Geometric Scaling of Bayesian Inference in LLMs
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
The Limits of Inference Scaling Through Resampling
von: Stroebl, Benedikt, et al.
Veröffentlicht: (2024)
von: Stroebl, Benedikt, et al.
Veröffentlicht: (2024)
Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
von: Vaidhya, Tejas, et al.
Veröffentlicht: (2025)
von: Vaidhya, Tejas, et al.
Veröffentlicht: (2025)
Industrializing Prediction-Powered Inference: The GLIDE Library for Reliable GenAI and Agentic Systems Evaluation
von: Martinon, Grégoire, et al.
Veröffentlicht: (2026)
von: Martinon, Grégoire, et al.
Veröffentlicht: (2026)
Winning the MIDST Challenge: New Membership Inference Attacks on Diffusion Models for Tabular Data Synthesis
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wu, Xiaoyu, et al.
Veröffentlicht: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Scalable Simulation-Based Model Inference with Test-Time Complexity Control
von: Gloeckler, Manuel, et al.
Veröffentlicht: (2026)
von: Gloeckler, Manuel, et al.
Veröffentlicht: (2026)
Scalable AI Inference: Performance Analysis and Optimization of AI Model Serving
von: Pham, Hung Cuong, et al.
Veröffentlicht: (2026)
von: Pham, Hung Cuong, et al.
Veröffentlicht: (2026)
BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
von: Wu, Xiaoyou, et al.
Veröffentlicht: (2026)
von: Wu, Xiaoyou, et al.
Veröffentlicht: (2026)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
Scaling On-Device GPU Inference for Large Generative Models
von: Tang, Jiuqiang, et al.
Veröffentlicht: (2025)
von: Tang, Jiuqiang, et al.
Veröffentlicht: (2025)
Diversified Scaling Inference in Time Series Foundation Models
von: Hua, Ruijin, et al.
Veröffentlicht: (2026)
von: Hua, Ruijin, et al.
Veröffentlicht: (2026)
Fuse It or Lose It: Deep Fusion for Multimodal Simulation-Based Inference
von: Schmitt, Marvin, et al.
Veröffentlicht: (2023)
von: Schmitt, Marvin, et al.
Veröffentlicht: (2023)
A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2025)
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2025)
Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo Methods
von: Puri, Isha, et al.
Veröffentlicht: (2025)
von: Puri, Isha, et al.
Veröffentlicht: (2025)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
von: Shrestha, Susav, et al.
Veröffentlicht: (2025)
von: Shrestha, Susav, et al.
Veröffentlicht: (2025)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
Adaptive Inference-Time Scaling via Cyclic Diffusion Search
von: Lee, Gyubin, et al.
Veröffentlicht: (2025)
von: Lee, Gyubin, et al.
Veröffentlicht: (2025)
It Just Takes Two: Scaling Amortized Inference to Large Sets
von: Wehenkel, Antoine, et al.
Veröffentlicht: (2026)
von: Wehenkel, Antoine, et al.
Veröffentlicht: (2026)
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
von: Riemer, Matthew, et al.
Veröffentlicht: (2024)
von: Riemer, Matthew, et al.
Veröffentlicht: (2024)
Fast Inference for Augmented Large Language Models
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
von: Shahout, Rana, et al.
Veröffentlicht: (2024)
Out of Context: Reliability in Multimodal Anomaly Detection Requires Contextual Inference
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
A Simple Model of Inference Scaling Laws
von: Levi, Noam
Veröffentlicht: (2024)
von: Levi, Noam
Veröffentlicht: (2024)
Learning to Inference Adaptively for Multimodal Large Language Models
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2025)
von: Xu, Zhuoyan, et al.
Veröffentlicht: (2025)
Mini-Giants: "Small" Language Models and Open Source Win-Win
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
von: Zhou, Zhengping, et al.
Veröffentlicht: (2023)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
von: Yue, Yang, et al.
Veröffentlicht: (2024)
von: Yue, Yang, et al.
Veröffentlicht: (2024)
Find A Winning Sign: Sign Is All We Need to Win the Lottery
von: Oh, Junghun, et al.
Veröffentlicht: (2025)
von: Oh, Junghun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning?
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2024) -
Recursive Inference Machines for Neural Reasoning
von: Komisarczyk, Mieszko, et al.
Veröffentlicht: (2026) -
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
von: Neth, Ashe, et al.
Veröffentlicht: (2025) -
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
von: Alabdulmohsin, Ibrahim, et al.
Veröffentlicht: (2023) -
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
von: Li, Bo, et al.
Veröffentlicht: (2026)