Efficient Reasoning on the Edge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bondarenko, Yelysei, Hehn, Thomas, Hesselink, Rob, Lepert, Romain, Massoli, Fabio Valerio, Mironov, Evgeny, Mirvakhabova, Leyla, Orekondy, Tribhuvanesh, Stasis, Spyridon, Kuzmin, Andrey, Kuzina, Anna, Nagel, Markus, Nayak, Ankita, Rainone, Corrado, de Rooij, Ork, Whatmough, Paul N, Behboodi, Arash, Bejnordi, Babak Ehteshami |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Simulating, Fast and Slow: Learning Policies for Black-Box Optimization
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2024)
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2024)
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
von: Rakhsha, Amin, et al.
Veröffentlicht: (2026)
von: Rakhsha, Amin, et al.
Veröffentlicht: (2026)
Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025)
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025)
Differentiable and Learnable Wireless Simulation with Geometric Transformers
von: Hehn, Thomas, et al.
Veröffentlicht: (2024)
von: Hehn, Thomas, et al.
Veröffentlicht: (2024)
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
von: Kuzina, Anna, et al.
Veröffentlicht: (2025)
von: Kuzina, Anna, et al.
Veröffentlicht: (2025)
Reinforcement Learning of Adaptive Acquisition Policies for Inverse Problems
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2024)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2024)
FPTQuant: Function-Preserving Transforms for LLM Quantization
von: van Breugel, Boris, et al.
Veröffentlicht: (2025)
von: van Breugel, Boris, et al.
Veröffentlicht: (2025)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
von: Skliar, Andrii, et al.
Veröffentlicht: (2024)
LaneRoPE: Positional Encoding for Collaborative Parallel Reasoning and Generation
von: Cesa, Gabriele, et al.
Veröffentlicht: (2026)
von: Cesa, Gabriele, et al.
Veröffentlicht: (2026)
Reasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2026)
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2026)
Low-Rank Quantization-Aware Training for LLMs
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
von: Bondarenko, Yelysei, et al.
Veröffentlicht: (2024)
When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems
von: Rainone, Corrado, et al.
Veröffentlicht: (2026)
von: Rainone, Corrado, et al.
Veröffentlicht: (2026)
Variational Learning ISTA
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2024)
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2024)
Replacing thinking with tool usage enables reasoning in small language models
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
Pruning vs Quantization: Which is Better?
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023)
von: Kuzmin, Andrey, et al.
Veröffentlicht: (2023)
Fundamental bounds on efficiency-confidence trade-off for transductive conformal prediction
von: Behboodi, Arash, et al.
Veröffentlicht: (2025)
von: Behboodi, Arash, et al.
Veröffentlicht: (2025)
An Information Theoretic Perspective on Conformal Prediction
von: Correia, Alvaro H. C., et al.
Veröffentlicht: (2024)
von: Correia, Alvaro H. C., et al.
Veröffentlicht: (2024)
Numerical study of plasma behavior in a disk‐shaped noble gas MHD generator
von: Ork Kimsor, et al.
Veröffentlicht: (2024)
von: Ork Kimsor, et al.
Veröffentlicht: (2024)
Vision-Assisted Digital Twin Creation for mmWave Beam Management
von: Arnold, Maximilian, et al.
Veröffentlicht: (2024)
von: Arnold, Maximilian, et al.
Veröffentlicht: (2024)
Neural Mesh Fusion: Unsupervised 3D Planar Surface Understanding
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2024)
von: Zanjani, Farhad G., et al.
Veröffentlicht: (2024)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
InterroGate: Learning to Share, Specialize, and Prune Representations for Multi-task Learning
von: Bejnordi, Babak Ehteshami, et al.
Veröffentlicht: (2024)
von: Bejnordi, Babak Ehteshami, et al.
Veröffentlicht: (2024)
Human-in-the-loop Reasoning For Traffic Sign Detection: Collaborative Approach Yolo With Video-llava
von: Azarafza, Mehdi, et al.
Veröffentlicht: (2024)
von: Azarafza, Mehdi, et al.
Veröffentlicht: (2024)
Dissecting Quantization Error: A Concentration-Alignment Perspective
von: Federici, Marco, et al.
Veröffentlicht: (2026)
von: Federici, Marco, et al.
Veröffentlicht: (2026)
Modelling Territorial Dynamics through Statistical Mechanics: An Application to Resident Foreign Population
von: Massoli, Pierpaolo
Veröffentlicht: (2025)
von: Massoli, Pierpaolo
Veröffentlicht: (2025)
Unveiling Complex Territorial Socio-Economic Dynamics: A Statistical Mechanics Approach
von: Massoli, Pierpaolo
Veröffentlicht: (2025)
von: Massoli, Pierpaolo
Veröffentlicht: (2025)
Learning Optical Flow Field via Neural Ordinary Differential Equation
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025)
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025)
Healers on the colonial market; Native doctors and midwives in the Dutch East Indies
von: Hesselink, Liesbeth
Veröffentlicht: (2011)
von: Hesselink, Liesbeth
Veröffentlicht: (2011)
Development of a College Success Management Course for York County Technical College.
von: Rainone, John J.
Veröffentlicht: (1997)
von: Rainone, John J.
Veröffentlicht: (1997)
GPTVQ: The Blessing of Dimensionality for LLM Quantization
von: van Baalen, Mart, et al.
Veröffentlicht: (2024)
von: van Baalen, Mart, et al.
Veröffentlicht: (2024)
AS MUDANÇAS LEGAIS NO AMBIENTE INSTITUCIONAL DO SETOR DE EDUCAÇÃO E AS ESTRATÉGIAS DE CRESCIMENTO DE UMA INSTITUIÇÃO DE ENSINO SUPERIOR
von: Renata Massoli Borges
Veröffentlicht: (2013)
von: Renata Massoli Borges
Veröffentlicht: (2013)
AN ANTI-CHRISTIAN REGISTER FROM NAGASAKI
von: Reinier H. Hesselink
Veröffentlicht: (2009)
von: Reinier H. Hesselink
Veröffentlicht: (2009)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
von: Cai, Ruisi, et al.
Veröffentlicht: (2024)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
Hierarchical VAE with a Diffusion-based VampPrior
von: Kuzina, Anna, et al.
Veröffentlicht: (2024)
von: Kuzina, Anna, et al.
Veröffentlicht: (2024)
STaMP: Sequence Transformation and Mixed Precision for Low-Precision Activation Quantization
von: Federici, Marco, et al.
Veröffentlicht: (2025)
von: Federici, Marco, et al.
Veröffentlicht: (2025)
HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations
von: Federici, Marco, et al.
Veröffentlicht: (2025)
von: Federici, Marco, et al.
Veröffentlicht: (2025)
Preliminary survey on the relationship between dimensions of pinctada radiata in Moqam
von: Ehteshami, Fariborz.
Veröffentlicht: (1994)
von: Ehteshami, Fariborz.
Veröffentlicht: (1994)
Anesthetizing Pinctada radiata with MS-222
von: Ehteshami, Fariborz
Veröffentlicht: (1993)
von: Ehteshami, Fariborz
Veröffentlicht: (1993)
Trendbericht Theoretische Chemie 2025 2/2: Theoretische Spektroskopie
von: Anna Hehn
Veröffentlicht: (2025)
von: Anna Hehn
Veröffentlicht: (2025)
Ähnliche Einträge
-
Simulating, Fast and Slow: Learning Policies for Black-Box Optimization
von: Massoli, Fabio Valerio, et al.
Veröffentlicht: (2024) -
LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents
von: Rakhsha, Amin, et al.
Veröffentlicht: (2026) -
Dirichlet-Prior Shaping: Guiding Expert Specialization in Upcycled MoEs
von: Mirvakhabova, Leyla, et al.
Veröffentlicht: (2025) -
Differentiable and Learnable Wireless Simulation with Geometric Transformers
von: Hehn, Thomas, et al.
Veröffentlicht: (2024) -
KaVa: Latent Reasoning via Compressed KV-Cache Distillation
von: Kuzina, Anna, et al.
Veröffentlicht: (2025)