Resource-Efficient Language Models: Quantization for Fast and Accessible Inference
Fuente:
arXiv
Guardado en:
| Autor principal: | Jørgensen, Tollef Emil |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
por: Spyra, Przemysław, et al.
Publicado: (2025)
por: Spyra, Przemysław, et al.
Publicado: (2025)
Large Language Models Report Subjective Experience Under Self-Referential Processing
por: Berg, Cameron, et al.
Publicado: (2025)
por: Berg, Cameron, et al.
Publicado: (2025)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
por: Shafieinejad, Masoumeh, et al.
Publicado: (2026)
por: Shafieinejad, Masoumeh, et al.
Publicado: (2026)
Interpretability Can Be Actionable
por: Orgad, Hadas, et al.
Publicado: (2026)
por: Orgad, Hadas, et al.
Publicado: (2026)
Improving Time Series Classification with Representation Soft Label Smoothing
por: Ma, Hengyi, et al.
Publicado: (2024)
por: Ma, Hengyi, et al.
Publicado: (2024)
Graph Neural Networks Need Cluster-Normalize-Activate Modules
por: Skryagin, Arseny, et al.
Publicado: (2024)
por: Skryagin, Arseny, et al.
Publicado: (2024)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
por: Badshah, Sher, et al.
Publicado: (2024)
por: Badshah, Sher, et al.
Publicado: (2024)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
por: Ged, François, et al.
Publicado: (2023)
por: Ged, François, et al.
Publicado: (2023)
The First MPDD Challenge: Multimodal Personality-aware Depression Detection
por: Fu, Changzeng, et al.
Publicado: (2025)
por: Fu, Changzeng, et al.
Publicado: (2025)
Deceptive Diffusion: Generating Synthetic Adversarial Examples
por: Beerens, Lucas, et al.
Publicado: (2024)
por: Beerens, Lucas, et al.
Publicado: (2024)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
por: Sarkar, Nilesh, et al.
Publicado: (2026)
por: Sarkar, Nilesh, et al.
Publicado: (2026)
ElegansNet: a brief scientific report and initial experiments
por: Bardozzo, Francesco, et al.
Publicado: (2023)
por: Bardozzo, Francesco, et al.
Publicado: (2023)
Surrealistic-like Image Generation with Vision-Language Models
por: Ayten, Elif, et al.
Publicado: (2024)
por: Ayten, Elif, et al.
Publicado: (2024)
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
por: Lin, Zezheng, et al.
Publicado: (2026)
por: Lin, Zezheng, et al.
Publicado: (2026)
Safe Continual Reinforcement Learning Methods for Nonstationary Environments. Towards a Survey of the State of the Art
por: Tomashevskiy, Timofey
Publicado: (2026)
por: Tomashevskiy, Timofey
Publicado: (2026)
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
por: Nayak, Nikhil, et al.
Publicado: (2026)
por: Nayak, Nikhil, et al.
Publicado: (2026)
UR4NNV: Neural Network Verification, Under-approximation Reachability Works!
por: Liang, Zhen, et al.
Publicado: (2024)
por: Liang, Zhen, et al.
Publicado: (2024)
Union of Experts: Adapting Hierarchical Routing to Equivalently Decomposed Transformer
por: Yang, Yujiao, et al.
Publicado: (2025)
por: Yang, Yujiao, et al.
Publicado: (2025)
The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning
por: Rajput, Vishal
Publicado: (2026)
por: Rajput, Vishal
Publicado: (2026)
From Language Models to Practical Self-Improving Computer Agents
por: Sheng, Alex
Publicado: (2024)
por: Sheng, Alex
Publicado: (2024)
The Curious Case of In-Training Compression of State Space Models
por: Chahine, Makram, et al.
Publicado: (2025)
por: Chahine, Makram, et al.
Publicado: (2025)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
por: Kajitsuka, Tokio, et al.
Publicado: (2023)
por: Kajitsuka, Tokio, et al.
Publicado: (2023)
Stealth edits to large language models
por: Sutton, Oliver J., et al.
Publicado: (2024)
por: Sutton, Oliver J., et al.
Publicado: (2024)
Structural Plasticity as Active Inference: A Biologically-Inspired Architecture for Homeostatic Control
por: Hill, Brennen A.
Publicado: (2025)
por: Hill, Brennen A.
Publicado: (2025)
Progressive Feedforward Collapse of ResNet Training
por: Wang, Sicong, et al.
Publicado: (2024)
por: Wang, Sicong, et al.
Publicado: (2024)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
por: Huang, Yingbing, et al.
Publicado: (2025)
por: Huang, Yingbing, et al.
Publicado: (2025)
Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior
por: Kim, Dongmin, et al.
Publicado: (2025)
por: Kim, Dongmin, et al.
Publicado: (2025)
Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration
por: Kabgere, Chethana Prasad
Publicado: (2025)
por: Kabgere, Chethana Prasad
Publicado: (2025)
Robust DDoS-Attack Classification with 3D CNNs Against Adversarial Methods
por: Bragg, Landon, et al.
Publicado: (2025)
por: Bragg, Landon, et al.
Publicado: (2025)
Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
por: Wang, Yongjie, et al.
Publicado: (2025)
por: Wang, Yongjie, et al.
Publicado: (2025)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
por: Platzer, André
Publicado: (2024)
por: Platzer, André
Publicado: (2024)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
por: Asperti, Andrea, et al.
Publicado: (2025)
por: Asperti, Andrea, et al.
Publicado: (2025)
A ZeNN architecture to avoid the Gaussian trap
por: Carvalho, Luís, et al.
Publicado: (2025)
por: Carvalho, Luís, et al.
Publicado: (2025)
Scalable Heterogeneous Graph Foundation Models for Data-Driven Optimal Power Flow in Smart Grids
por: Pasini, Massimiliano Lupo, et al.
Publicado: (2026)
por: Pasini, Massimiliano Lupo, et al.
Publicado: (2026)
Customizing Graph Neural Networks using Path Reweighting
por: Chen, Jianpeng, et al.
Publicado: (2021)
por: Chen, Jianpeng, et al.
Publicado: (2021)
Feature Relevancy, Necessity and Usefulness: Complexity and Algorithms
por: Capdevielle, Tomás, et al.
Publicado: (2025)
por: Capdevielle, Tomás, et al.
Publicado: (2025)
A Taxonomy of Omnicidal Futures Involving Artificial Intelligence
por: Critch, Andrew, et al.
Publicado: (2025)
por: Critch, Andrew, et al.
Publicado: (2025)
A Theoretical Framework for Adaptive Utility-Weighted Benchmarking
por: Waggoner, Philip
Publicado: (2026)
por: Waggoner, Philip
Publicado: (2026)
Adaptive Latent-Space Constraints in Personalized Federated Learning
por: Ayromlou, Sana, et al.
Publicado: (2025)
por: Ayromlou, Sana, et al.
Publicado: (2025)
Ejemplares similares
-
Energy-Efficient Deep Learning Without Backpropagation: A Rigorous Evaluation of Forward-Only Algorithms
por: Spyra, Przemysław, et al.
Publicado: (2025) -
Large Language Models Report Subjective Experience Under Self-Referential Processing
por: Berg, Cameron, et al.
Publicado: (2025) -
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
por: Shafieinejad, Masoumeh, et al.
Publicado: (2026) -
Interpretability Can Be Actionable
por: Orgad, Hadas, et al.
Publicado: (2026) -
Improving Time Series Classification with Representation Soft Label Smoothing
por: Ma, Hengyi, et al.
Publicado: (2024)