Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs
Fuente:
arXiv
Saved in:
| Main Authors: | Seymour, Liam, Kutukcu, Basar, Baidya, Sabur |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JExplore: Design Space Exploration Tool for Nvidia Jetson Boards
by: Kutukcu, Basar, et al.
Published: (2025)
by: Kutukcu, Basar, et al.
Published: (2025)
CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
by: Uddin, Mohammad Helal, et al.
Published: (2025)
by: Uddin, Mohammad Helal, et al.
Published: (2025)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
by: Hazra, Rishi, et al.
Published: (2025)
by: Hazra, Rishi, et al.
Published: (2025)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Diffusion Language Models are Provably Optimal Parallel Samplers
by: Jiang, Haozhe, et al.
Published: (2025)
by: Jiang, Haozhe, et al.
Published: (2025)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
by: Uddin, Mohammad Helal, et al.
Published: (2025)
by: Uddin, Mohammad Helal, et al.
Published: (2025)
Smoothed Analysis for Learning Concepts with Low Intrinsic Dimension
by: Chandrasekaran, Gautam, et al.
Published: (2024)
by: Chandrasekaran, Gautam, et al.
Published: (2024)
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
by: Fan, Lizhou, et al.
Published: (2023)
by: Fan, Lizhou, et al.
Published: (2023)
Emissions and Performance Trade-off Between Small and Large Language Models
by: Garg, Anandita, et al.
Published: (2025)
by: Garg, Anandita, et al.
Published: (2025)
Additive Models Explained: A Computational Complexity Approach
by: Bassan, Shahaf, et al.
Published: (2025)
by: Bassan, Shahaf, et al.
Published: (2025)
The Fine-Grained Complexity of Gradient Computation for Training Large Language Models
by: Alman, Josh, et al.
Published: (2024)
by: Alman, Josh, et al.
Published: (2024)
On Efficiently Representing Regular Languages as RNNs
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent
by: Chen, Bo, et al.
Published: (2025)
by: Chen, Bo, et al.
Published: (2025)
Inference Scaling vs Reasoning: An Empirical Analysis of Compute-Optimal LLM Problem-Solving
by: AbdElhameed, Marwan, et al.
Published: (2024)
by: AbdElhameed, Marwan, et al.
Published: (2024)
Learnability of Parameter-Bounded Bayes Nets
by: Bhattacharyya, Arnab, et al.
Published: (2024)
by: Bhattacharyya, Arnab, et al.
Published: (2024)
Statistical and Computational Guarantees of Kernel Max-Sliced Wasserstein Distances
by: Wang, Jie, et al.
Published: (2024)
by: Wang, Jie, et al.
Published: (2024)
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Ask, and it shall be given: On the Turing completeness of prompting
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
Data Debugging is NP-hard for Classifiers Trained with SGD
by: Guo, Zizheng, et al.
Published: (2024)
by: Guo, Zizheng, et al.
Published: (2024)
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformers
by: Merrill, William, et al.
Published: (2025)
by: Merrill, William, et al.
Published: (2025)
How Hard Is Continuous Clustering? Lower Bounds from the Existential Theory of the Reals
by: Majumdar, Angshul
Published: (2026)
by: Majumdar, Angshul
Published: (2026)
Spiky Rank and Its Applications to Rigidity and Circuits
by: Hambardzumyan, Lianna, et al.
Published: (2026)
by: Hambardzumyan, Lianna, et al.
Published: (2026)
Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete
by: Li, Qian, et al.
Published: (2026)
by: Li, Qian, et al.
Published: (2026)
Low-Rank Matrix Approximation for Neural Network Compression
by: Cherukuri, Kalyan, et al.
Published: (2025)
by: Cherukuri, Kalyan, et al.
Published: (2025)
Proximity to Losslessly Compressible Parameters
by: Farrugia-Roberts, Matthew
Published: (2023)
by: Farrugia-Roberts, Matthew
Published: (2023)
Decision Tree Learning on Product Spaces
by: Moakahr, Arshia Soltani, et al.
Published: (2026)
by: Moakahr, Arshia Soltani, et al.
Published: (2026)
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers
by: Amiri, Alireza, et al.
Published: (2025)
by: Amiri, Alireza, et al.
Published: (2025)
Fundamental Limits of Crystalline Equivariant Graph Neural Networks: A Circuit Complexity Perspective
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
A Logic for Expressing Log-Precision Transformers
by: Merrill, William, et al.
Published: (2022)
by: Merrill, William, et al.
Published: (2022)
Optimizing Computational-Statistical Runtime for Wasserstein Distance Estimation
by: Jacobs, Peter Matthew, et al.
Published: (2026)
by: Jacobs, Peter Matthew, et al.
Published: (2026)
Distribution-Specific Agnostic Conditional Classification With Halfspaces
by: Huang, Jizhou, et al.
Published: (2025)
by: Huang, Jizhou, et al.
Published: (2025)
From Pseudorandomness to Multi-Group Fairness and Back
by: Dwork, Cynthia, et al.
Published: (2023)
by: Dwork, Cynthia, et al.
Published: (2023)
Necessary and Sufficient Oracles: Toward a Computational Taxonomy For Reinforcement Learning
by: Rohatgi, Dhruv, et al.
Published: (2025)
by: Rohatgi, Dhruv, et al.
Published: (2025)
How Global Calibration Strengthens Multiaccuracy
by: Casacuberta, Sílvia, et al.
Published: (2025)
by: Casacuberta, Sílvia, et al.
Published: (2025)
On the Computational Hardness of Transformers
by: Saha, Barna, et al.
Published: (2026)
by: Saha, Barna, et al.
Published: (2026)
Certifiable Boolean Reasoning Is Universal
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Constant Bit-size Transformers Are Turing Complete
by: Li, Qian, et al.
Published: (2025)
by: Li, Qian, et al.
Published: (2025)
New Hardness Results for Low-Rank Matrix Completion
by: Chawin, Dror, et al.
Published: (2025)
by: Chawin, Dror, et al.
Published: (2025)
Reachability In Simple Neural Networks
by: Sälzer, Marco, et al.
Published: (2022)
by: Sälzer, Marco, et al.
Published: (2022)
Similar Items
-
JExplore: Design Space Exploration Tool for Nvidia Jetson Boards
by: Kutukcu, Basar, et al.
Published: (2025) -
CAMP-HiVe: Cyclic Pair Merging based Efficient DNN Pruning with Hessian-Vector Approximation for Resource-Constrained Systems
by: Uddin, Mohammad Helal, et al.
Published: (2025) -
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
by: Hazra, Rishi, et al.
Published: (2025) -
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
by: Chen, Yifang, et al.
Published: (2025) -
Diffusion Language Models are Provably Optimal Parallel Samplers
by: Jiang, Haozhe, et al.
Published: (2025)