LLMs for Analog Circuit Design Continuum (ACDC)
Fuente:
arXiv
Salvato in:
| Autori principali: | Esfandiari, Yasaman, Rego, Jocelyn, Meyer, Austin, Gallagher, Jonathan, Levy, Mia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EXAQ: Exponent Aware Quantization For LLMs Acceleration
di: Shkolnik, Moran, et al.
Pubblicazione: (2024)
di: Shkolnik, Moran, et al.
Pubblicazione: (2024)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
di: Hourri, Younes, et al.
Pubblicazione: (2025)
di: Hourri, Younes, et al.
Pubblicazione: (2025)
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
di: Knoop, Jonathan, et al.
Pubblicazione: (2026)
di: Knoop, Jonathan, et al.
Pubblicazione: (2026)
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026)
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2025)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2025)
MoEITS: A Green AI approach for simplifying MoE-LLMs
di: Balderas, Luis, et al.
Pubblicazione: (2026)
di: Balderas, Luis, et al.
Pubblicazione: (2026)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
di: Huang, Zixiao, et al.
Pubblicazione: (2025)
di: Huang, Zixiao, et al.
Pubblicazione: (2025)
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
di: Ellis-Mohr, Austin R., et al.
Pubblicazione: (2025)
di: Ellis-Mohr, Austin R., et al.
Pubblicazione: (2025)
REAM: Merging Improves Pruning of Experts in LLMs
di: Jha, Saurav, et al.
Pubblicazione: (2026)
di: Jha, Saurav, et al.
Pubblicazione: (2026)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
di: Israel, Daniel, et al.
Pubblicazione: (2025)
di: Israel, Daniel, et al.
Pubblicazione: (2025)
KernelBench: Can LLMs Write Efficient GPU Kernels?
di: Ouyang, Anne, et al.
Pubblicazione: (2025)
di: Ouyang, Anne, et al.
Pubblicazione: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
di: Tran, Nguyen Phuc, et al.
Pubblicazione: (2025)
di: Tran, Nguyen Phuc, et al.
Pubblicazione: (2025)
SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025)
di: AbouElhamayed, Ahmed F., et al.
Pubblicazione: (2025)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
di: Bi, Zhen, et al.
Pubblicazione: (2026)
di: Bi, Zhen, et al.
Pubblicazione: (2026)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models
di: Shao, Zishan, et al.
Pubblicazione: (2025)
di: Shao, Zishan, et al.
Pubblicazione: (2025)
Quantum Neural Networks for Wind Energy Forecasting: A Comparative Study of Performance and Scalability with Classical Models
di: Hangun, Batuhan, et al.
Pubblicazione: (2025)
di: Hangun, Batuhan, et al.
Pubblicazione: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
di: Hossain, Md Arafat, et al.
Pubblicazione: (2025)
di: Hossain, Md Arafat, et al.
Pubblicazione: (2025)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
di: Yin, Yishu, et al.
Pubblicazione: (2025)
di: Yin, Yishu, et al.
Pubblicazione: (2025)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
di: Zhao, Yushang, et al.
Pubblicazione: (2025)
di: Zhao, Yushang, et al.
Pubblicazione: (2025)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
di: Kermani, Arshia, et al.
Pubblicazione: (2025)
di: Kermani, Arshia, et al.
Pubblicazione: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
di: Qiao, Liang, et al.
Pubblicazione: (2025)
di: Qiao, Liang, et al.
Pubblicazione: (2025)
Exchangeability in Neural Network and its Application to Dynamic Pruning
di: Pu, et al.
Pubblicazione: (2025)
di: Pu, et al.
Pubblicazione: (2025)
Knowledge Grafting: A Mechanism for Optimizing AI Model Deployment in Resource-Constrained Environments
di: Almurshed, Osama, et al.
Pubblicazione: (2025)
di: Almurshed, Osama, et al.
Pubblicazione: (2025)
The Race to Efficiency: A New Perspective on AI Scaling Laws
di: Lu, Chien-Ping
Pubblicazione: (2025)
di: Lu, Chien-Ping
Pubblicazione: (2025)
Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
di: Avinash, MSR
Pubblicazione: (2025)
di: Avinash, MSR
Pubblicazione: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
AdaGradSelect: An adaptive gradient-guided layer selection method for efficient fine-tuning of SLMs
di: Kumar, Anshul, et al.
Pubblicazione: (2025)
di: Kumar, Anshul, et al.
Pubblicazione: (2025)
On the Sustainability of AI Inferences in the Edge
di: Sobhani, Ghazal, et al.
Pubblicazione: (2025)
di: Sobhani, Ghazal, et al.
Pubblicazione: (2025)
Knowledge Distillation for Reservoir-based Classifier: Human Activity Recognition
di: Kagiyama, Masaharu, et al.
Pubblicazione: (2025)
di: Kagiyama, Masaharu, et al.
Pubblicazione: (2025)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
di: Estevez, Melissa, et al.
Pubblicazione: (2025)
di: Estevez, Melissa, et al.
Pubblicazione: (2025)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
di: Yi, Qingao, et al.
Pubblicazione: (2025)
di: Yi, Qingao, et al.
Pubblicazione: (2025)
Towards Generalized Parameter Tuning in Coherent Ising Machines: A Portfolio-Based Approach
di: Hanyu, Tatsuro, et al.
Pubblicazione: (2025)
di: Hanyu, Tatsuro, et al.
Pubblicazione: (2025)
Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial
di: Cortes, David, et al.
Pubblicazione: (2025)
di: Cortes, David, et al.
Pubblicazione: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
di: Liu, Jiashuo, et al.
Pubblicazione: (2025)
di: Liu, Jiashuo, et al.
Pubblicazione: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
di: Georganas, Evangelos, et al.
Pubblicazione: (2025)
EcoSpa: Efficient Transformer Training with Coupled Sparsity
di: Xiao, Jinqi, et al.
Pubblicazione: (2025)
di: Xiao, Jinqi, et al.
Pubblicazione: (2025)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
di: Yin, Wangsong, et al.
Pubblicazione: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2024)
Deploying Open-Source Large Language Models: A performance Analysis
di: Bendi-Ouis, Yannis, et al.
Pubblicazione: (2024)
di: Bendi-Ouis, Yannis, et al.
Pubblicazione: (2024)
Documenti analoghi
-
EXAQ: Exponent Aware Quantization For LLMs Acceleration
di: Shkolnik, Moran, et al.
Pubblicazione: (2024) -
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
di: Hourri, Younes, et al.
Pubblicazione: (2025) -
Private LLM Inference on Consumer Blackwell GPUs: A Practical Guide for Cost-Effective Local Deployment in SMEs
di: Knoop, Jonathan, et al.
Pubblicazione: (2026) -
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs
di: Yun, Vincent-Daniel, et al.
Pubblicazione: (2026) -
OPTIMA: Optimal One-shot Pruning for LLMs via Quadratic Programming Reconstruction
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2025)