Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability
Fuente:
arXiv
Saved in:
| Main Author: | Huang, Hen-Hsen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Known Facts: Generating Unseen Temporal Knowledge to Address Data Contamination in LLM Evaluation
by: Amalvy, Arthur, et al.
Published: (2026)
by: Amalvy, Arthur, et al.
Published: (2026)
Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
by: Amalvy, Arthur, et al.
Published: (2026)
by: Amalvy, Arthur, et al.
Published: (2026)
No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs
by: Wu, Wei-Chi, et al.
Published: (2026)
by: Wu, Wei-Chi, et al.
Published: (2026)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
by: Lin, Wei-Hsiang, et al.
Published: (2025)
by: Lin, Wei-Hsiang, et al.
Published: (2025)
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
by: Chuang, Ko-Wei, et al.
Published: (2025)
by: Chuang, Ko-Wei, et al.
Published: (2025)
Diverge to Induce Prompting: Multi-Rationale Induction for Zero-Shot Reasoning
by: Chen, Po-Chun, et al.
Published: (2026)
by: Chen, Po-Chun, et al.
Published: (2026)
Strategy-Induct: Task-Level Strategy Induction for Instruction Generation
by: Chen, Po-Chun, et al.
Published: (2026)
by: Chen, Po-Chun, et al.
Published: (2026)
Personalized Graph-Empowered Large Language Model for Proactive Information Access
by: Chang, Chia Cheng, et al.
Published: (2026)
by: Chang, Chia Cheng, et al.
Published: (2026)
Co-Trained Retriever-Generator Framework for Question Generation in Earnings Calls
by: Juan, Yining, et al.
Published: (2024)
by: Juan, Yining, et al.
Published: (2024)
Diagnosing Model Editing via Knowledge Spectrum
by: Pan, Tsung-Hsuan, et al.
Published: (2025)
by: Pan, Tsung-Hsuan, et al.
Published: (2025)
Pre-Finetuning with Impact Duration Awareness for Stock Movement Prediction
by: Chiu, Chr-Jr, et al.
Published: (2024)
by: Chiu, Chr-Jr, et al.
Published: (2024)
"Why" Has the Least Side Effect on Model Editing
by: Pan, Tsung-Hsuan, et al.
Published: (2024)
by: Pan, Tsung-Hsuan, et al.
Published: (2024)
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
by: Chan, Brian J, et al.
Published: (2024)
by: Chan, Brian J, et al.
Published: (2024)
Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models
by: Wei, Sheng-Lun, et al.
Published: (2024)
by: Wei, Sheng-Lun, et al.
Published: (2024)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
by: Shih, Yu-Fei, et al.
Published: (2025)
by: Shih, Yu-Fei, et al.
Published: (2025)
Efficient Beam Search for Large Language Models Using Trie-Based Decoding
by: Chan, Brian J, et al.
Published: (2025)
by: Chan, Brian J, et al.
Published: (2025)
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models
by: Yeh, Cheng-Kai, et al.
Published: (2025)
by: Yeh, Cheng-Kai, et al.
Published: (2025)
Enhancing Investment Opinion Ranking through Argument-Based Sentiment Analysis
by: Chen, Chung-Chi, et al.
Published: (2024)
by: Chen, Chung-Chi, et al.
Published: (2024)
From Opinion Mining to Financial Argument Mining
by: Chen, Chung-Chi, et al.
Published: (2021)
by: Chen, Chung-Chi, et al.
Published: (2021)
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
by: Wei, Sheng-Lun, et al.
Published: (2026)
by: Wei, Sheng-Lun, et al.
Published: (2026)
SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework
by: Huang, Mao Xun, et al.
Published: (2025)
by: Huang, Mao Xun, et al.
Published: (2025)
Vocabulary Customization for Efficient Domain-Specific LLM Deployment
by: Herold, Christian, et al.
Published: (2025)
by: Herold, Christian, et al.
Published: (2025)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
AE-LLM: Adaptive Efficiency Optimization for Large Language Models
by: Tanaka, Kaito, et al.
Published: (2026)
by: Tanaka, Kaito, et al.
Published: (2026)
Grounded in Reality: Learning and Deploying Proactive LLM from Offline Logs
by: Wei, Fei, et al.
Published: (2025)
by: Wei, Fei, et al.
Published: (2025)
The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management
by: Shen, Binqi, et al.
Published: (2026)
by: Shen, Binqi, et al.
Published: (2026)
RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment
by: Luo, Yingfeng, et al.
Published: (2026)
by: Luo, Yingfeng, et al.
Published: (2026)
Hierarchical Chain-of-Thought Prompting: Enhancing LLM Reasoning Performance and Efficiency
by: Huang, Xingshuai, et al.
Published: (2026)
by: Huang, Xingshuai, et al.
Published: (2026)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
by: Chen, Weize, et al.
Published: (2024)
by: Chen, Weize, et al.
Published: (2024)
LLM-Based Insight Extraction for Contact Center Analytics and Cost-Efficient Deployment
by: Embar, Varsha, et al.
Published: (2025)
by: Embar, Varsha, et al.
Published: (2025)
OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Comparative Efficiency Analysis of Lightweight Transformer Models: A Multi-Domain Empirical Benchmark for Enterprise NLP Deployment
by: Khan, Muhammad Shahmeer
Published: (2026)
by: Khan, Muhammad Shahmeer
Published: (2026)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025)
by: Mekky, Ali, et al.
Published: (2025)
Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
by: Bansal, Mayank, et al.
Published: (2026)
by: Bansal, Mayank, et al.
Published: (2026)
Understanding "Democratization" in NLP and ML Research
by: Subramonian, Arjun, et al.
Published: (2024)
by: Subramonian, Arjun, et al.
Published: (2024)
MoECollab: Democratizing LLM Development Through Collaborative Mixture of Experts
by: Harshit
Published: (2025)
by: Harshit
Published: (2025)
From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization
by: Cui, Chaoqun, et al.
Published: (2026)
by: Cui, Chaoqun, et al.
Published: (2026)
Tuning-Free Accountable Intervention for LLM Deployment -- A Metacognitive Approach
by: Tan, Zhen, et al.
Published: (2024)
by: Tan, Zhen, et al.
Published: (2024)
Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
by: He, Bowei, et al.
Published: (2025)
by: He, Bowei, et al.
Published: (2025)
Similar Items
-
Beyond Known Facts: Generating Unseen Temporal Knowledge to Address Data Contamination in LLM Evaluation
by: Amalvy, Arthur, et al.
Published: (2026) -
Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing
by: Amalvy, Arthur, et al.
Published: (2026) -
No One Fits All: From Fixed Prompting to Learned Routing in Multilingual LLMs
by: Wu, Wei-Chi, et al.
Published: (2026) -
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
by: Lin, Wei-Hsiang, et al.
Published: (2025) -
DINA: A Dual Defense Framework Against Internal Noise and External Attacks in Natural Language Processing
by: Chuang, Ko-Wei, et al.
Published: (2025)