Towards Optimizing the Costs of LLM Usage
Fuente:
arXiv
Saved in:
| Main Authors: | Shekhar, Shivanshu, Dubey, Tanishq, Mukherjee, Koyel, Saxena, Apoorv, Tyagi, Atharv, Kotla, Nishanth |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
by: Anand, Nikhil, et al.
Published: (2026)
by: Anand, Nikhil, et al.
Published: (2026)
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
by: Mahapatra, Debabrata, et al.
Published: (2025)
by: Mahapatra, Debabrata, et al.
Published: (2025)
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
by: Lee, Jaeho, et al.
Published: (2025)
by: Lee, Jaeho, et al.
Published: (2025)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
ROCM: RLHF on consistency models
by: Shekhar, Shivanshu, et al.
Published: (2025)
by: Shekhar, Shivanshu, et al.
Published: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
PLD+: Accelerating LLM inference by leveraging Language Model Artifacts
by: Somasundaram, Shwetha, et al.
Published: (2024)
by: Somasundaram, Shwetha, et al.
Published: (2024)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
by: Pandey, Atharva, et al.
Published: (2025)
by: Pandey, Atharva, et al.
Published: (2025)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
by: Xia, Yuchen, et al.
Published: (2024)
by: Xia, Yuchen, et al.
Published: (2024)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
by: Salinas, David, et al.
Published: (2025)
by: Salinas, David, et al.
Published: (2025)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
by: Saunshi, Nikunj, et al.
Published: (2025)
by: Saunshi, Nikunj, et al.
Published: (2025)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
by: Adamska, Marta, et al.
Published: (2025)
by: Adamska, Marta, et al.
Published: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
by: Zhang, Wenbo, et al.
Published: (2026)
by: Zhang, Wenbo, et al.
Published: (2026)
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
by: Vatsal, Shubham, et al.
Published: (2025)
by: Vatsal, Shubham, et al.
Published: (2025)
SciNets: Graph-Constrained Multi-Hop Reasoning for Scientific Literature Synthesis
by: Dubey, Sauhard
Published: (2025)
by: Dubey, Sauhard
Published: (2025)
A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning
by: Hong, Guan Zhe, et al.
Published: (2024)
by: Hong, Guan Zhe, et al.
Published: (2024)
Towards Multilingual LLM Evaluation for European Languages
by: Thellmann, Klaudia, et al.
Published: (2024)
by: Thellmann, Klaudia, et al.
Published: (2024)
Pyramid MoA: A Probabilistic Framework for Cost-Optimized Anytime Inference
by: Khaled, Arindam
Published: (2026)
by: Khaled, Arindam
Published: (2026)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)
by: Kim, Taeho, et al.
Published: (2024)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
by: Rao, Varun, et al.
Published: (2025)
by: Rao, Varun, et al.
Published: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
MPO: Boosting LLM Agents with Meta Plan Optimization
by: Xiong, Weimin, et al.
Published: (2025)
by: Xiong, Weimin, et al.
Published: (2025)
Contextually Entangled Gradient Mapping for Optimized LLM Comprehension
by: Sisate, Colin, et al.
Published: (2025)
by: Sisate, Colin, et al.
Published: (2025)
CAPO: Cost-Aware Prompt Optimization
by: Zehle, Tom, et al.
Published: (2025)
by: Zehle, Tom, et al.
Published: (2025)
CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
by: Xie, Guofu, et al.
Published: (2025)
by: Xie, Guofu, et al.
Published: (2025)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
by: Bhattacharjee, Amrita, et al.
Published: (2023)
by: Bhattacharjee, Amrita, et al.
Published: (2023)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
by: Song, Yifan, et al.
Published: (2024)
by: Song, Yifan, et al.
Published: (2024)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
ReDit: Reward Dithering for Improved LLM Policy Optimization
by: Wei, Chenxing, et al.
Published: (2025)
by: Wei, Chenxing, et al.
Published: (2025)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
Quantifying Modality Contributions via Disentangling Multimodal Representations
by: Amit, Padegal, et al.
Published: (2025)
by: Amit, Padegal, et al.
Published: (2025)
Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
by: An, Jisu, et al.
Published: (2025)
by: An, Jisu, et al.
Published: (2025)
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
by: Chen, Hongzheng, et al.
Published: (2025)
by: Chen, Hongzheng, et al.
Published: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
by: Panaganti, Kishan, et al.
Published: (2026)
by: Panaganti, Kishan, et al.
Published: (2026)
Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
by: Wang, Ziyan, et al.
Published: (2025)
by: Wang, Ziyan, et al.
Published: (2025)
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
by: Wu, Siyang, et al.
Published: (2025)
by: Wu, Siyang, et al.
Published: (2025)
Similar Items
-
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
by: Anand, Nikhil, et al.
Published: (2026) -
RCStat: A Statistical Framework for using Relative Contextualization in Transformers
by: Mahapatra, Debabrata, et al.
Published: (2025) -
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
by: Lee, Jaeho, et al.
Published: (2025) -
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025) -
ROCM: RLHF on consistency models
by: Shekhar, Shivanshu, et al.
Published: (2025)