LOCUS: A System and Method for Low-Cost Customization for Universal Specialization
Fuente:
arXiv
Salvato in:
| Autori principali: | Sundararaman, Dhanasekar, Li, Keying, Xiong, Wayne, Garg, Aashna |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
di: Datta, Shounak, et al.
Pubblicazione: (2025)
di: Datta, Shounak, et al.
Pubblicazione: (2025)
RexBERT: Context Specialized Bidirectional Encoders for E-commerce
di: Bajaj, Rahul, et al.
Pubblicazione: (2026)
di: Bajaj, Rahul, et al.
Pubblicazione: (2026)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
di: Garg, Aashna, et al.
Pubblicazione: (2026)
di: Garg, Aashna, et al.
Pubblicazione: (2026)
CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
di: Yuan, Lifan, et al.
Pubblicazione: (2023)
di: Yuan, Lifan, et al.
Pubblicazione: (2023)
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
di: Fu, Yu, et al.
Pubblicazione: (2024)
di: Fu, Yu, et al.
Pubblicazione: (2024)
One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
di: Gu, Yiyang, et al.
Pubblicazione: (2026)
di: Gu, Yiyang, et al.
Pubblicazione: (2026)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
di: Shi, Luohe, et al.
Pubblicazione: (2024)
di: Shi, Luohe, et al.
Pubblicazione: (2024)
CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs
di: Shi, Jingzhe, et al.
Pubblicazione: (2024)
di: Shi, Jingzhe, et al.
Pubblicazione: (2024)
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems
di: Wang, Qianli, et al.
Pubblicazione: (2025)
di: Wang, Qianli, et al.
Pubblicazione: (2025)
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
di: Zhang, Kaiyan, et al.
Pubblicazione: (2024)
di: Zhang, Kaiyan, et al.
Pubblicazione: (2024)
Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design
di: Gao, Ming, et al.
Pubblicazione: (2024)
di: Gao, Ming, et al.
Pubblicazione: (2024)
AXE: Low-Cost Cross-Domain Web Structured Information Extraction
di: Mansour, Abdelrahman, et al.
Pubblicazione: (2026)
di: Mansour, Abdelrahman, et al.
Pubblicazione: (2026)
A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
di: Labat, Sofie, et al.
Pubblicazione: (2025)
di: Labat, Sofie, et al.
Pubblicazione: (2025)
Evaluating, Synthesizing, and Enhancing for Customer Support Conversation
di: Zhu, Jie, et al.
Pubblicazione: (2025)
di: Zhu, Jie, et al.
Pubblicazione: (2025)
Scalable Model Editing via Customized Expert Networks
di: Yao, Zihan, et al.
Pubblicazione: (2024)
di: Yao, Zihan, et al.
Pubblicazione: (2024)
Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment
di: Chen, Zhipeng, et al.
Pubblicazione: (2024)
di: Chen, Zhipeng, et al.
Pubblicazione: (2024)
Improving Value-based Process Verifier via Low-Cost Variance Reduction
di: Sun, Zetian, et al.
Pubblicazione: (2025)
di: Sun, Zetian, et al.
Pubblicazione: (2025)
Can Linguistically Related Languages Guide LLM Translation in Low-Resource Settings?
di: Ramasethu, Aishwarya, et al.
Pubblicazione: (2026)
di: Ramasethu, Aishwarya, et al.
Pubblicazione: (2026)
Continuous Semantic Caching for Low-Cost LLM Serving
di: Atalar, Baran, et al.
Pubblicazione: (2026)
di: Atalar, Baran, et al.
Pubblicazione: (2026)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs
di: Zhong, Wei, et al.
Pubblicazione: (2024)
di: Zhong, Wei, et al.
Pubblicazione: (2024)
LatentMem: Customizing Latent Memory for Multi-Agent Systems
di: Fu, Muxin, et al.
Pubblicazione: (2026)
di: Fu, Muxin, et al.
Pubblicazione: (2026)
GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
di: Du, Cunxiao, et al.
Pubblicazione: (2024)
di: Du, Cunxiao, et al.
Pubblicazione: (2024)
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs
di: Nag, Arijit, et al.
Pubblicazione: (2024)
di: Nag, Arijit, et al.
Pubblicazione: (2024)
Improvement in Semantic Address Matching using Natural Language Processing
di: Gupta, Vansh, et al.
Pubblicazione: (2024)
di: Gupta, Vansh, et al.
Pubblicazione: (2024)
Position: Vector Prompt Interfaces Should Be Exposed to Enable Customization of Large Language Models
di: Yang, Liangwei, et al.
Pubblicazione: (2026)
di: Yang, Liangwei, et al.
Pubblicazione: (2026)
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
di: Zhu, Lei, et al.
Pubblicazione: (2024)
di: Zhu, Lei, et al.
Pubblicazione: (2024)
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
di: Liu, Peiyu, et al.
Pubblicazione: (2024)
di: Liu, Peiyu, et al.
Pubblicazione: (2024)
A Survey on Long Text Modeling with Transformers
di: Dong, Zican, et al.
Pubblicazione: (2023)
di: Dong, Zican, et al.
Pubblicazione: (2023)
SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
di: Shi, Dachuan, et al.
Pubblicazione: (2025)
Chunking Methods on Retrieval-Augmented Generation - Effectiveness Evaluation Against Computational Cost and Limitations
di: Śmigielski, Mateusz, et al.
Pubblicazione: (2026)
di: Śmigielski, Mateusz, et al.
Pubblicazione: (2026)
Watermarking Low-entropy Generation for Large Language Models: An Unbiased and Low-risk Method
di: Mao, Minjia, et al.
Pubblicazione: (2024)
di: Mao, Minjia, et al.
Pubblicazione: (2024)
NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
di: Zhang, Yulong, et al.
Pubblicazione: (2025)
di: Zhang, Yulong, et al.
Pubblicazione: (2025)
An Iris for Expected Cost Analysis
di: Lohse, Janine, et al.
Pubblicazione: (2024)
di: Lohse, Janine, et al.
Pubblicazione: (2024)
When Models Reason in Your Language: Controlling Thinking Language Comes at the Cost of Accuracy
di: Qi, Jirui, et al.
Pubblicazione: (2025)
di: Qi, Jirui, et al.
Pubblicazione: (2025)
Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
di: Baghel, Shruti Singh, et al.
Pubblicazione: (2025)
di: Baghel, Shruti Singh, et al.
Pubblicazione: (2025)
Spec-TOD: A Specialized Instruction-Tuned LLM Framework for Efficient Task-Oriented Dialogue Systems
di: Nguyen, Quang-Vinh, et al.
Pubblicazione: (2025)
di: Nguyen, Quang-Vinh, et al.
Pubblicazione: (2025)
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
di: Li, Qintong, et al.
Pubblicazione: (2023)
di: Li, Qintong, et al.
Pubblicazione: (2023)
1+1>2: A Synergistic Sparse and Low-Rank Compression Method for Large Language Models
di: Zong, Zeliang, et al.
Pubblicazione: (2025)
di: Zong, Zeliang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
di: Datta, Shounak, et al.
Pubblicazione: (2025) -
RexBERT: Context Specialized Bidirectional Encoders for E-commerce
di: Bajaj, Rahul, et al.
Pubblicazione: (2026) -
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
di: Garg, Aashna, et al.
Pubblicazione: (2026) -
CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets
di: Yuan, Lifan, et al.
Pubblicazione: (2023) -
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
di: Fu, Yu, et al.
Pubblicazione: (2024)