Crayon: Customized On-Device LLM via Instant Adapter Blending and Edge-Server Hybrid Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Bang, Jihwan, Lee, Juntae, Shim, Kyuhong, Yang, Seunghan, Chang, Simyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Chain-of-Rank: Enhancing Large Language Models for Domain-Specific RAG in Edge Device
di: Lee, Juntae, et al.
Pubblicazione: (2025)
di: Lee, Juntae, et al.
Pubblicazione: (2025)
CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM
di: Lee, Juntae, et al.
Pubblicazione: (2025)
di: Lee, Juntae, et al.
Pubblicazione: (2025)
Feedback Adaptation for Retrieval-Augmented Generation
di: Bang, Jihwan, et al.
Pubblicazione: (2026)
di: Bang, Jihwan, et al.
Pubblicazione: (2026)
Learning Contextual Retrieval for Robust Conversational Search
di: Yang, Seunghan, et al.
Pubblicazione: (2025)
di: Yang, Seunghan, et al.
Pubblicazione: (2025)
Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
di: Bang, Jihwan, et al.
Pubblicazione: (2025)
di: Bang, Jihwan, et al.
Pubblicazione: (2025)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
di: Kim, Byeonggeun, et al.
Pubblicazione: (2024)
di: Kim, Byeonggeun, et al.
Pubblicazione: (2024)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
di: Kim, Minsoo, et al.
Pubblicazione: (2024)
di: Kim, Minsoo, et al.
Pubblicazione: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech
di: Lee, Yejin, et al.
Pubblicazione: (2025)
di: Lee, Yejin, et al.
Pubblicazione: (2025)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
di: Kim, Donghoon, et al.
Pubblicazione: (2024)
di: Kim, Donghoon, et al.
Pubblicazione: (2024)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
di: Kang, Jaehoon, et al.
Pubblicazione: (2026)
di: Kang, Jaehoon, et al.
Pubblicazione: (2026)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
di: Kim, Donghoon, et al.
Pubblicazione: (2025)
di: Kim, Donghoon, et al.
Pubblicazione: (2025)
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
di: Lee, Hanna, et al.
Pubblicazione: (2026)
di: Lee, Hanna, et al.
Pubblicazione: (2026)
ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback
di: Yun, Taewon, et al.
Pubblicazione: (2025)
di: Yun, Taewon, et al.
Pubblicazione: (2025)
Problem-Solving Guide: Predicting the Algorithm Tags and Difficulty for Competitive Programming Problems
di: Kim, Juntae, et al.
Pubblicazione: (2023)
di: Kim, Juntae, et al.
Pubblicazione: (2023)
RevMUX: Data Multiplexing with Reversible Adapters for Efficient LLM Batch Inference
di: Xu, Yige, et al.
Pubblicazione: (2024)
di: Xu, Yige, et al.
Pubblicazione: (2024)
Feature Diversification and Adaptation for Federated Domain Generalization
di: Yang, Seunghan, et al.
Pubblicazione: (2024)
di: Yang, Seunghan, et al.
Pubblicazione: (2024)
Token Level Routing Inference System for Edge Devices
di: She, Jianshu, et al.
Pubblicazione: (2025)
di: She, Jianshu, et al.
Pubblicazione: (2025)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
di: Wang, Zhenting, et al.
Pubblicazione: (2025)
Learning Primitive Relations for Compositional Zero-Shot Learning
di: Lee, Insu, et al.
Pubblicazione: (2025)
di: Lee, Insu, et al.
Pubblicazione: (2025)
Contextually Guided Transformers via Low-Rank Adaptation
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
References Indeed Matter? Reference-Free Preference Optimization for Conversational Query Reformulation
di: Kim, Doyoung, et al.
Pubblicazione: (2025)
di: Kim, Doyoung, et al.
Pubblicazione: (2025)
Learning to Summarize from LLM-generated Feedback
di: Song, Hwanjun, et al.
Pubblicazione: (2024)
di: Song, Hwanjun, et al.
Pubblicazione: (2024)
Personal Intelligence System UniLM: Hybrid On-Device Small Language Model and Server-Based Large Language Model for Malay Nusantara
di: Nazri, Azree, et al.
Pubblicazione: (2024)
di: Nazri, Azree, et al.
Pubblicazione: (2024)
EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
di: Chen, Jiyu, et al.
Pubblicazione: (2025)
di: Chen, Jiyu, et al.
Pubblicazione: (2025)
Resolving Word Vagueness with Scenario-guided Adapter for Natural Language Inference
di: Liu, Yonghao, et al.
Pubblicazione: (2024)
di: Liu, Yonghao, et al.
Pubblicazione: (2024)
Instant Personalized Large Language Model Adaptation via Hypernetwork
di: Tan, Zhaoxuan, et al.
Pubblicazione: (2025)
di: Tan, Zhaoxuan, et al.
Pubblicazione: (2025)
BlendX: Complex Multi-Intent Detection with Blended Patterns
di: Yoon, Yejin, et al.
Pubblicazione: (2024)
di: Yoon, Yejin, et al.
Pubblicazione: (2024)
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
di: Zhmoginov, Andrey, et al.
Pubblicazione: (2025)
Adapters for Altering LLM Vocabularies: What Languages Benefit the Most?
di: Han, HyoJung, et al.
Pubblicazione: (2024)
di: Han, HyoJung, et al.
Pubblicazione: (2024)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
di: Svirschevski, Ruslan, et al.
Pubblicazione: (2024)
di: Svirschevski, Ruslan, et al.
Pubblicazione: (2024)
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
di: Trivedi, Aakash, et al.
Pubblicazione: (2026)
di: Trivedi, Aakash, et al.
Pubblicazione: (2026)
Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices
di: Zai, Liu, et al.
Pubblicazione: (2026)
di: Zai, Liu, et al.
Pubblicazione: (2026)
PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch Collaboration
di: Zhan, Junfei, et al.
Pubblicazione: (2025)
di: Zhan, Junfei, et al.
Pubblicazione: (2025)
LLM-based User Profile Management for Recommender System
di: Bang, Seunghwan, et al.
Pubblicazione: (2025)
di: Bang, Seunghwan, et al.
Pubblicazione: (2025)
Faster LLM Inference via Sequential Monte Carlo
di: Emara, Yahya, et al.
Pubblicazione: (2026)
di: Emara, Yahya, et al.
Pubblicazione: (2026)
From Belief Entrenchment to Robust Reasoning in LLM Agents
di: Oh, Jihwan, et al.
Pubblicazione: (2025)
di: Oh, Jihwan, et al.
Pubblicazione: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
di: Guo, Jinyu, et al.
Pubblicazione: (2026)
di: Guo, Jinyu, et al.
Pubblicazione: (2026)
Pluggable Neural Machine Translation Models via Memory-augmented Adapters
di: Xu, Yuzhuang, et al.
Pubblicazione: (2023)
di: Xu, Yuzhuang, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Chain-of-Rank: Enhancing Large Language Models for Domain-Specific RAG in Edge Device
di: Lee, Juntae, et al.
Pubblicazione: (2025) -
CIFLEX: Contextual Instruction Flow for Sub-task Execution in Multi-Turn Interactions with a Single On-Device LLM
di: Lee, Juntae, et al.
Pubblicazione: (2025) -
Feedback Adaptation for Retrieval-Augmented Generation
di: Bang, Jihwan, et al.
Pubblicazione: (2026) -
Learning Contextual Retrieval for Robust Conversational Search
di: Yang, Seunghan, et al.
Pubblicazione: (2025) -
Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
di: Bang, Jihwan, et al.
Pubblicazione: (2025)