Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Shuming, Zhao, Enbo, Cai, Deng, Cui, Leyang, Huang, Xinting, Li, Huayang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
von: Yang, Sen, et al.
Veröffentlicht: (2024)
von: Yang, Sen, et al.
Veröffentlicht: (2024)
Knowledge Fusion of Large Language Models
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
von: Zhang, Yue, et al.
Veröffentlicht: (2023)
Knowledge Verification to Nip Hallucination in the Bud
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)
On the Transformations across Reward Model, Parameter Update, and In-Context Prompt
von: Cai, Deng, et al.
Veröffentlicht: (2024)
von: Cai, Deng, et al.
Veröffentlicht: (2024)
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2024)
von: Chung, Tsz Ting, et al.
Veröffentlicht: (2024)
A Frustratingly Simple Decoding Method for Neural Text Generation
von: Yang, Haoran, et al.
Veröffentlicht: (2023)
von: Yang, Haoran, et al.
Veröffentlicht: (2023)
CORM: Cache Optimization with Recent Message for Large Language Model Inference
von: Dai, Jincheng, et al.
Veröffentlicht: (2024)
von: Dai, Jincheng, et al.
Veröffentlicht: (2024)
Retrieval is Accurate Generation
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
von: Cao, Bowen, et al.
Veröffentlicht: (2024)
Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
von: Huang, Jianheng, et al.
Veröffentlicht: (2024)
von: Huang, Jianheng, et al.
Veröffentlicht: (2024)
RePo: Language Models with Context Re-Positioning
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving
von: Gao, Chang, et al.
Veröffentlicht: (2023)
von: Gao, Chang, et al.
Veröffentlicht: (2023)
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
von: Li, Qintong, et al.
Veröffentlicht: (2023)
von: Li, Qintong, et al.
Veröffentlicht: (2023)
DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
von: Chen, Yongrui, et al.
Veröffentlicht: (2023)
von: Chen, Yongrui, et al.
Veröffentlicht: (2023)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
Reasons to Reject? Aligning Language Models with Judgments
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
von: Xu, Weiwen, et al.
Veröffentlicht: (2023)
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
von: Dong, Yixin, et al.
Veröffentlicht: (2024)
von: Dong, Yixin, et al.
Veröffentlicht: (2024)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
von: Li, Yafu, et al.
Veröffentlicht: (2024)
von: Li, Yafu, et al.
Veröffentlicht: (2024)
TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild
von: Li, Huayang, et al.
Veröffentlicht: (2023)
von: Li, Huayang, et al.
Veröffentlicht: (2023)
DROJ: A Prompt-Driven Attack against Large Language Models
von: Hu, Leyang, et al.
Veröffentlicht: (2024)
von: Hu, Leyang, et al.
Veröffentlicht: (2024)
A Survey on Inference Engines for Large Language Models: Perspectives on Optimization and Efficiency
von: Park, Sihyeong, et al.
Veröffentlicht: (2025)
von: Park, Sihyeong, et al.
Veröffentlicht: (2025)
MAGE: Machine-generated Text Detection in the Wild
von: Li, Yafu, et al.
Veröffentlicht: (2023)
von: Li, Yafu, et al.
Veröffentlicht: (2023)
Efficient Tuning and Inference for Large Language Models on Textual Graphs
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
von: Yu, Pengfei, et al.
Veröffentlicht: (2025)
von: Yu, Pengfei, et al.
Veröffentlicht: (2025)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
Cross-lingual Contextualized Phrase Retrieval
von: Li, Huayang, et al.
Veröffentlicht: (2024)
von: Li, Huayang, et al.
Veröffentlicht: (2024)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
von: Huang, You-Liang, et al.
Veröffentlicht: (2026)
von: Huang, You-Liang, et al.
Veröffentlicht: (2026)
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects
von: Zhang, Jun, et al.
Veröffentlicht: (2026)
von: Zhang, Jun, et al.
Veröffentlicht: (2026)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
von: Li, Qintong, et al.
Veröffentlicht: (2024)
von: Li, Qintong, et al.
Veröffentlicht: (2024)
A Survey on Efficient Inference for Large Language Models
von: Zhou, Zixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Zixuan, et al.
Veröffentlicht: (2024)
Large Language Models as Search Engines: Societal Challenges
von: Sadeddine, Zacchary, et al.
Veröffentlicht: (2025)
von: Sadeddine, Zacchary, et al.
Veröffentlicht: (2025)
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
von: Pang, Jianhui, et al.
Veröffentlicht: (2024)
Dual-Density Inference for Efficient Language Model Reasoning
von: Zhao, Zhengyi, et al.
Veröffentlicht: (2025)
von: Zhao, Zhengyi, et al.
Veröffentlicht: (2025)
The End of Manual Decoding: Towards Truly End-to-End Language Models
von: Wang, Zhichao, et al.
Veröffentlicht: (2025)
von: Wang, Zhichao, et al.
Veröffentlicht: (2025)
Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via Alignment Tax Reduction
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
Efficient Inference for Large Reasoning Models: A Survey
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
von: Yang, Sen, et al.
Veröffentlicht: (2024) -
Knowledge Fusion of Large Language Models
von: Wan, Fanqi, et al.
Veröffentlicht: (2024) -
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
von: Zhang, Yue, et al.
Veröffentlicht: (2023) -
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2023) -
Knowledge Verification to Nip Hallucination in the Bud
von: Wan, Fanqi, et al.
Veröffentlicht: (2024)