The Solution for the AIGC Inference Performance Optimization Competition
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Sishun, Xu, Haonan, Wan, Zhonghua, Yang, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Solution for the sequential task continual learning track of the 2nd Greater Bay Area International Algorithm Competition
by: Pan, Sishun, et al.
Published: (2024)
by: Pan, Sishun, et al.
Published: (2024)
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
by: Huang, Longfei, et al.
Published: (2024)
by: Huang, Longfei, et al.
Published: (2024)
The Solution for Language-Enhanced Image New Category Discovery
by: Xu, Haonan, et al.
Published: (2024)
by: Xu, Haonan, et al.
Published: (2024)
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
by: Wu, Jie, et al.
Published: (2026)
by: Wu, Jie, et al.
Published: (2026)
The Solution for The PST-KDD-2024 OAG-Challenge
by: Zhong, Shupeng, et al.
Published: (2024)
by: Zhong, Shupeng, et al.
Published: (2024)
Evolving LLMs' Self-Refinement Capability via Synergistic Training-Inference Optimization
by: Zeng, Yongcheng, et al.
Published: (2025)
by: Zeng, Yongcheng, et al.
Published: (2025)
Optimizing Temperature for Language Models with Multi-Sample Inference
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
New Solutions on LLM Acceleration, Optimization, and Application
by: Huang, Yingbing, et al.
Published: (2024)
by: Huang, Yingbing, et al.
Published: (2024)
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
by: Guo, Qingpei, et al.
Published: (2025)
by: Guo, Qingpei, et al.
Published: (2025)
Pre-trained Language Models Improve the Few-shot Prompt Ability of Decision Transformer
by: Yang, Yu, et al.
Published: (2024)
by: Yang, Yu, et al.
Published: (2024)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
by: Łazuka, Małgorzata, et al.
Published: (2024)
by: Łazuka, Małgorzata, et al.
Published: (2024)
Neuro-Symbolic Contrastive Learning for Cross-domain Inference
by: Liu, Mingyue, et al.
Published: (2025)
by: Liu, Mingyue, et al.
Published: (2025)
Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling
by: Luo, Xianzhen, et al.
Published: (2024)
by: Luo, Xianzhen, et al.
Published: (2024)
Task Vector Geometry Underlies Dual Modes of Task Inference in Transformers
by: Yan, Hao, et al.
Published: (2026)
by: Yan, Hao, et al.
Published: (2026)
Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo
by: Wang, Weixin, et al.
Published: (2026)
by: Wang, Weixin, et al.
Published: (2026)
Optimal Stopping vs Best-of-$N$ for Inference Time Optimization
by: Kalayci, Yusuf, et al.
Published: (2025)
by: Kalayci, Yusuf, et al.
Published: (2025)
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
by: Fernandez, Jared, et al.
Published: (2025)
by: Fernandez, Jared, et al.
Published: (2025)
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
Research on Multi-hop Inference Optimization of LLM Based on MQUAKE Framework
by: Liang, Zucheng, et al.
Published: (2025)
by: Liang, Zucheng, et al.
Published: (2025)
Blackbox Model Provenance via Palimpsestic Membership Inference
by: Kuditipudi, Rohith, et al.
Published: (2025)
by: Kuditipudi, Rohith, et al.
Published: (2025)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
by: Ye, Jiancai, et al.
Published: (2026)
by: Ye, Jiancai, et al.
Published: (2026)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution
by: Li, Kaiyuan, et al.
Published: (2026)
by: Li, Kaiyuan, et al.
Published: (2026)
AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer
by: Yang, Leixin, et al.
Published: (2023)
by: Yang, Leixin, et al.
Published: (2023)
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
Cascade Speculative Drafting for Even Faster LLM Inference
by: Chen, Ziyi, et al.
Published: (2023)
by: Chen, Ziyi, et al.
Published: (2023)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
The Solution for the CVPR 2023 1st foundation model challenge-Track2
by: Xu, Haonan, et al.
Published: (2024)
by: Xu, Haonan, et al.
Published: (2024)
OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling
by: Lou, Yuxuan, et al.
Published: (2026)
by: Lou, Yuxuan, et al.
Published: (2026)
Calibrate to Discriminate: Improve In-Context Learning with Label-Free Comparative Inference
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
Continuous Diffusion Scales Competitively with Discrete Diffusion for Language
by: Yang, Zhihan, et al.
Published: (2026)
by: Yang, Zhihan, et al.
Published: (2026)
NJUST-KMG at TRAC-2024 Tasks 1 and 2: Offline Harm Potential Identification
by: Wang, Jingyuan, et al.
Published: (2024)
by: Wang, Jingyuan, et al.
Published: (2024)
Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
by: Wang, Guoli, et al.
Published: (2026)
by: Wang, Guoli, et al.
Published: (2026)
Faster MoE LLM Inference for Extremely Large Models
by: Yang, Haoqi, et al.
Published: (2025)
by: Yang, Haoqi, et al.
Published: (2025)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
by: Park, Core Francisco, et al.
Published: (2024)
by: Park, Core Francisco, et al.
Published: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
by: Si, Chongjie, et al.
Published: (2025)
by: Si, Chongjie, et al.
Published: (2025)
Similar Items
-
The Solution for the sequential task continual learning track of the 2nd Greater Bay Area International Algorithm Competition
by: Pan, Sishun, et al.
Published: (2024) -
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
by: Huang, Longfei, et al.
Published: (2024) -
The Solution for Language-Enhanced Image New Category Discovery
by: Xu, Haonan, et al.
Published: (2024) -
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
by: Wu, Jie, et al.
Published: (2026) -
The Solution for The PST-KDD-2024 OAG-Challenge
by: Zhong, Shupeng, et al.
Published: (2024)