Saved in:
Bibliographic Details
Main Authors: Wu, Haoyuan, Ming, Rui, Zheng, Haisheng, He, Zhuolun, Yu, Bei
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.12502
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915157660139520
author Wu, Haoyuan
Ming, Rui
Zheng, Haisheng
He, Zhuolun
Yu, Bei
author_facet Wu, Haoyuan
Ming, Rui
Zheng, Haisheng
He, Zhuolun
Yu, Bei
contents Large language models (LLMs) have shown significant promise in question-answering (QA) tasks, particularly in retrieval-augmented generation (RAG) scenarios and long-context applications. However, their performance is hindered by noisy reference documents, which often distract from essential information. Despite fine-tuning efforts, Transformer-based architectures struggle to prioritize relevant content. This is evidenced by their tendency to allocate disproportionate attention to irrelevant or later-positioned documents. Recent work proposes the differential attention mechanism to address this issue, but this mechanism is limited by an unsuitable common-mode rejection ratio (CMRR) and high computational costs. Inspired by the operational amplifier (OpAmp), we propose the OpAmp adaptation to address these challenges, which is implemented with adapters efficiently. By integrating the adapter into pre-trained Transformer blocks, our approach enhances focus on the golden context without costly training from scratch. Empirical evaluations on noisy-context benchmarks reveal that our Qwen2.5-OpAmp-72B model, trained with our OpAmp adaptation, surpasses the performance of state-of-the-art LLMs, including DeepSeek-V3 and GPT-4o.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
Wu, Haoyuan
Ming, Rui
Zheng, Haisheng
He, Zhuolun
Yu, Bei
Computation and Language
Large language models (LLMs) have shown significant promise in question-answering (QA) tasks, particularly in retrieval-augmented generation (RAG) scenarios and long-context applications. However, their performance is hindered by noisy reference documents, which often distract from essential information. Despite fine-tuning efforts, Transformer-based architectures struggle to prioritize relevant content. This is evidenced by their tendency to allocate disproportionate attention to irrelevant or later-positioned documents. Recent work proposes the differential attention mechanism to address this issue, but this mechanism is limited by an unsuitable common-mode rejection ratio (CMRR) and high computational costs. Inspired by the operational amplifier (OpAmp), we propose the OpAmp adaptation to address these challenges, which is implemented with adapters efficiently. By integrating the adapter into pre-trained Transformer blocks, our approach enhances focus on the golden context without costly training from scratch. Empirical evaluations on noisy-context benchmarks reveal that our Qwen2.5-OpAmp-72B model, trained with our OpAmp adaptation, surpasses the performance of state-of-the-art LLMs, including DeepSeek-V3 and GPT-4o.
title Efficient OpAmp Adaptation for Zoom Attention to Golden Contexts
topic Computation and Language
url https://arxiv.org/abs/2502.12502