Uncovering the Role of Initial Saliency in U-Shaped Attention Bias: Scaling Initial Token Weight for Enhanced Long-Text Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Qiang, Zewen, Zhao, Sendong, Wang, Haochun, Qin, Bing, Liu, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis
by: Wang, Haochun, et al.
Published: (2024)
by: Wang, Haochun, et al.
Published: (2024)
LLMs May Perform MCQA by Selecting the Least Incorrect Option
by: Wang, Haochun, et al.
Published: (2024)
by: Wang, Haochun, et al.
Published: (2024)
Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection
by: Wang, Haochun, et al.
Published: (2026)
by: Wang, Haochun, et al.
Published: (2026)
MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text Prompts
by: Guo, Haoqiang, et al.
Published: (2024)
by: Guo, Haoqiang, et al.
Published: (2024)
AS-ES Learning: Towards Efficient CoT Learning in Small Models
by: Xi, Nuwa, et al.
Published: (2024)
by: Xi, Nuwa, et al.
Published: (2024)
Beyond Frameworks: Unpacking Collaboration Strategies in Multi-Agent Systems
by: Wang, Haochun, et al.
Published: (2025)
by: Wang, Haochun, et al.
Published: (2025)
META-RAG: Meta-Analysis-Inspired Evidence-Re-Ranking Method for Retrieval-Augmented Generation in Evidence-Based Medicine
by: Sun, Mengzhou, et al.
Published: (2025)
by: Sun, Mengzhou, et al.
Published: (2025)
Manifold-based Verbalizer Space Re-embedding for Tuning-free Prompt-based Classification
by: Wang, Haochun, et al.
Published: (2023)
by: Wang, Haochun, et al.
Published: (2023)
MolFusion: Multimodal Fusion Learning for Molecular Representations via Multi-granularity Views
by: Cai, Muzhen, et al.
Published: (2024)
by: Cai, Muzhen, et al.
Published: (2024)
ArcAligner: Adaptive Recursive Aligner for Compressed Context Embeddings in RAG
by: Li, Jianbo, et al.
Published: (2026)
by: Li, Jianbo, et al.
Published: (2026)
M-Eval: A Heterogeneity-Based Framework for Multi-evidence Validation in Medical RAG Systems
by: Sun, Mengzhou, et al.
Published: (2025)
by: Sun, Mengzhou, et al.
Published: (2025)
Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
by: Wang, Haochun, et al.
Published: (2023)
by: Wang, Haochun, et al.
Published: (2023)
CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery
by: Chen, Yuhan, et al.
Published: (2023)
by: Chen, Yuhan, et al.
Published: (2023)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
by: Xiao, Boyu, et al.
Published: (2026)
by: Xiao, Boyu, et al.
Published: (2026)
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
by: Du, Yanrui, et al.
Published: (2025)
by: Du, Yanrui, et al.
Published: (2025)
Orchestrating Intelligence: Confidence-Aware Routing for Efficient Multi-Agent Collaboration across Multi-Scale Models
by: Wang, Jingbo, et al.
Published: (2026)
by: Wang, Jingbo, et al.
Published: (2026)
An Analysis for Reasoning Bias of Language Models with Small Initialization
by: Yao, Junjie, et al.
Published: (2025)
by: Yao, Junjie, et al.
Published: (2025)
From Latent Signals to Reflection Behavior: Tracing Meta-Cognitive Activation Trajectory in R1-Style LLMs
by: Du, Yanrui, et al.
Published: (2026)
by: Du, Yanrui, et al.
Published: (2026)
SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation
by: Wang, Haochun, et al.
Published: (2026)
by: Wang, Haochun, et al.
Published: (2026)
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
by: Du, Yanrui, et al.
Published: (2023)
by: Du, Yanrui, et al.
Published: (2023)
GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis
by: Jiang, Yi, et al.
Published: (2025)
by: Jiang, Yi, et al.
Published: (2025)
GSEM: Graph-based Self-Evolving Memory for Experience Augmented Clinical Reasoning
by: Han, Xiao, et al.
Published: (2026)
by: Han, Xiao, et al.
Published: (2026)
Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs
by: Chen, Hanting, et al.
Published: (2025)
by: Chen, Hanting, et al.
Published: (2025)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
by: Zeris, Athanasios
Published: (2026)
by: Zeris, Athanasios
Published: (2026)
Re-Initialization Token Learning for Tool-Augmented Large Language Models
by: Li, Chenghao, et al.
Published: (2025)
by: Li, Chenghao, et al.
Published: (2025)
Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models
by: Kim, Jiyeon, et al.
Published: (2026)
by: Kim, Jiyeon, et al.
Published: (2026)
Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
by: Du, Yanrui, et al.
Published: (2024)
by: Du, Yanrui, et al.
Published: (2024)
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
by: Nagaraj, Manish, et al.
Published: (2025)
by: Nagaraj, Manish, et al.
Published: (2025)
PICOs-RAG: PICO-supported Query Rewriting for Retrieval-Augmented Generation in Evidence-Based Medicine
by: Sun, Mengzhou, et al.
Published: (2025)
by: Sun, Mengzhou, et al.
Published: (2025)
MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability
by: Du, Yanrui, et al.
Published: (2024)
by: Du, Yanrui, et al.
Published: (2024)
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
by: Du, Yanrui, et al.
Published: (2025)
by: Du, Yanrui, et al.
Published: (2025)
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
by: Du, Yanrui, et al.
Published: (2023)
by: Du, Yanrui, et al.
Published: (2023)
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
by: Chen, Daiwei, et al.
Published: (2026)
by: Chen, Daiwei, et al.
Published: (2026)
On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty Perspective
by: Guo, Yikai, et al.
Published: (2026)
by: Guo, Yikai, et al.
Published: (2026)
Multilingual Test-Time Scaling via Initial Thought Transfer
by: Bajpai, Prasoon, et al.
Published: (2025)
by: Bajpai, Prasoon, et al.
Published: (2025)
Merging Text Transformer Models from Different Initializations
by: Verma, Neha, et al.
Published: (2024)
by: Verma, Neha, et al.
Published: (2024)
Token Weighting for Long-Range Language Modeling
by: Helm, Falko, et al.
Published: (2025)
by: Helm, Falko, et al.
Published: (2025)
Problematic Tokens: Tokenizer Bias in Large Language Models
by: Yang, Jin, et al.
Published: (2024)
by: Yang, Jin, et al.
Published: (2024)
Similar Items
-
Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis
by: Wang, Haochun, et al.
Published: (2024) -
LLMs May Perform MCQA by Selecting the Least Incorrect Option
by: Wang, Haochun, et al.
Published: (2024) -
Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection
by: Wang, Haochun, et al.
Published: (2026) -
MolTailor: Tailoring Chemical Molecular Representation to Specific Tasks via Text Prompts
by: Guo, Haoqiang, et al.
Published: (2024) -
AS-ES Learning: Towards Efficient CoT Learning in Small Models
by: Xi, Nuwa, et al.
Published: (2024)