BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ghaddar, Abbas, Kobyzev, Ivan, Chen, Boxing, Cui, Yufei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
by: Kobyzev, Ivan, et al.
Published: (2025)
by: Kobyzev, Ivan, et al.
Published: (2025)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
by: Huang, Chenyang, et al.
Published: (2024)
by: Huang, Chenyang, et al.
Published: (2024)
ReGLA: Refining Gated Linear Attention
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
by: Metel, Michael R., et al.
Published: (2026)
by: Metel, Michael R., et al.
Published: (2026)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
BinarySelect to Improve Accessibility of Black-Box Attack Research
by: Ghosh, Shatarupa, et al.
Published: (2024)
by: Ghosh, Shatarupa, et al.
Published: (2024)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
by: Wang, Suyuchen, et al.
Published: (2024)
by: Wang, Suyuchen, et al.
Published: (2024)
Hyperparameter Optimization for Large Language Model Instruction-Tuning
by: Tribes, Christophe, et al.
Published: (2023)
by: Tribes, Christophe, et al.
Published: (2023)
Resona: Improving Context Copying in Linear Recurrence Models with Retrieval
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
by: Lu, Peng, et al.
Published: (2023)
by: Lu, Peng, et al.
Published: (2023)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
by: Heuillet, Maxime, et al.
Published: (2025)
by: Heuillet, Maxime, et al.
Published: (2025)
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
by: Zhou, Ziyu, et al.
Published: (2025)
by: Zhou, Ziyu, et al.
Published: (2025)
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs
by: Vazhentsev, Artem, et al.
Published: (2025)
by: Vazhentsev, Artem, et al.
Published: (2025)
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
by: Lin, Xihui, et al.
Published: (2024)
by: Lin, Xihui, et al.
Published: (2024)
A Survey of Calibration Process for Black-Box LLMs
by: Xie, Liangru, et al.
Published: (2024)
by: Xie, Liangru, et al.
Published: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
Knocking-Heads Attention
by: Zhou, Zhanchao, et al.
Published: (2025)
by: Zhou, Zhanchao, et al.
Published: (2025)
Debiasing LLMs by Masking Unfairness-Driving Attention Heads
by: Han, Tingxu, et al.
Published: (2025)
by: Han, Tingxu, et al.
Published: (2025)
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
by: Lu, Yi, et al.
Published: (2024)
by: Lu, Yi, et al.
Published: (2024)
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
by: Li, Zhuo, et al.
Published: (2024)
by: Li, Zhuo, et al.
Published: (2024)
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
by: Hua, Kai, et al.
Published: (2025)
by: Hua, Kai, et al.
Published: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
Which Attention Heads Matter for In-Context Learning?
by: Yin, Kayo, et al.
Published: (2025)
by: Yin, Kayo, et al.
Published: (2025)
Why Does the Effective Context Length of LLMs Fall Short?
by: An, Chenxin, et al.
Published: (2024)
by: An, Chenxin, et al.
Published: (2024)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
by: Zhong, Meizhi, et al.
Published: (2024)
by: Zhong, Meizhi, et al.
Published: (2024)
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
by: Zhao, Fufangchen, et al.
Published: (2024)
by: Zhao, Fufangchen, et al.
Published: (2024)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
AdpQ: A Zero-shot Calibration Free Adaptive Post Training Quantization Method for LLMs
by: Ghaffari, Alireza, et al.
Published: (2024)
by: Ghaffari, Alireza, et al.
Published: (2024)
Topic Modelling Black Box Optimization
by: Akramov, Roman, et al.
Published: (2025)
by: Akramov, Roman, et al.
Published: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
by: Donhauser, Konstantin, et al.
Published: (2025)
by: Donhauser, Konstantin, et al.
Published: (2025)
EWEK-QA: Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems
by: Dehghan, Mohammad, et al.
Published: (2024)
by: Dehghan, Mohammad, et al.
Published: (2024)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
Similar Items
-
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
by: Kobyzev, Ivan, et al.
Published: (2025) -
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
by: Huang, Chenyang, et al.
Published: (2024) -
ReGLA: Refining Gated Linear Attention
by: Lu, Peng, et al.
Published: (2025) -
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024) -
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024)