Watch Your Head: Assembling Projection Heads to Save the Reliability of Federated Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Jinqian, Zhu, Jihua, Zheng, Qinghai, Li, Zhongyu, Tian, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Factor Aggregation: Gauge-Aware Low-Rank Server Representations for Federated LoRA
von: Chen, Jinqian, et al.
Veröffentlicht: (2026)
von: Chen, Jinqian, et al.
Veröffentlicht: (2026)
HFedCKD: Toward Robust Heterogeneous Federated Learning via Data-free Knowledge Distillation and Two-way Contrast
von: Zheng, Yiting, et al.
Veröffentlicht: (2025)
von: Zheng, Yiting, et al.
Veröffentlicht: (2025)
FedPSA: Modeling Behavioral Staleness in Asynchronous Federated Learning
von: Lu, Chaoyi, et al.
Veröffentlicht: (2026)
von: Lu, Chaoyi, et al.
Veröffentlicht: (2026)
AFBS:Buffer Gradient Selection in Semi-asynchronous Federated Learning
von: Lu, Chaoyi, et al.
Veröffentlicht: (2025)
von: Lu, Chaoyi, et al.
Veröffentlicht: (2025)
TPFL: A Trustworthy Personalized Federated Learning Framework via Subjective Logic
von: Chen, Jinqian, et al.
Veröffentlicht: (2024)
von: Chen, Jinqian, et al.
Veröffentlicht: (2024)
Deep Fusion: Capturing Dependencies in Contrastive Learning via Transformer Projection Heads
von: Li, Huanran, et al.
Veröffentlicht: (2024)
von: Li, Huanran, et al.
Veröffentlicht: (2024)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
von: Luo, Cheng, et al.
Veröffentlicht: (2025)
Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning
von: Raje, Arian, et al.
Veröffentlicht: (2025)
von: Raje, Arian, et al.
Veröffentlicht: (2025)
Two Heads Are Better than One: Model-Weight and Latent-Space Analysis for Federated Learning on Non-iid Data against Poisoning Attacks
von: Lyu, Xingyu, et al.
Veröffentlicht: (2025)
von: Lyu, Xingyu, et al.
Veröffentlicht: (2025)
MoH: Multi-Head Attention as Mixture-of-Head Attention
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
von: Liao, Xutao, et al.
Veröffentlicht: (2024)
von: Liao, Xutao, et al.
Veröffentlicht: (2024)
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
von: Tranheden, Wilhelm, et al.
Veröffentlicht: (2026)
von: Tranheden, Wilhelm, et al.
Veröffentlicht: (2026)
Watch Your Steps: Observable and Modular Chains of Thought
von: Cohen, Cassandra A., et al.
Veröffentlicht: (2024)
von: Cohen, Cassandra A., et al.
Veröffentlicht: (2024)
FedAH: Aggregated Head for Personalized Federated Learning
von: Zhou, Pengzhan, et al.
Veröffentlicht: (2024)
von: Zhou, Pengzhan, et al.
Veröffentlicht: (2024)
The Gaussian-Head OFL Family: One-Shot Federated Learning from Client Global Statistics
von: Turazza, Fabio, et al.
Veröffentlicht: (2026)
von: Turazza, Fabio, et al.
Veröffentlicht: (2026)
Fast and Lightweight Backdoor Detection via Head Random Probing
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
von: Yu, Yinbo, et al.
Veröffentlicht: (2026)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
Cultural Binding Heads in Language Models
von: Floro, Avrile, et al.
Veröffentlicht: (2026)
von: Floro, Avrile, et al.
Veröffentlicht: (2026)
STAMP: Spatial-Temporal Adapter with Multi-Head Pooling
von: Shook, Brad, et al.
Veröffentlicht: (2025)
von: Shook, Brad, et al.
Veröffentlicht: (2025)
Multi-Head Spectral-Adaptive Graph Anomaly Detection
von: Cao, Qingyue, et al.
Veröffentlicht: (2025)
von: Cao, Qingyue, et al.
Veröffentlicht: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
von: Chen, Yilong, et al.
Veröffentlicht: (2024)
Singular Vectors of Attention Heads Align with Features
von: Franco, Gabriel, et al.
Veröffentlicht: (2026)
von: Franco, Gabriel, et al.
Veröffentlicht: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
Geometric Analysis of Token Selection in Multi-Head Attention
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2026)
Do Attention Heads Compete or Cooperate during Counting?
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
von: Zsámboki, Pál, et al.
Veröffentlicht: (2025)
Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries
von: Li, Yanhang, et al.
Veröffentlicht: (2026)
von: Li, Yanhang, et al.
Veröffentlicht: (2026)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
von: Williams, Jorge L. Ruiz
Veröffentlicht: (2026)
von: Williams, Jorge L. Ruiz
Veröffentlicht: (2026)
Not All Models Suit Expert Offloading: On Local Routing Consistency of Mixture-of-Expert Models
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
von: Liang, Jingcong, et al.
Veröffentlicht: (2025)
Continual Domain Adversarial Adaptation via Double-Head Discriminators
von: Shen, Yan, et al.
Veröffentlicht: (2024)
von: Shen, Yan, et al.
Veröffentlicht: (2024)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
von: Bali, Karan, et al.
Veröffentlicht: (2026)
von: Bali, Karan, et al.
Veröffentlicht: (2026)
The Strong Lottery Ticket Hypothesis for Multi-Head Attention Mechanisms
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)
von: Otsuka, Hikari, et al.
Veröffentlicht: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Fourier Head: Helping Large Language Models Learn Complex Probability Distributions
von: Gillman, Nate, et al.
Veröffentlicht: (2024)
von: Gillman, Nate, et al.
Veröffentlicht: (2024)
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
von: Wu, Zhoutong, et al.
Veröffentlicht: (2025)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
von: Ghodsi, Ali
Veröffentlicht: (2025)
von: Ghodsi, Ali
Veröffentlicht: (2025)
Prospector Heads: Generalized Feature Attribution for Large Models & Data
von: Machiraju, Gautam, et al.
Veröffentlicht: (2024)
von: Machiraju, Gautam, et al.
Veröffentlicht: (2024)
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
von: Wang, George, et al.
Veröffentlicht: (2024)
von: Wang, George, et al.
Veröffentlicht: (2024)
TransMLA: Multi-Head Latent Attention Is All You Need
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Factor Aggregation: Gauge-Aware Low-Rank Server Representations for Federated LoRA
von: Chen, Jinqian, et al.
Veröffentlicht: (2026) -
HFedCKD: Toward Robust Heterogeneous Federated Learning via Data-free Knowledge Distillation and Two-way Contrast
von: Zheng, Yiting, et al.
Veröffentlicht: (2025) -
FedPSA: Modeling Behavioral Staleness in Asynchronous Federated Learning
von: Lu, Chaoyi, et al.
Veröffentlicht: (2026) -
AFBS:Buffer Gradient Selection in Semi-asynchronous Federated Learning
von: Lu, Chaoyi, et al.
Veröffentlicht: (2025) -
TPFL: A Trustworthy Personalized Federated Learning Framework via Subjective Logic
von: Chen, Jinqian, et al.
Veröffentlicht: (2024)