Abstract:Adaptive moment estimation (Adam), as a Stochastic Gradient Descent (SGD) variant, has gained widespread popularity in federated learning (FL) due to its fast convergence. However, federated Adam (FedAdam) algorithms suffer from a threefold increase in uplink communication overhead compared to federated SGD (FedSGD) algorithms, which arises from the necessity to transmit both local model updates and first and second moment estimates from distributed devices to the centralized server for aggregation. Driven by this issue, we propose a novel sparse FedAdam algorithm called FedAdam-SSM, wherein distributed devices sparsify the updates of local model parameters and moment estimates and subsequently upload the sparse representations to the centralized server. To further reduce the communication overhead, the updates of local model parameters and moment estimates incorporate a shared sparse mask (SSM) into the sparsification process, eliminating the need for three separate sparse masks. Theoretically, we develop an upper bound on the divergence between the local model trained by FedAdam-SSM and the desired model trained by centralized Adam, which is related to sparsification error and imbalanced data distribution. By minimizing the divergence bound between the model trained by FedAdam-SSM and centralized Adam, we optimize the SSM to mitigate the learning performance degradation caused by sparsification error. Additionally, we provide convergence bounds for FedAdam-SSM in both convex and non-convex objective function settings, and investigate the impact of local epoch, learning rate and sparsification ratio on the convergence rate of FedAdam-SSM. Experimental results show that FedAdam-SSM outperforms baselines in terms of convergence rate (over 1.1$\times$ faster than the sparse FedAdam baselines) and test accuracy (over 14.5\% ahead of the quantized FedAdam baselines).

Optimizing Parameter Mixing under Constrained Communications in Parallel Federated Learning

Projected Federated Averaging with Heterogeneous Differential Privacy.

FedPSE: Personalized Sparsification with Element-wise Aggregation for Federated Learning

FedPD: A Federated Learning Framework with Optimal Rates and Adaptivity to Non-IID Data.

FedPA: An adaptively partial model aggregation strategy in Federated Learning

Preconditioned Federated Learning

Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization

Conquering the Communication Constraints to Enable Large Pre-Trained Models in Federated Learning

Federated Learning Hyper-Parameter Tuning From A System Perspective

Computation and Communication Efficient Federated Learning With Adaptive Model Pruning

FedMCP: Parameter-Efficient Federated Learning with Model-Contrastive Personalization

SA-FedLora: Adaptive Parameter Allocation for Efficient Federated Learning with LoRA Tuning

Accelerating Hybrid Federated Learning Convergence under Partial Participation

Efficient Wireless Federated Learning with Partial Model Aggregation

FedPD: A Federated Learning Framework With Adaptivity to Non-IID Data

FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization

FedLPS: Heterogeneous Federated Learning for Multiple Tasks with Local Parameter Sharing

pFedAFM: Adaptive Feature Mixture for Batch-Level Personalization in Heterogeneous Federated Learning

FedMix: Boosting with Data Mixture for Vertical Federated Learning

Towards Communication-efficient Federated Learning via Sparse and Aligned Adaptive Optimization

No One Idles: Efficient Heterogeneous Federated Learning with Parallel Edge and Server Computation.