Communication-Efficient and Robust Multi-Modal Federated Learning via Latent-Space Consensus

arXiv cs.LG / 3/20/2026

📰 NewsModels & Research

共有:

Key Points

Introduces CoMFed, a communication-efficient multi-modal federated learning framework that uses learnable projection matrices to create compressed latent representations.
A latent-space regularizer aligns representations across clients to improve cross-modal consistency and robustness to outliers.
The approach addresses heterogeneity in modalities and model architectures while preserving privacy and reducing communication overhead.
Experimental results on human activity recognition benchmarks show competitive accuracy with minimal overhead.

Abstract

Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, but applying FL to multi-modal settings introduces significant challenges. Clients typically possess heterogeneous modalities and model architectures, making it difficult to align feature spaces efficiently while preserving privacy and minimizing communication costs. To address this, we introduce CoMFed, a Communication-Efficient Multi-Modal Federated Learning framework that uses learnable projection matrices to generate compressed latent representations. A latent-space regularizer aligns these representations across clients, improving cross-modal consistency and robustness to outliers. Experiments on human activity recognition benchmarks show that CoMFed achieves competitive accuracy with minimal overhead.

[R] Combining Identity Anchors + Permission Hierarchies achieves 100% refusal in abliterated LLMs — system prompt only, no fine-tuning

Reddit r/MachineLearning

[P] Vibecoded on a home PC: building a ~2700 Elo browser-playable neural chess engine with a Karpathy-inspired AI-assisted research loop

Reddit r/MachineLearning

Meet DuckLLM 1.0 My First Model!

Reddit r/LocalLLaMA

Since FastFlowLM added support for Linux, I decided to benchmark all the models they support, here are some results

Reddit r/LocalLLaMA

What measure do I use to compare nested models and non nested models in high dimensional survival analysis [D]

Reddit r/MachineLearning

Communication-Efficient and Robust Multi-Modal Federated Learning via Latent-Space Consensus

Key Points

Abstract

Related Articles

[R] Combining Identity Anchors + Permission Hierarchies achieves 100% refusal in abliterated LLMs — system prompt only, no fine-tuning

[P] Vibecoded on a home PC: building a ~2700 Elo browser-playable neural chess engine with a Karpathy-inspired AI-assisted research loop

Meet DuckLLM 1.0 My First Model!

Since FastFlowLM added support for Linux, I decided to benchmark all the models they support, here are some results

What measure do I use to compare nested models and non nested models in high dimensional survival analysis [D]

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer