Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training

arXiv cs.AI / 4/28/2026

💬 OpinionDeveloper Stack & InfrastructureModels & Research

共有:

Key Points

The paper addresses inefficiencies in geo-distributed ML training, mainly caused by missing elastic scheduling for multi-regional cloud resources and WAN communication overhead with bandwidth limits and fluctuations.
It proposes “Cloudless-Training,” a PS-based framework that uses a two-layer control/physical training architecture to enable elastic scheduling and WAN-aware communication in a serverless manner.
Cloudless-Training introduces an elastic scheduling strategy that adapts training workflows to heterogeneous cloud resources and to the location/distribution of pre-existing training datasets.
It also contributes two synchronization methods across clouds—ASGD-GA (asynchronous SGD with gradient accumulation) and inter-PS model averaging (MA)—to improve partition coordination while maintaining model correctness.
Implemented with OpenFaaS and evaluated on Tencent Cloud, the approach shows substantial gains, including 9.2%–24.0% training cost reduction and up to 1.7× synchronization/training speedup versus baseline.

Abstract

Geo-distributed ML training can benefit many emerging ML scenarios (e.g., large model training, federated learning) with multi-regional cloud resources and wide area network. However, its efficiency is limited due to 2 challenges. First, efficient elastic scheduling of multi-regional cloud resources is usually missing, affecting resource utilization and performance of training. Second, training communication on WAN is still the main overhead, easily subjected to low bandwidth and high fluctuations of WAN. In this paper, we propose a framework, Cloudless-Training, to realize efficient PS-based geo-distributed ML training in 3 aspects. First, it uses a two-layer architecture with control and physical training planes to support elastic scheduling and communication for multi-regional clouds in a serverless maner.Second, it provides an elastic scheduling strategy that can deploy training workflows adaptively according to the heterogeneity of available cloud resources and distribution of pre-existing training datasets. Third, it provides 2 new synchronization strategies for training partitions among clouds, including asynchronous SGD with gradient accumulation (ASGD-GA) and inter-PS model averaging (MA). It is implemented with OpenFaaS and evaluated on Tencent Cloud. Experiments show that Cloudless-Training can support general ML training in a geo-distributed way, greatly improve resource utilization (e.g., 9.2%-24.0% training cost reduction) and synchronization efficiency (e.g., 1.7x training speedup over baseline at most) with model correctness guarantees.

How to Build Traceable and Evaluated LLM Workflows Using Promptflow, Prompty, and OpenAI

MarkTechPost

An improvement of the convergence proof of the ADAM-Optimizer

Dev.to

Claude Code 会话历史在哪里？如何找回你的 AI 编程对话记录

Dev.to

We built an AI that runs an entire business autonomously. Not a demo. Not a prototype. Actually running. YC-backed, here's what we learned.

Reddit r/artificial

langchain-tests==1.1.7

LangChain Releases

Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training

Key Points

Abstract

Related Articles

How to Build Traceable and Evaluated LLM Workflows Using Promptflow, Prompty, and OpenAI

An improvement of the convergence proof of the ADAM-Optimizer

Claude Code 会话历史在哪里？如何找回你的 AI 编程对话记录

We built an AI that runs an entire business autonomously. Not a demo. Not a prototype. Actually running. YC-backed, here's what we learned.

langchain-tests==1.1.7

関連おすすめサービス

Notta搭載AI議事録イヤホン ZENCHORD1

AI搭載ボイスレコーダー Plaud

画像高画質化AIツール Aiarty Image Enhancer