Optimizing Network Performance in Distributed Machine Learning

Luo Mai; Chuntao Hong; Paolo Costa

help promote

HotCloud '16 button

USENIX Conference Policies

Optimizing Network Performance in Distributed Machine Learning

Luo Mai, Imperial College London; Chuntao Hong and Paolo Costa, Microsoft Research

To cope with the ever growing availability of training data, there have been several proposals to scale machine learning computation beyond a single server and distribute it across a cluster. While this enables reducing the training time, the observed speed up is often limited by network bottlenecks.

To address this, we design MLNET, a host-based communication layer that aims to improve the network performance of distributed machine learning systems. This is achieved through a combination of traffic reduction techniques (to diminish network load in the core and at the edges) and traffic management (to reduce average training time). A key feature of MLNET, is its compatibility with existing hardware and software infrastructure so it can be immediately deployed.

We describe the main techniques underpinning MLNET, and show through simulation that the overall training time can be reduced by up to 78%. While preliminary, our results indicate the critical role played by the network and the benefits of introducing a new communication layer to increase the performance of distributed machine learning systems.

Luo Mai, Imperial College London

Chuntao Hong, Microsoft Research

Paolo Costa, Microsoft Research

Open Access Media

USENIX is committed to Open Access to the research presented at our events. Papers and proceedings are freely available to everyone once the event begins. Any video, audio, and/or slides that are posted after the event are also free and open to everyone. Support USENIX and our commitment to Open Access.

BibTeX

@inproceedings {190634,
author = {Luo Mai and Chuntao Hong and Paolo Costa},
title = {Optimizing Network Performance in Distributed Machine Learning},
booktitle = {7th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 15)},
year = {2015},
address = {Santa Clara, CA},
url = {https://www.usenix.org/conference/hotcloud15/workshop-program/presentation/mai},
publisher = {USENIX Association},
month = jul
}

Download

Mai PDF

View the slides

help promote

USENIX Conference Policies

Optimizing Network Performance in Distributed Machine Learning

Luo Mai, Imperial College London

Chuntao Hong, Microsoft Research

Paolo Costa, Microsoft Research

Open Access Media

Silver Sponsors

Bronze Sponsors

Media Sponsors & Industry Partners

sponsors

help promote

USENIX Conference Policies

Optimizing Network Performance in Distributed Machine Learning

Luo Mai, Imperial College London

Chuntao Hong, Microsoft Research

Paolo Costa, Microsoft Research

Open Access Media

Silver Sponsors

Bronze Sponsors

Media Sponsors & Industry Partners