INVITED: Accelerator Design for Deep Learning Training: Extended Abstract: Invited

Ankur Agrawal; Chia-Yu Chen; Jungwook Choi; Kailash Gopalakrishnan; Jinwook Oh; Sunil Shukla; Vijayalakshmi Srinivasan; Swagath Venkataramani; Wei Zhang

doi:10.1145/3061639.3072944

DAC 2017

Conference paper

18 Jun 2017

INVITED: Accelerator Design for Deep Learning Training: Extended Abstract: Invited

View publication

Abstract

Deep Neural Networks (DNNs) have emerged as a powerful and versatile set of techniques showing successes on challenging artificial intelligence (AI) problems. Applications in domains such as image/video processing, autonomous cars, natural language processing, speech synthesis and recognition, genomics and many others have embraced deep learning as the foundation. DNNs achieve superior accuracy for these applications with high computational complexity using very large models which require 100s of MBs of data storage, exaops of computation and high bandwidth for data movement. In spite of these impressive advances, it still takes days to weeks to train state of the art Deep Networks on large datasets-which directly limits the pace of innovation and adoption. In this paper, we present a multi-pronged approach to address the challenges in meeting both the throughput and the energy efficiency goals for DNN training.

Paper