Multi-objective autotuning of mobile nets across the full software/hardware stack

9Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We present a customizable Collective Knowledge workflow to study the execution time vs. accuracy trade-offs for the MobileNets CNN family. We use this workflow to evaluate MobileNets on Arm Cortex CPUs using TensorFlow and Arm Mali GPUs using several versions of the Arm Compute Library. Our optimizations for the Arm Bifrost GPU architecture reduce the execution time by 2-3 times, while lying on a Pareto-optimal frontier. We also highlight the challenge of maintaining the accuracy when deploying CNN models across diverse platforms. We make all the workflow components (models, programs, scripts, etc.) publicly available to encourage further exploration by the community.

Cite

CITATION STYLE

APA

Lokhmotov, A., Vella, F., Chunosov, N., & Fursin, G. (2018). Multi-objective autotuning of mobile nets across the full software/hardware stack. In Proceedings of the 1st Reproducible Quality-Efficient Systems Tournament on Co-Designing Pareto-Efficient Deep Learning, ReQuEST 2018 - Co-located with ACM ASPLOS 2018. Association for Computing Machinery. https://doi.org/10.1145/3229762.3229767

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free