Abstract
Inference-as-a-service (IAAS) has been recently launched by cloud service providers to support on-demand AI applications. Many natural language processing (NLP) services are based on the Transformer Sequence Transduction model. However, the inference process of the Transformer model consumes a significant amount of energy due to the large model size (e.g., billions of parameters) and tremendous computations. How to reduce the energy consumption of IAAS without violating the service-level agreement (SLA) becomes a practical challenge for service providers. In this work, we conduct a comprehensive study on the inference performance and energy efficiency of a Transformer model trained for the language translation service. First, we empirically characterize some essential performance metrics, including latency, throughput, and energy consumption on three different GPUs with diversified workload configurations. The detailed workload separation facilitates a thorough and deep understanding of the inference process of the Transformer model. Second, we provide an energy consumption model for the Transformer based on the observed data. Finally, we propose the Aligned scheduling scheme that optimizes throughput and energy efficiency with up to 2.86× and 2.73× improvement at the cost of 40% average latency loss. Our findings provide a full scope of Transformer inference, and suggest that the workload balancing and scheduling have great potentials to offer energy-efficient Transformer inference services.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - IEEE Congress on Cybermatics |
| Subtitle of host publication | 2020 IEEE International Conferences on Internet of Things, iThings 2020, IEEE Green Computing and Communications, GreenCom 2020, IEEE Cyber, Physical and Social Computing, CPSCom 2020 and IEEE Smart Data, SmartData 2020 |
| Publisher | IEEE |
| Pages | 323-331 |
| Number of pages | 9 |
| ISBN (Electronic) | 9781728176475 |
| DOIs | |
| Publication status | Published - Nov 2020 |
| Event | 2020 IEEE Congress on Cybermatics: 13th IEEE International Conferences on Internet of Things, iThings 2020, 16th IEEE International Conference on Green Computing and Communications, GreenCom 2020, 13th IEEE International Conference on Cyber, Physical and Social Computing, CPSCom 2020 and 6th IEEE International Conference on Smart Data, SmartData 2020 - Rhodes Island, Greece Duration: 2 Nov 2020 → 6 Nov 2020 |
Publication series
| Name | Proceedings - IEEE Congress on Cybermatics: 2020 IEEE International Conferences on Internet of Things, iThings 2020, IEEE Green Computing and Communications, GreenCom 2020, IEEE Cyber, Physical and Social Computing, CPSCom 2020 and IEEE Smart Data, SmartData 2020 |
|---|
Conference
| Conference | 2020 IEEE Congress on Cybermatics: 13th IEEE International Conferences on Internet of Things, iThings 2020, 16th IEEE International Conference on Green Computing and Communications, GreenCom 2020, 13th IEEE International Conference on Cyber, Physical and Social Computing, CPSCom 2020 and 6th IEEE International Conference on Smart Data, SmartData 2020 |
|---|---|
| Country/Territory | Greece |
| City | Rhodes Island |
| Period | 2/11/20 → 6/11/20 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 7 Affordable and Clean Energy
User-Defined Keywords
- Batch Inference
- Cloud Service
- Energy Efficiency
- Graphics Processing Units
- Inference Scheduling
- Transformer Model
Fingerprint
Dive into the research topics of 'Energy-efficient Inference Service of Transformer-based Deep Learning Models on GPUs'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver