Exploring the Knowledge Transferred by Response-Based Teacher-Student Distillation

23Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Response-based Knowledge Distillation refers to the technique of supervising the student network with the teacher networks' predictions. The method is motivated by observing that the predicted probabilities reflect the relation among labels, which is the knowledge to be transferred. This paper explores the transferred knowledge from a novel perspective: comparing the knowledge transferred through different teachers. Two intriguing properties are observed. First, higher confidence scores of teachers' predictions lead to better distillation results, and second, teachers' incorrectly predicted training samples should be kept for distillation. We then analyze the phenomenon by studying teachers' decision boundaries, of which some can help the student generalize while some may not. Based on the observations, we further propose an embarrassingly simple distillation framework named Efficient Distillation, which is effective on ImageNet with different teacher-student pairs: When using ResNet34 as the teacher, the student ResNet18 trained from scratch reaches 74.07% Top-1 accuracy within 98 GPU hours (RTX 3090), outperforming current state-of-the-art result (73.19%) by a large margin. Our code is available at https://github.com/lsongx/EffDstl.

Cite

CITATION STYLE

APA

Song, L., Gong, X., Zhou, H., Chen, J., Zhang, Q., Doermann, D., & Yuan, J. (2023). Exploring the Knowledge Transferred by Response-Based Teacher-Student Distillation. In MM 2023 - Proceedings of the 31st ACM International Conference on Multimedia (pp. 2704–2713). Association for Computing Machinery, Inc. https://doi.org/10.1145/3581783.3612162

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free