Abstract
Convolutional Neural Networks (CNNs) surpass human-level performance on visual object recognition, yet their behavior differs from humans in important ways. One prominent example is that CNNs trained on ImageNet exhibit a texture bias, while humans show a strong shape bias. Although CNN shape bias can be increased through e.g., data augmentation or additional training, the source of this discrepancy remains unclear. Developmental research suggests that one factor driving human shape bias is that during early childhood, toddlers tend to fill their field-of-view with close-up objects. We operationalize this close-up as a zoom-in on objects during CNN training, which increases shape bias without additional training or data augmentation. Systematic manipulation of background-object ratios during training reveals a strong inverse correlation with shape bias. Notably, zooming in on objects, thereby more closely emulating child vision, aligns classification accuracy and shape bias between humans and CNNs. Finally, we achieve a near human-like shape bias when using a developmentally-inspired background-object ratio for training and shape bias assessment. These findings demonstrate that a simple adjustment to image datasets — zooming in on objects — can produce human-like shape bias. This suggests that human learning strategies offer a promising avenue for developing human-aligned, efficient, and robust vision CNNs.
Author supplied keywords
Cite
CITATION STYLE
Müller, N., Snoek, C. G. M., Groen, I. I. A., & Scholte, H. S. (2026). Object-zoomed training of convolutional neural networks inspired by toddler development improves shape bias. Connection Science, 38(1). https://doi.org/10.1080/09540091.2026.2694256
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.