Multi-encoder attention-based architectures for sound recognition with partial visual assistance

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be exploited to improve the outputs of models designed for associated tasks. Frequently, however, not all contents are available for all samples of such a collection: For example, the original material may have been removed from the source platform at some point, and therefore, non-auditory features can no longer be acquired. We demonstrate that a multi-encoder framework can be employed to deal with this issue by applying this method to attention-based deep learning systems, which are currently part of the state of the art in the domain of sound recognition. More specifically, we show that the proposed model extension can successfully be utilized to incorporate partially available visual information into the operational procedures of such networks, which normally only use auditory features during training and inference. Experimentally, we verify that the considered approach leads to improved predictions in a number of evaluation scenarios pertaining to audio tagging and sound event detection. Additionally, we scrutinize some properties and limitations of the presented technique.

References Powered by Scopus

Deep residual learning for image recognition

174051Citations
N/AReaders
Get full text

ImageNet Large Scale Visual Recognition Challenge

30411Citations
N/AReaders
Get full text

Audio Set: An ontology and human-labeled dataset for audio events

2291Citations
N/AReaders
Get full text

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Boes, W., & Van hamme, H. (2022). Multi-encoder attention-based architectures for sound recognition with partial visual assistance. Eurasip Journal on Audio, Speech, and Music Processing, 2022(1). https://doi.org/10.1186/s13636-022-00252-9

Readers over time

‘22‘23‘2400.751.52.253

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 2

100%

Readers' Discipline

Tooltip

Computer Science 1

50%

Linguistics 1

50%

Save time finding and organizing research with Mendeley

Sign up for free
0