An analysis of audio classification techniques using deep learning architectures

Citation

Abstract

Failure to classify audio data with high efficiency causes major setbacks in audio processing, voice recognition and noise cancellation. In order to find the best possible neural network models for audio classi cation, this paper shows the steps in the experiments done on our newly designed CF Model and CFClean Model in both CNN and RNN, and compares the results with some existing models such as DCNN and Piczak-CNN. To get a clear view on the consistency of the results, three di erent datasets have been experimented on: UrbanSound8k, FSDKaggle2018 and ESC-50. This paper also sheds light on which dataset performs best in terms of train and test accuracy and loss percentage. Moreover, this paper also dives deep into the reasons behind particular models and datasets performing better than the others. Finally, this paper shows what influence envelope function, normalization, segmentation, regularization techniques and dropout layers have in the overall progress.

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 26-27).
This thesis is submitted in partial fulfillment of the requirements for the degree of Bachelor of Science in Computer Science, 2020.

Publisher Link

Type

Thesis