A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise
A Lightweight 1D-CNN for Bark and Howl Classification from Raw Audio Waveforms Under Controlled Additive Noise
Abstract
Automatic classification of dog vocalizations can support bioacoustic monitoring and animal welfare, but many systems require spectral or cepstral preprocessing. This study evaluates a lightweight one-dimensional convolutional neural network (1D-CNN) for Bark and Howl classification directly from raw waveforms under controlled additive noise. The dataset comprised 46 Bark and 57 Howl recordings. Audio was converted to mono, resampled to 16 kHz, and standardized to 2.0 s. The network contains 7130 trainable parameters, occupies 27.85 KB with 32-bit weights, and requires 18.35 MFLOPs. The complete five-fold cross-validation procedure was repeated ten times with independently generated run-specific seeds and newly shuffled partitions. Under the no-added-noise condition, mean accuracy was 93.40 ± 2.62%, and macro F1-score was 93.20 ± 2.75%. Performance remained within run-to-run variability between 30 and 5 dB SNR for Gaussian and uniform additive noise, whereas mean accuracy decreased to 79.42% at 0 dB. In the seed-42 reference ablation, removing noise augmentation preserved no-added-noise accuracy but reduced 5 dB accuracy by approximately 20 percentage points. The findings provide preliminary recording-level evidence for efficient Bark and Howl classification under controlled conditions. Generalization to unseen dogs and field recordings remains unverified.
Description
ORCID
Keywords
Bioacoustics, Lightweight Deep Learning, Bark and Howl, Raw Audio Waveform, Dog Vocalization Classification, Repeated Cross-Validation, Pattern Recognition (Psychology), Speech Recognition, Computer Science, One-Dimensional Convolutional Neural Network, Controlled Additive Noise, Waveform, Noise (Video)
Fields of Science
Citation
WoS Q
Scopus Q
Source
Volume
16
Issue
13
Start Page
6819
End Page
6819

