• No se han encontrado resultados

Humberto Ramírez Gómez

The FFT of the power spectrum captures the magnitude of both the frequency at which an insect is stridulating, and of the frequency of slower, repeating patterns. For example, the dark bush-cricket, common in the New Forest, chirps at a frequency of approximately 1 Hz, but each chirp is composed of three phrases, repeating at about 40 Hz, and in a clean recording each of these reveals a further 600 Hz amplitude mod- ulation (for a clear example refer to Figure2.8ain Chapter2). Averaging over time or taking the maximum value of each frequency bin would discard this information. Fig- ure5.2 shows an artificial signal that exemplifies the power of this feature extraction method. The sample signal is composed of two sine waves, one at 8 kHz and one at 16 kHz, silenced 6 and 10 times per second respectively. The FFT of the mixed signal shows, like the clean carrier wave, the two peak frequencies, and the spectrogram also highlights the alternating chirps. However, these four frequencies are perfectly sum- marised only in the FFT of the power spectrum, which in Figure5.2 is capped to the first 50 Hz. On the y-axis, the frequency domain is represented as in the spectrogram, while the x-axis shows the frequency of the repeated patterns (6 and 10 Hz). Finally Figure5.2shows the same feature, but represented with one line per frequency band, capped at 300 and 30 Hz respectively.

This artificial signal exemplifies the effectiveness of the FFT of the power spectrum. However, a question remains as to how to represent this feature compactly while re- maining maximally general about the features extracted—the method should capture, for example, all the 1, 40 and 600 Hz modulations of the dark bush-cricket’s call, and equally whatever other insect is given to the classifier. To do so, this work proposes to uniformly resample the FFT bins in the log-frequency space. This is motivated by the fact that determining the probability with which a bin contains the feature is

0 5000 10000 15000 20000 Frequency (Hz) 0 20000 40000 60000 80000 100000

120000 FFT of the signal (entire)

mixed sinewave 15600 15700 15800 15900 16000 16100 16200 16300 16400 Frequency (Hz) 0 20000 40000 60000 80000 100000

120000 FFT of the signal (zoomed at peak ±1%)

mixed sinewave 0 50 100 150 200 250 300 Frequency of chirps (Hz) 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7

FFT of hertz spectrum (lines, capped to 300 Hz)

0 5 10 15 20 25 30 Frequency of chirps (Hz) 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7

FFT of hertz spectrum (lines, capped to 30 Hz)

0 2000 4000 6000 8000 10000 Time (µsec) 0 5 10 15 20 Fr equency (kHz) Hertz spectrogram 0 10 20 30 40 50 Frequency of chirps (Hz) 0 5 10 15 20 Fr equency (kHz)

FFT of hertz spectrum – capped to 50 Hz

Figure 5.2: Modulation coefficients of a sample sine wave. The signal is composed of 5 seconds of a 8 kHz and 5 seconds of a 16 kHz sine waves, repeating at 6 and 10 Hz respectively. The FFT shows a clear peak at the two carrier frequencies, but only

equivalent to trying to determine an uninformative prior in Bayesian statistics. Max- imum entropy theory says that the correct uninformative prior for a scale parameter, such as a time or a length, is given by a probability distribution proportional to 1/x (Jaynes, 2003). Therefore, since it cannot be determined a priori whether the repeti- tion of phrases happens at say, 1, 40 or 600 Hz, resampling the FFT bins uniformly in the log-frequency space in the range provided by the power spectrum and the FFT discards minimal information, independently from the scale.

A few parameters determine the final output feature. Firstly, the maximum frequency fmax that the modulation coefficients can detect is determined by the window size w of the STFT and the sampling rate fsof the recording, such that:

fmax= 2fs

·w

In the example above, with sampling rate of 44,100 Hz and a window size of 256, fmax is only ≈ 86 Hz, which is insufficient for many of the insects considered. Reducing the window size degrades the accuracy of the spectrum, but tests have shown that a window size of 64 samples, which gives≈ 344 Hz output, performs well for the data sets analysed by this research (see Section5.4below). A second parameter is the lowest frequency of interest fmin. This is limited by the length of the recording considered, as the lowest component a recording will contain is inversely proportional to its length s, such that:

fmin = 21

·s

Therefore, a 1-second recording will not contain any component below 0.5 Hz and a 30-second recording will not contain anything below 1/60 Hz. However, since this value must remain consistent across the feature set given to the classifier, the low- est useful value will be selected, albeit being wasteful in sampling space for shorter recordings. Finally, the number of modulation coefficients for the log-frequency space can be selected arbitrarily, with the lowest number being maximally efficient and the entire range of bins in the input being maximally accurate. Empirical tests show that 48 is a good compromise on a modern desktop computer. Figure 5.3 shows the log-frequency modulation coefficients for a sample call of New Forest cicada, dark bush-cricket and Roesel’s bush-cricket side-by-side. The three insects are quite similar in this feature space, though clear differences can still be seen.

In conclusion, the combination of these feature extraction and aggregation methods is used in the system here proposed. The power spectrum is used either raw or logged, cosine-transformed and translated onto the mel scale, as well as each permutation of these transformations. It is then summarised by mean and standard deviation over time, maximum value over time, or with the modulation coefficients sampled on the log-frequency scale describe thus far, leading to a total 24 different feature sets. For each of these, the entire set of recordings is classified both with a decision tree and with a random forest classifier.