Comparing Machine Learning Models for Audio Classification of Engines
Evaluating effectiveness and training time of neural networks, SVMs, and other models for engine sound classification
Introduction
Predictive maintenance is a critical component of modern industrial operations, enabling early detection of equipment faults and reducing costly downtime. Audio analysis is a powerful, non-invasive technique for condition monitoring because the sounds produced by a machine can reveal subtle changes in its operational state. This report investigates machine learning models for classifying electric motor status from audio signatures.
The objective is to distinguish three motor states: Good (normal operation), Broken (faulty), and Heavy Load. The analysis uses the IDMT-ISA-Electric-Engine dataset, which contains recordings of brushless DC motors under these conditions.
Data Exploration
The dataset, released by the Fraunhofer Institute for Digital Media Technology, contains approximately 42 minutes of mono WAV files recorded at 44.1 kHz and 32-bit depth. The recordings cover normal, heavy-load, and faulty operating conditions.
Sample Balance
The pre-segmented version contains three-second audio files. Its class distribution is slightly unbalanced:
| Class | Sample count | Proportion |
|---|---|---|
| Good | 105 | 29.4% |
| Broken | 124 | 34.7% |
| Heavy Load | 128 | 35.9% |
Feature Exploration
Time-domain plots revealed minor visual differences between motor states but were generally not informative for classification. Frequency-domain FFT plots displayed distinct spectral peaks characteristic of each state. Spectrograms provided detailed time-frequency representations but were more computationally demanding.
RMS reflects overall energy, which is higher for heavy-load motors. Zero-crossing rate indicates high-frequency activity and is lower for steady, loaded motors. Crest factor highlights occasional peaks, which are more pronounced in broken motors. Spectral centroid, bandwidth, and rolloff describe the frequency distribution. Together, these features provide discriminative information aligned with the motor operating condition. For simplicity and expandability, the final pipeline used FFT features instead.
Data Preprocessing
Audio preprocessing was applied consistently to training and test data. This reduces background noise, balances loudness, and removes irrelevant frequency components so that the model receives stable, standardized inputs.
Bandpass Filter
A bandpass filter isolates the desired frequency range, removing low-frequency hums such as electrical noise or microphone rumble and high-frequency artifacts such as hiss or sensor interference. Based on FFT visualizations, a range of 5-18,000 Hz was considered for this dataset.
Normalization
Normalization adjusts audio amplitude to ensure consistent loudness across samples. RMS normalization reflects average signal power more closely than peak normalization, producing more stable feature extraction and generally improving model convergence.
transforms = Compose([
BandpassFilter(sample_rate=SAMPLE_RATE, low_freq=LOW_CUT_FILTER,
high_freq=HIGH_CUT_FILTER, Q=6),
RMSNormalize(target_rms=0.1),
FFTTransform(n_features=FFT_FEATURES, return_db=True),
])
Feature Extraction
MFCCs are often used for non-stationary signals such as speech or music. Motor audio is comparatively steady and periodic, so time-sensitive extraction is less beneficial. Frequency-domain features derived from the FFT emphasize the harmonic and tonal structures that distinguish motor states. Unlike an STFT, a single FFT captures the essential stationary spectral characteristics with less computational cost.
Models
Neural Network
The neural network is a fully connected feedforward model implemented in PyTorch. It reduces the input from 1024 FFT features through hidden layers of 512 and 256 units to three output classes. It uses ReLU activations, cross-entropy loss, and the Adam optimizer with a learning rate of 0.001. Training uses batches of 16 and typically converges in approximately 10 epochs.
class SimpleNN(nn.Module):
def __init__(self):
super(SimpleNN, self).__init__()
self.fc1 = nn.Linear(FFT_FEATURES, FFT_FEATURES // 2)
self.fc2 = nn.Linear(FFT_FEATURES // 2, FFT_FEATURES // 4)
self.fc3 = nn.Linear(FFT_FEATURES // 4, NUM_CLASSES)
def forward(self, x):
if x.ndim > 2:
x = x.view(x.size(0), -1)
x = torch.relu(self.fc1(x))
x = torch.relu(self.fc2(x))
return self.fc3(x)
Support Vector Machine
An SVM was trained using the same 1024-dimensional FFT features and transforms. Features were standardized with Scikit-learn’s StandardScaler. GridSearchCV with three-fold cross-validation varied the regularization strength and RBF kernel coefficient.
param_grid = {
'C': [0.0001, 0.001, 0.01, 0.1, 1, 10],
'gamma': [0.0001, 0.001, 0.01, 0.1, 1, 10],
'kernel': ['rbf']
}
svm = SVC(probability=True, random_state=42)
grid_search = GridSearchCV(svm, param_grid, cv=3, scoring='accuracy', n_jobs=-1)
grid_search.fit(X_train_scaled, y_train)
# Best params: {'C': 0.1, 'gamma': 0.001, 'kernel': 'rbf'}
Results
Neural Network
The neural network achieved 98.61% accuracy on the test set.
| True class | Good | Broken | Heavy Load | Total |
|---|---|---|---|---|
| Good | 655 | 0 | 14 | 669 |
| Broken | 0 | 658 | 7 | 665 |
| Heavy Load | 12 | 0 | 675 | 687 |
| Total | 667 | 658 | 696 | 2021 |
| Class | Precision | Recall | F1-score | Support |
|---|---|---|---|---|
| Good | 0.9969 | 0.9761 | 0.9864 | 669 |
| Broken | 1.0000 | 0.9850 | 0.9924 | 665 |
| Heavy Load | 0.9634 | 0.9971 | 0.9800 | 687 |
| Accuracy | 0.9861 | 2021 | ||
| Macro avg | 0.9868 | 0.9860 | 0.9863 | 2021 |
| Weighted avg | 0.9866 | 0.9861 | 0.9862 | 2021 |
Misclassified files were distributed across stresstest (17), atmo_high (9), and talking_2 (2).
Support Vector Machine
The SVM achieved 98.81% accuracy.
| True class | Good | Broken | Heavy Load | Total |
|---|---|---|---|---|
| Good | 653 | 0 | 16 | 669 |
| Broken | 0 | 657 | 8 | 665 |
| Heavy Load | 0 | 0 | 687 | 687 |
| Total | 667 | 658 | 696 | 2021 |
| Class | Precision | Recall | F1-score | Support |
|---|---|---|---|---|
| Good | 1.0000 | 0.9761 | 0.9879 | 669 |
| Broken | 1.0000 | 0.9880 | 0.9939 | 665 |
| Heavy Load | 0.9662 | 1.0000 | 0.9828 | 687 |
| Accuracy | 0.9881 | 2021 | ||
| Macro avg | 0.9887 | 0.9880 | 0.9882 | 2021 |
| Weighted avg | 0.9885 | 0.9881 | 0.9882 | 2021 |
Misclassified files occurred in stresstest (16) and atmo_high (6).
Preprocessing Comparison
| Preprocessing pipeline | NN accuracy | SVM accuracy |
|---|---|---|
| Bandpass + Normalize + FFT | 98.61% | 98.81% |
| FFT only | 98.86% | 97.72% |
| Bandpass + FFT | 99.01% | 98.96% |
| Normalize + FFT | 80.41% | 88.27% |
The SVM achieved its highest accuracy with Bandpass + FFT, while the neural network remained highly accurate across tested pipelines except Normalize + FFT.
Discussion
Preprocessing Importance
The full pipeline provides strong and robust performance for both models. Both perform poorly when normalization and FFT are used without bandpass filtering, likely because normalization amplifies the noise floor when the signal-to-noise ratio is low. This highlights the importance of removing noise before normalization.
The SVM performs best without normalization on this dataset, although that approach may be less robust when recording levels vary. The neural network performs well with FFT alone, but training convergence is significantly slower without normalization.
Model Performance
Both models performed well, with the SVM achieving a slight advantage over the neural network. The SVM showed excellent recall for Heavy Load (1.00) and strong recall for Good (0.98) and Broken (0.99). The neural network achieved recall of 0.98, 0.99, and 1.00 for Good, Broken, and Heavy Load respectively.
The most common confusions involved Good and Heavy Load. Misclassifications corresponded to extreme stress tests and high atmospheric noise, where spectra differ most from the training conditions.
Future Generalization
Real-world deployment will involve domain shifts caused by different sensors, microphone placements, and ambient noise. The neural network could be adapted by freezing early layers and fine-tuning a final layer with a small amount of target-environment data. The SVM could be refit using representative samples and a scaler calibrated for the target device.
Both models would also benefit from confidence handling. Softmax entropy for the neural network or a margin/decision score for the SVM could allow the system to abstain or trigger a re-check for low-confidence or out-of-distribution inputs.
Conclusion
This study demonstrates that machine learning can classify electric motor states from audio recordings with high accuracy. The SVM achieved 98.81% accuracy and the neural network achieved 98.61%, making both viable options depending on deployment constraints and interpretability requirements.
Bandpass filtering proved to be a critical preprocessing step, particularly before normalization. Future work could explore convolutional neural networks operating on spectrograms and evaluate performance across a wider range of motor types, sensors, and operating conditions.
References
- Grollmisch, S., Abeßer, J., Liebetrau, J., and Lukashevich. IDMT-ISA-Electric-Engine Dataset, 2019.