Comparing Machine Learning Models for Audio Classification of Engines

Evaluating effectiveness and training time of neural networks, SVMs, and other models for engine sound classification

Introduction

Predictive maintenance is a critical component of modern industrial operations, enabling early detection of equipment faults and reducing costly downtime. Audio analysis is a powerful, non-invasive technique for condition monitoring because the sounds produced by a machine can reveal subtle changes in its operational state. This report investigates machine learning models for classifying electric motor status from audio signatures.

The objective is to distinguish three motor states: Good (normal operation), Broken (faulty), and Heavy Load. The analysis uses the IDMT-ISA-Electric-Engine dataset, which contains recordings of brushless DC motors under these conditions.


Data Exploration

The dataset, released by the Fraunhofer Institute for Digital Media Technology, contains approximately 42 minutes of mono WAV files recorded at 44.1 kHz and 32-bit depth. The recordings cover normal, heavy-load, and faulty operating conditions.

Sample Balance

The pre-segmented version contains three-second audio files. Its class distribution is slightly unbalanced:

ClassSample countProportion
Good10529.4%
Broken12434.7%
Heavy Load12835.9%

Feature Exploration

Time-domain plots revealed minor visual differences between motor states but were generally not informative for classification. Frequency-domain FFT plots displayed distinct spectral peaks characteristic of each state. Spectrograms provided detailed time-frequency representations but were more computationally demanding.

RMS reflects overall energy, which is higher for heavy-load motors. Zero-crossing rate indicates high-frequency activity and is lower for steady, loaded motors. Crest factor highlights occasional peaks, which are more pronounced in broken motors. Spectral centroid, bandwidth, and rolloff describe the frequency distribution. Together, these features provide discriminative information aligned with the motor operating condition. For simplicity and expandability, the final pipeline used FFT features instead.


Data Preprocessing

Audio preprocessing was applied consistently to training and test data. This reduces background noise, balances loudness, and removes irrelevant frequency components so that the model receives stable, standardized inputs.

Bandpass Filter

A bandpass filter isolates the desired frequency range, removing low-frequency hums such as electrical noise or microphone rumble and high-frequency artifacts such as hiss or sensor interference. Based on FFT visualizations, a range of 5-18,000 Hz was considered for this dataset.

Normalization

Normalization adjusts audio amplitude to ensure consistent loudness across samples. RMS normalization reflects average signal power more closely than peak normalization, producing more stable feature extraction and generally improving model convergence.

transforms = Compose([
		BandpassFilter(sample_rate=SAMPLE_RATE, low_freq=LOW_CUT_FILTER,
									 high_freq=HIGH_CUT_FILTER, Q=6),
		RMSNormalize(target_rms=0.1),
		FFTTransform(n_features=FFT_FEATURES, return_db=True),
])

Feature Extraction

MFCCs are often used for non-stationary signals such as speech or music. Motor audio is comparatively steady and periodic, so time-sensitive extraction is less beneficial. Frequency-domain features derived from the FFT emphasize the harmonic and tonal structures that distinguish motor states. Unlike an STFT, a single FFT captures the essential stationary spectral characteristics with less computational cost.


Models

Neural Network

The neural network is a fully connected feedforward model implemented in PyTorch. It reduces the input from 1024 FFT features through hidden layers of 512 and 256 units to three output classes. It uses ReLU activations, cross-entropy loss, and the Adam optimizer with a learning rate of 0.001. Training uses batches of 16 and typically converges in approximately 10 epochs.

class SimpleNN(nn.Module):
	def __init__(self):
		super(SimpleNN, self).__init__()
		self.fc1 = nn.Linear(FFT_FEATURES, FFT_FEATURES // 2)
		self.fc2 = nn.Linear(FFT_FEATURES // 2, FFT_FEATURES // 4)
		self.fc3 = nn.Linear(FFT_FEATURES // 4, NUM_CLASSES)

	def forward(self, x):
		if x.ndim > 2:
				x = x.view(x.size(0), -1)
		x = torch.relu(self.fc1(x))
		x = torch.relu(self.fc2(x))
		return self.fc3(x)

Support Vector Machine

An SVM was trained using the same 1024-dimensional FFT features and transforms. Features were standardized with Scikit-learn’s StandardScaler. GridSearchCV with three-fold cross-validation varied the regularization strength and RBF kernel coefficient.

param_grid = {
	'C': [0.0001, 0.001, 0.01, 0.1, 1, 10],
	'gamma': [0.0001, 0.001, 0.01, 0.1, 1, 10],
	'kernel': ['rbf']
}

svm = SVC(probability=True, random_state=42)
grid_search = GridSearchCV(svm, param_grid, cv=3, scoring='accuracy', n_jobs=-1)
grid_search.fit(X_train_scaled, y_train)

# Best params: {'C': 0.1, 'gamma': 0.001, 'kernel': 'rbf'}

Results

Neural Network

The neural network achieved 98.61% accuracy on the test set.

True classGoodBrokenHeavy LoadTotal
Good655014669
Broken06587665
Heavy Load120675687
Total6676586962021
ClassPrecisionRecallF1-scoreSupport
Good0.99690.97610.9864669
Broken1.00000.98500.9924665
Heavy Load0.96340.99710.9800687
Accuracy0.98612021
Macro avg0.98680.98600.98632021
Weighted avg0.98660.98610.98622021

Misclassified files were distributed across stresstest (17), atmo_high (9), and talking_2 (2).

Support Vector Machine

The SVM achieved 98.81% accuracy.

True classGoodBrokenHeavy LoadTotal
Good653016669
Broken06578665
Heavy Load00687687
Total6676586962021
ClassPrecisionRecallF1-scoreSupport
Good1.00000.97610.9879669
Broken1.00000.98800.9939665
Heavy Load0.96621.00000.9828687
Accuracy0.98812021
Macro avg0.98870.98800.98822021
Weighted avg0.98850.98810.98822021

Misclassified files occurred in stresstest (16) and atmo_high (6).

Preprocessing Comparison

Preprocessing pipelineNN accuracySVM accuracy
Bandpass + Normalize + FFT98.61%98.81%
FFT only98.86%97.72%
Bandpass + FFT99.01%98.96%
Normalize + FFT80.41%88.27%

The SVM achieved its highest accuracy with Bandpass + FFT, while the neural network remained highly accurate across tested pipelines except Normalize + FFT.


Discussion

Preprocessing Importance

The full pipeline provides strong and robust performance for both models. Both perform poorly when normalization and FFT are used without bandpass filtering, likely because normalization amplifies the noise floor when the signal-to-noise ratio is low. This highlights the importance of removing noise before normalization.

The SVM performs best without normalization on this dataset, although that approach may be less robust when recording levels vary. The neural network performs well with FFT alone, but training convergence is significantly slower without normalization.

Model Performance

Both models performed well, with the SVM achieving a slight advantage over the neural network. The SVM showed excellent recall for Heavy Load (1.00) and strong recall for Good (0.98) and Broken (0.99). The neural network achieved recall of 0.98, 0.99, and 1.00 for Good, Broken, and Heavy Load respectively.

The most common confusions involved Good and Heavy Load. Misclassifications corresponded to extreme stress tests and high atmospheric noise, where spectra differ most from the training conditions.

Future Generalization

Real-world deployment will involve domain shifts caused by different sensors, microphone placements, and ambient noise. The neural network could be adapted by freezing early layers and fine-tuning a final layer with a small amount of target-environment data. The SVM could be refit using representative samples and a scaler calibrated for the target device.

Both models would also benefit from confidence handling. Softmax entropy for the neural network or a margin/decision score for the SVM could allow the system to abstain or trigger a re-check for low-confidence or out-of-distribution inputs.


Conclusion

This study demonstrates that machine learning can classify electric motor states from audio recordings with high accuracy. The SVM achieved 98.81% accuracy and the neural network achieved 98.61%, making both viable options depending on deployment constraints and interpretability requirements.

Bandpass filtering proved to be a critical preprocessing step, particularly before normalization. Future work could explore convolutional neural networks operating on spectrograms and evaluate performance across a wider range of motor types, sensors, and operating conditions.

References