Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Binary classification assigns each sample to one of two classes. Our task in this chapter is a standard one in seismology: given a short window of seismogram data, decide whether it contains a seismic event (label 1) or only noise (label 0). Instead of raw waveforms, we work with four physical features computed from each window:

  • sta_lta: the ratio of a short-term average amplitude to a long-term average amplitude. An impulsive arrival raises the short-term average and pushes the ratio well above 1. The distribution is heavy-tailed, spanning roughly 0.4 to 40.
  • kurtosis: how impulsive the amplitude distribution is; spiky arrivals raise it.
  • spectral_centroid_hz: the amplitude-weighted mean frequency of the window, roughly 0.2 to 24 Hz. Local events carry more high-frequency energy than ocean-generated microseism noise.
  • dominant_freq_hz: the frequency of the spectral peak.

For the two-dimensional plots in this notebook we use sta_lta and spectral_centroid_hz: an event window tends to have both a high STA/LTA ratio and a high spectral centroid, while noise sits at low values of both.

1. Event versus noise data

We load a balanced set of 2000 windows, half events and half noise, from the course’s synthetic data package.

🖥️ Lecture slides — Session 14 (Fri Oct 30)

Loading...
<Figure size 640x480 with 1 Axes>

The log-scaled x-axis already tells us how to build the feature matrix: taking the logarithm of a heavy-tailed ratio like STA/LTA is standard feature engineering, and we do it before handing the features to any model.

Before any classifier, the baseline. We split the data first: everything that gets fit — a scaler, a trivial baseline, a classifier — sees only the training split, and the test split stays untouched until scoring. The trivial predictor assigns every window to the most frequent class; on this balanced set that gives close to 0.5 accuracy, and every classifier below must beat it.

Majority-class baseline accuracy: 0.49

We will start with the fundamental LDA.

The mean accuracy on the given test and labels is 0.952500
<Figure size 640x480 with 1 Axes>

LDA draws a single straight line through the feature space. It does well here because the two classes are roughly separable by a line.

Let’s try a different classifier: KNN.

The mean accuracy on the given test and labels is 0.960000
<Figure size 640x480 with 1 Axes>

Features with different physical units

Now we test what happens when we do not normalize the features. Our two features have different physical meanings: log-STA/LTA is dimensionless, the spectral centroid is a frequency. Suppose a colleague hands us the centroid in millihertz instead of hertz. Nothing physical has changed, but one axis is now numerically a thousand times larger.

The mean accuracy without scaling is 0.865000
<Figure size 640x480 with 1 Axes>

Accuracy drops and the decision boundary is now a set of horizontal bands: the classifier is effectively using the centroid alone. Distance-based classifiers see only numbers, not units. Euclidean distance in the (log-STA/LTA, millihertz) plane is dominated by the millihertz axis, so the STA/LTA information is ignored. Scaling with StandardScaler puts both features on equal footing, which is why we normalized before fitting.

2. Classifier Performance Metrics

In a binary classifier, we label one of the two classes as positive, the other class as negative. Let’s consider N data samples.

True Class \ predicted ClassNegativePositiveTotal
NegativeTrue NegativeFalse Positiven
PositiveFalse NegativeTrue Positivep
Totaln’p’N

In this example, there were originally a total of p=FN+TPp=FN+TP positive labels and n=FP+TNn=FP+TN negative labels. We ended up with p′=TP+FPp'=TP+FP predicted as positive and n′=FN+TNn'=FN+TN predicted as negative.

True positive TP: the number of data predicted as positive that were originally positive.

True negative TN: the number of data predicted as negative that were originally negative.

False positive FP: the number of data predicted as positive but that were originally negative.

False negative FN: the number of data predicted as negative but that were originally positive.

Confusion matrix: Count the instances that an element of class A is classified in class B: entry (i,j)(i, j) is the number of samples of true class ii predicted as class jj. scikit-learn sorts the classes by label, so with noise = 0 (negative) and event = 1 (positive) the matrix confusion_matrix prints is:

C=TNFPFNTP C = \begin{array}{|cc|} TN & FP \\ FN & TP \end{array}

The first row is the true-noise windows, the second row the true events; correct predictions sit on the diagonal. The confusion matrix can be extended for a multi-class classification and the matrix is KxK instead of 2x2. The best confusion matrix is one that is close to diagonal, with little off diagonal terms.

Other model performance metrics Model performance can be assessed with the following:

  • Error : the fraction of the data that was misclassified

    err=FP+FNNerr = \frac{FP+FN}{N} -> 0

  • Accuracy: the fraction of the data that was correctly classified:

    acc=TP+TNN=1−erracc = \frac{TP+TN}{N} = 1 - err --> 1

  • TP-rate: the fraction of the truly positive samples that are correctly classified:

    TPR=TPTP+FNTPR = \frac{TP}{TP+FN} --> 1

    This ratio is also the recall value or sensitivity.

  • TN-rate: the fraction of the truly negative samples that are correctly classified:

    TNR=TNTN+FPTNR = \frac{TN}{TN+FP} --> 1

    This ratio is also the specificity.

  • Precision: the ratio of samples predicted in the positive class that were indeed positive to the total number of samples predicted as positive.

    pr=TPTP+FPpr = \frac{TP}{TP+FP} --> 1

  • F1 score:

    F1=2(1/precision+1/recall)=TPTP+(FN+FP)/2F_1 = \frac{2}{(1/ precision + 1/recall)} = \frac{TP}{TP + (FN+FP)/2} --> 1.

The harmonic mean gives more weight to the lower of the two, so the F1 score is high only if both recall and precision are high.

How do precision and recall co-vary?

<Figure size 640x480 with 1 Axes>

From above, we see that a wide range of behavior between precision and recall can be expected. It is not because a classifier might do well in predicting one class than it does predicting the other class. This demonstrates the important of reporting both metrics.

Let’s print these measures for our scaled KNN classifier, evaluated on the test set, using scikit-learn.

confusion matrix
[[384  12]
 [ 20 384]]
precison, recall
0.9696969696969697 0.9504950495049505
F1 score
0.96

A complete well-formatted report of the performance can be called using the function classification_report:

Classification report for classifier KNeighborsClassifier():
              precision    recall  f1-score   support

           0       0.95      0.97      0.96       396
           1       0.97      0.95      0.96       404

    accuracy                           0.96       800
   macro avg       0.96      0.96      0.96       800
weighted avg       0.96      0.96      0.96       800


Precision and recall trade off: increasing precision reduces recall.

precision=TPTP+FPprecision = \frac{TP}{TP+FP}

recall=TPTP+FNrecall = \frac{TP}{TP+FN}

The classifier uses a threshold value to decide whether a data belongs to a class. Increasing the threshold gives higher precision score, decreasing the thresholds gives higher recall scores. Let’s look at the various score values.

Receiver Operating Characteristics ROC

It plots the true positive rate against the false positive rate. The ROC curve is visual, but we can quantify the classifier performance using the area under the curve (aka AUC). Ideally, AUC is 1.

ROC curve

[source: https://commons.wikimedia.org/wiki/File:Roc-draft-xkcd-style.svg]

We compute the curve on the test set: earlier editions of this book plotted the training-set ROC, and that flatters the model.

[[1.  0. ]
 [1.  0. ]
 [0.  1. ]
 [0.8 0.2]
 [0.  1. ]]
<Figure size 640x480 with 1 Axes>

We now explore the different classifiers packaged in scikit-learn. We can systematically test their performance and save the precision, recall, and F1 scores.

A word of caution on one entry: GaussianProcessClassifier training scales as O(n3)O(n^3) with the number of samples. It is fine here at n=2000n=2000, but it is a trap on real detection catalogs, which routinely hold 105 to 106 windows.

3. Model exploration

  • Explore How each of these models perform on the synthetic data.

  • Save in an array the precision, recall, F1 score values.

  • Find the best performing model

            CLF name  precision    recall  f1_score
0  Nearest Neighbors   0.969620  0.948020  0.958698
1         Linear SVM   0.969309  0.938119  0.953459
2            RBF SVM   0.967005  0.943069  0.954887
3   Gaussian Process   0.971939  0.943069  0.957286
4      Decision Tree   0.966837  0.938119  0.952261
5      Random Forest   0.964824  0.950495  0.957606
6           AdaBoost   0.959900  0.948020  0.953923
7        Naive Bayes   0.966921  0.940594  0.953576
8                QDA   0.966921  0.940594  0.953576

Imbalanced detection: when accuracy stops meaning anything

The balanced set above was a convenient classroom fiction. On a real station, event windows are rare: most of the day is noise. We now draw 5000 windows with an event fraction of 2 percent, about 100 events among 4900 noise windows — a 1:50 imbalance. Same features, same log transform, same scaling. The split is stratified so that the rare class keeps the same proportion in train and test.

label
0    4900
1     100
Name: count, dtype: int64
Majority-class baseline accuracy: 0.980

The trivial predictor that never declares an event already scores 0.98. Accuracy is now useless as a metric. Let’s train two real classifiers and look at the metrics that matter for the event class.

Logistic regression
  accuracy         : 0.992
  precision (event): 0.929
  recall (event)   : 0.650
  F1 (event)       : 0.765
Random forest
  accuracy         : 0.992
  precision (event): 0.853
  recall (event)   : 0.725
  F1 (event)       : 0.784

Both models beat the baseline accuracy by less than one percentage point, yet they miss roughly a third of the events. Recall on the event class is the number the analyst cares about: it counts the earthquakes the detector actually catches. The trade-off is operational. A detector tuned for high recall floods the analyst with false triggers to review; a detector tuned for high precision keeps the trigger list clean but misses small events. Where to sit on that curve is a decision about analyst time and scientific cost, not a modeling detail.

We can see the whole trade-off at once by plotting the ROC curve and the precision-recall curve side by side.

<Figure size 1100x450 with 2 Axes>

The ROC curve looks excellent, and that is exactly the problem. With a 1:50 imbalance the false-positive rate divides by the roughly 1960 noise windows in the test split, so it stays tiny even when false positives outnumber true detections. The precision-recall curve divides by the number of predicted positives instead, so it tracks the analyst’s actual workload: it drops as soon as the trigger list fills with noise. When the positive class is rare, read the PR curve, not the ROC.

Treating the imbalance

Diagnosing the imbalance is only half the job. Two standard knobs move the operating point without collecting new data:

  • class_weight="balanced" reweights the training loss so that each event error counts as much as roughly 49 noise errors, pushing the fitted boundary toward the rare class.
  • Threshold moving: predict cuts the predicted probability at 0.5, but nothing forces that choice. Lowering the threshold trades precision for recall along the curve we just plotted.
Loading...

Both knobs raise event recall by paying with precision: the trigger list gets longer, but fewer events are missed. Neither is “better” — they slide along the same precision-recall curve, and the choice of operating point belongs to the analyst, not the model.

Closing

The choice of metric is a design decision, not an afterthought. It encodes the operational cost of each error type: a missed event costs science, a false trigger costs analyst time. Decide what each error costs in your application, pick the metric that prices it, and only then compare models.