Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

The perceptron is the building block of neural networks. It computes a weighted sum of its inputs, adds a bias, and passes the result through an activation function. In this notebook we:

  1. Build a perceptron from scratch with NumPy.
  2. Train it with the classic perceptron learning rule to separate seismic events from noise.
  3. Reuse the same perceptron as a linear regressor trained by gradient descent, and compare it to ordinary least squares (OLS).
  4. Study how the learning rate changes training.

🖥️ Lecture slides — Session 21 (Wed Nov 18)

1. A perceptron from scratch

A perceptron with inputs x1,…,xnx_1, \dots, x_n, weights w1,…,wnw_1, \dots, w_n, and bias weight bb computes

o=ϕ(∑jwjxj+b),o = \phi\left(\sum_j w_j x_j + b\right),

where ϕ is the activation function. We implement three activations: none (identity), sigmoid, and step.

1.1 Test the perceptron

With one weight set to 1, bias 0, and no activation, the perceptron should return its input unchanged.

array([100. , -1. , 0.1])

2. Classifying seismic detections

Seismic detectors summarize short windows of a seismogram with a few features. The mlgeo_synth.detector_features table contains four such features per window: the STA/LTA ratio (sta_lta), the kurtosis of the waveform, the spectral centroid in Hz, and the dominant frequency in Hz. The label is 1 for an earthquake window and 0 for a noise window, with balanced classes.

We use two features, sta_lta and spectral_centroid_hz. Events have impulsive arrivals (high STA/LTA) and more high-frequency energy (high spectral centroid) than background noise.

     sta_lta   kurtosis  spectral_centroid_hz  dominant_freq_hz  label
0  10.419133  22.113184              9.556484          7.786286      1
1   3.314406   2.773164              1.716694          1.009758      0
2   2.336543   4.639630              1.271397          0.251278      0
3   3.976867  13.591792              6.692700          4.682170      1
4   1.151811   3.131239              4.474578          1.426544      0
<Figure size 600x450 with 1 Axes>

2.1 Why an easy two-feature slice?

The classic perceptron learning rule only converges when the data are linearly separable: some line must classify every training sample correctly. If the classes overlap, the rule keeps updating forever. The scatter above shows a small overlap band between the two clouds, so the full dataset is not linearly separable.

We therefore select a near-linearly-separable subset: we compute a simple combined score, sta_lta + spectral_centroid_hz, and drop the samples in the overlap band. We also stick to exactly two features so we can draw the decision boundary as a line in the plane. Real detector data is messier; later notebooks handle the overlap with multi-layer networks and gradient-based losses.

<Figure size 600x450 with 1 Axes>

2.2 The perceptron learning rule

We start with all weights set to 0 and use the step activation, so the perceptron outputs 0 or 1. We then sweep through the dataset. For each sample we:

  1. Make a prediction with the current weights.

  2. Update each weight with the rule

    wj=wj+Δwjw_j = w_j + \Delta w_j, where Δwj=η (ti−oi) xji\Delta w_j = \eta\,(t^i - o^i)\,x^i_j

    Here η is the learning rate, tit^i and oio^i are the target and the output for sample ii, and xjix^i_j is input jj of that sample. The bias weight gets the same update with input 1. Note that correctly classified samples (ti=oit^i = o^i) leave the weights unchanged.

  3. Repeat full sweeps (epochs) over the data until one epoch produces zero mistakes.

One detail: with zero initial weights, the learning rate only scales the final weight vector. The step activation ignores that scale, so the sequence of predictions, and the number of epochs to converge, is the same for any η>0\eta > 0.

Converged after 2 epochs; mistakes per epoch: [42, 0]
Weights: [0.03415388 0.03796047], bias: -0.220, training accuracy: 1.000

The rule converges in a few epochs and classifies every training sample correctly. Because we use two features, the decision boundary is the line w1x1+w2x2+b=0w_1 x_1 + w_2 x_2 + b = 0, which we can draw directly.

<Figure size 600x450 with 1 Axes>

We just implemented an online algorithm: the weights update after every sample, not after seeing the whole dataset.

Exercise. Swap spectral_centroid_hz for kurtosis in the subset-selection cell (keep the same score-based filter on the original features). Does the learning rule still converge within 50 epochs? Why or why not?

3. Fitting a line with gradient descent

The same perceptron, with no activation, is a linear model o=wx+bo = w x + b. Instead of the perceptron rule, we can train it by gradient descent on a cost function. We generate a noisy linear dataset to fit.

<Figure size 600x400 with 1 Axes>

3.1 Cost function

We use the mean squared error (MSE) as the cost.

np.float64(0.0)

3.2 Gradient descent

Gradient descent repeats three steps: predict, measure the cost, and move the weights a small step against the cost gradient. For MSE and a linear model the weight update is Δw=η XT(t−o)\Delta w = \eta\, X^T (t - o) and the bias update is Δb=η∑i(ti−oi)\Delta b = \eta \sum_i (t^i - o^i). We stop when the cost change falls below a tolerance, the cost stops being finite (divergence), or we hit the iteration limit.

<Figure size 1200x400 with 3 Axes>
'The mean squared error of our perceptron on the test half is: 1.122684'

3.3 Benchmark against ordinary least squares

OLS solves the same problem in closed form. A gradient-descent perceptron that works should land on nearly the same line.

<Figure size 1200x400 with 1 Axes>

4. How the learning rate changes training

The learning rate η controls the step size of gradient descent. Too small and training crawls; too large and the cost oscillates or diverges. The grid below trains the same perceptron from the same zero start with four learning rates and plots the cost curve for each (log scale on the cost axis).

<Figure size 1000x700 with 4 Axes>

Read the four panels: at η=10−5\eta = 10^{-5} the cost is still falling after 2000 iterations (too slow). At η=10−4\eta = 10^{-4} it converges but takes over a thousand iterations. At η=10−3\eta = 10^{-3} it converges in a few hundred iterations. At η=10−2\eta = 10^{-2} the cost grows without bound: the steps overshoot the minimum and training diverges (the curve stops where the cost becomes infinite).

4.1 Interactive explorer

The widget below lets you vary the learning rate, iteration count, stopping criterion, and initial weights, then retrain and replot. Widgets do not run in the static rendered book: download and run this notebook in Jupyter to use it. The static grid above shows the same lesson.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...

Summary

  • A perceptron is a weighted sum plus a bias, passed through an activation function.
  • The classic perceptron learning rule converges only on linearly separable data. It separated our event and noise windows using two detector features, and the two-feature choice let us draw the decision boundary.
  • With no activation, the same perceptron is a linear regressor. Gradient descent on the MSE cost recovers nearly the same line as OLS.
  • The learning rate sets the trade-off between slow convergence and divergence.

Next: 4.1 Neural Networks stacks perceptrons into layers to handle data that a single line cannot separate.