Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

1. Introduction

Feature engineering is the process of transforming raw data into quantities — features — that can be used in statistical analysis and prediction.

Key types of features include:

Statistical features

Statistical features summarize the distribution of data values: central tendency, dispersion, and shape.

  • Mean: average value of the data points.
  • Variance/Standard Deviation: measure of the spread of the data.
  • Skewness: asymmetry in the distribution of values.
  • Kurtosis: “tailedness” of the distribution.
  • Percentiles: threshold values for different segments of the distribution.

Temporal features

Temporal features describe the time-dependent patterns within a series.

  • Autocorrelation: correlation of a time series with a lagged version of itself.
  • Trend: long-term increase or decrease in the data.
  • Seasonality: regular patterns that repeat over a specific period.
  • Change points: locations where the statistical properties of the series change.

Below we build a synthetic daily series with a trend and a yearly cycle, remove the trend, and measure the lag-1 autocorrelation of what remains.

🖥️ Lecture slides — Session 09 (Mon Oct 19)

Lag-1 autocorrelation of the detrended seasonal series: 0.68
<Figure size 1000x600 with 1 Axes>

The detrended series still carries the yearly cycle, so neighboring days are similar and the lag-1 autocorrelation is high. Compare with pure white noise of the same length: white noise has no memory, so its lag-1 autocorrelation is near zero.

Lag-1 autocorrelation of the detrended seasonal series: 0.68
Lag-1 autocorrelation of white noise:                   -0.00

Spatial features

For geospatial data, spatial features capture the relationships between different locations in an image or dataset.

  • Texture: patterns in the local intensity variations (e.g., smooth, rough).
  • Spatial correlation: measures how similar nearby locations are in terms of intensity values.
<Figure size 1200x600 with 4 Axes>

Spatial autocorrelation of the white noise, computed via the power spectrum (Wiener-Khinchin theorem):

Spatial autocorrelation of the correlated noise:

<Figure size 1100x600 with 8 Axes>

Estimate the feature “correlation length” of the two images. The correlation length is often measured as the distance at which the autocorrelation function decays to 1/e of its maximum value.

Estimated correlation length for white noise: 0.07 and for spatially correlated noise: 0.61

Fractal features

Fractal features describe the self-similarity or complexity of the data across scales. A common one is the fractal dimension: a smooth curve has dimension close to 1, while a rough curve that fills more of the plane has a dimension closer to 2.

The Higuchi method estimates the fractal dimension of a time series. It measures the average “length” of the curve at coarser and coarser subsamplings; the slope of log(length) versus log(scale) gives the dimension.

Higuchi fractal dimension of white noise:         1.99
Higuchi fractal dimension of the seasonal series: 1.98
Higuchi fractal dimension of smoothed noise:      1.10

White noise is maximally rough, so its dimension is near 2. Smoothing lowers the dimension toward 1. The fractal dimension is a single number that captures roughness, which makes it a compact feature for classifying signals.

2. Seismic Time Series Data Example

This section demonstrates how to automatically extract features using a simple Python package, tsfel, applied to seismic waveforms recorded in the Pacific Northwest.

The miniPNW subset includes labeled seismic waveforms for events of various origins:

  • Earthquakes
  • Explosions (mostly quarry blasts)
  • Surface events (such as avalanches and landslides)
  • Sonic booms
  • Thunder

We will explore how features vary across these classes of seismic events.

Dataset citation. The data is a subset of the PNW-ML benchmark: Ni, Y., Hutko, A., Skene, F., Denolle, M., Malone, S., Bodin, P., Hartog, R., & Wright, A. (2023). Curated Pacific Northwest AI-ready Seismic Dataset. Seismica, 2(1). doi:Ni et al. (2023). The archival access path to the full dataset is seisbench.data.PNW(); here we use a small subset (“miniPNW”) hosted on a class server.

We download two files from the class storage into ./data/: the waveforms as an HDF5 file and their associated metadata as a CSV file. The metadata file is small. The waveform file is large (about 1 GB), so we guard its download: if it fails, the notebook prints a message and skips the waveform-dependent cells.

Metadata file present: True
Downloading waveform file (large, be patient)...
Waveform file present: True

Metadata

We first read the metadata and arrange it into a pandas DataFrame.

Loading...

The nature of the event source is stored in one of the metadata attributes.

<ArrowStringArray> ['earthquake', 'explosion', 'sonic_boom', 'thunder', 'surface_event'] Length: 5, dtype: str

Assume that we are exploring features to classify the waveforms into the categories of event types. We take the source_type attribute as the label.

source_type
earthquake       500
explosion        500
surface_event    500
sonic_boom       126
thunder           94
Name: count, dtype: int64

How many seismic waveforms are there in each category?

     earthquake:   500 waveforms (29.1%)
      explosion:   500 waveforms (29.1%)
  surface_event:   500 waveforms (29.1%)
     sonic_boom:   126 waveforms (7.3%)
        thunder:    94 waveforms (5.5%)

Would you say that this is a balanced dataset with respect to the classes of interest? Class imbalance matters: a classifier trained on this set will see many more of the dominant class, and accuracy alone becomes a misleading score.

<Figure size 640x480 with 1 Axes>

Waveform data

Now we read the waveform data. It is stored in an HDF5 file under a finite number of groups. Each group has an array of datasets that correspond to the waveforms. To link the metadata to the waveform files, the key trace_name has the dataset ID. The address is labeled as follows:

bucketX$i,:3,:n

where X is the HDF5 group number and i is the index. The file typically has 3 waveforms, one from each direction of ground motion: N, E, Z. In the following, we focus on the vertical (Z) waveforms.

All cells that need the waveform file are gated behind a file-existence check, so the notebook still runs top to bottom if the download failed.

Below, a function to read a waveform out of the file.

The trace name is stored as a data attribute in the metadata.

'bucket1$0,:3,:15001'
The first dimension of the data is 3
The second dimension of the data is 15001
<Figure size 1000x600 with 3 Axes>

We extract the Z component of every waveform and stack them into a single array.

We have a total of 1720 data samples and each has 15001 data points

Automatic feature extraction with tsfel

Now we have data and its attributes, in particular the label as source type. We are going to extract features automatically with tsfel and explore how they vary across classes.

tsfel organizes features by domain (statistical, temporal, spectral). We load the default configuration:

/home/runner/work/mlgeo-book/mlgeo-book/.pixi/envs/default/lib/python3.12/site-packages/tsfel/feature_extraction/calc_features.py:195: SyntaxWarning: invalid escape sequence '\*'
  \**kwargs:
['spectral', 'statistical', 'temporal', 'fractal']

tsfel takes a 1D array and the sampling rate, and returns a one-row DataFrame of features. We wrap the extraction into a function that loops over a set of waveforms, attaches the label, and cleans up the column names (tsfel prefixes every column with 0_).

Note on runtime: extracting the full feature set for every waveform in miniPNW takes too long for class. We cap the extraction to a random subset of about 200 waveforms. Feel free to raise the cap outside of class.

Extracting features from sample 0/200
Extracting features from sample 20/200
Extracting features from sample 40/200
Extracting features from sample 60/200
Extracting features from sample 80/200
Extracting features from sample 100/200
Extracting features from sample 120/200
Extracting features from sample 140/200
Extracting features from sample 160/200
Extracting features from sample 180/200
Time taken to calculate features for 200 waveforms: 33.51 seconds
Loading...

New DataFrame clean up

Removing the samples with NaN or infinity.

no of samples in the dataframe: 200
no of samples in the new dataframe: 200

Explore correlation among features

<Figure size 1500x1000 with 2 Axes>

Blocks of highly correlated features are redundant: they carry the same information. Dimensionality reduction (next lesson) or feature selection can prune them.

Exploring the feature space for classification

Here we plot distributions of selected features among the classes. A feature is useful for classification when its distributions separate between classes.

<Figure size 640x480 with 1 Axes>
/home/runner/work/mlgeo-book/mlgeo-book/.pixi/envs/default/lib/python3.12/site-packages/pandas/core/arraylike.py:402: RuntimeWarning: invalid value encountered in log10
  result = getattr(ufunc, method)(*inputs, **kwargs)
<Figure size 640x480 with 1 Axes>
<Figure size 640x480 with 1 Axes>

Student exercise

  1. Which of the features will be best to classify among the event type classes?

Use the scaffold below. The idea: a feature separates two classes well when the difference between the class means is large compared to the spread within each class. Rank the features by that criterion.

ECDF_9                     inf
ECDF_5                     inf
ECDF_6                     inf
ECDF_2                     inf
ECDF_4                     inf
Entropy                    inf
Spectral variation    2.033088
MFCC_0                1.909093
MFCC_3                1.392157
LPCC_11               1.376106
dtype: float64
References
  1. Ni, Y., Hutko, A., Skene, F., Denolle, M., Malone, S., Bodin, P., Hartog, R., & Wright, A. (2023). Curated Pacific Northwest AI-ready Seismic Dataset. Seismica, 2(1). 10.26443/seismica.v2i1.368