STATSEI14 · PRACTICAL TUTORIAL

Deep Learning
Meets ETAS

Practical Hybrid Workflows for Seismic Forecasting

Francisco Plaza-Vega
Universidad de Santiago de Chile

13 October 2026 · Chile

30 minutes of concepts and research · 30 minutes of guided practice

What can we learn from seismicity?

Recorded training-period events in the Chilean study region, with coastlines and a locator.

Signals reveal recorded events.

Catalogues describe event histories.

Models turn histories into questions about what may follow.

What information do we represent, and what do we ask a model to learn?

Map: frozen USGS training catalogue and Natural Earth. Illustrates today’s regional setting.

Deep learning in seismology: inputs and targets

Example Input Learning task
ConvNetQuake · 2018 Recorded waveforms Detect events and classify their source region
RECAST · 2023 Past event times and magnitudes Estimate the next-event time distribution
Delgado et al. · 2025 Features of catalogued events Classify foreshock, mainshock and aftershock labels
Farfán et al. · 2025 Sequences of daily spatial grids Forecast gridded seismic activity

Today’s practical: 30 days of history and a seven-day occurrence probability.

A network learns a mapping from examples

Represented data → Neural network → Defined output

Architecture determines how information is combined.

Training adjusts weights to reduce the loss.

The output and the loss must match the scientific question.

MLP: combine measured features

English adaptation of the course multilayer perceptron diagram: input features, hidden layers and output.

A multilayer perceptron learns nonlinear combinations of a feature vector.

Temporal laboratory: the MLP uses 90 values at fixed lag positions. Diagram adapted from the author’s DL course, Unit 2.

CNN: learn local patterns with shared filters

English course-style diagram of local convolution filters, feature maps and an output.

The same filter examines different neighborhoods.

Temporal laboratory: CNN1D shares filters across days. Diagram adapted from the author’s DL course, Unit 3.

LSTM: carry information through a sequence

English course-style recurrent diagram showing inputs through time, shared LSTM updates and a sequence summary.

A learned state summarizes the history for a defined output.

Same learned update across time. Adapted from the author’s DL course, Unit 4.

Generative learning: produce new samples

English adaptation of a GAN learning diagram: noise feeds a generator; generated and observed examples inform an adversarial training signal.

A generator learns to produce samples from a learned distribution.

GAN is one approach to generative learning. Adapted from the author’s DL course, Unit 5.

Our first case: learning from an ETAS representation

Original published-project workflow linking data processing, LSTM and CNN to temporal and spatial summaries.

Scientific question

Can neural networks learn the evolution of a statistical representation of seismic activity?

Two views

A daily time series and a sequence of spatial maps.

Nicolis, Plaza & Salas · Spatial Statistics 42 (2021), 100442 · CSN Chile, 2000–2017

ETAS supplies a statistical description of activity

\lambda(t\mid\mathcal H_t)=\mu+\sum_{t_i<t}\text{contribution from event }i

Background + magnitude-dependent contributions that decay with time

Spatial ETAS also describes how contributions vary over space.

This conditional occurrence rate becomes a representation for learning.

Conceptual sketch. Process intensity is different from magnitude and shaking intensity.

The temporal task: tomorrow’s maximum intensity

30 daily maxima → LSTM → Next-day maximum

Input and target: summaries of the ETAS-derived intensity.

Learning question: how much of its evolution can a temporal network capture?

Specify what is being predicted before interpreting a fit statistic.

Published case, Nicolis, Plaza & Salas (2021). The target is a model-derived quantity.

The spatial task: the macrozone of the maximum

30 daily ETAS maps → CNN → Six zone probabilities

Which macrozone contains tomorrow’s maximum estimated intensity?

The representation determines which information the network can use.

Published case, Nicolis, Plaza & Salas (2021). These probabilities concern six macrozones.

With Isidora: can we generate aftershock sequences?

Mainshock conditions + noise → Generator → A possible sequence

Thesis question

Can conditional generative models reproduce the temporal pattern of aftershock sequences?

Model families explored

Conditional GAN · WGAN-GP · attention-based generator

A richer output requires us to decide what a sequence actually represents.

Isidora Jara Muñoz (2025), thesis · Subsequent manuscript by Plaza-Vega, Jara, Nicolis & Salinas (2026), under review

One sequence mixes absence and magnitude

Representation in the manuscript

14 days → 224 intervals of 90 minutes → maximum magnitude per occupied interval

2.5 2.5 4.1 2.5 3.4 2.5 2.5 5.0

Schematic values. Gray: absence marker. Orange: maximum magnitude in an occupied interval.

“No recorded event” and “event magnitude” are different outcomes.

Generated sequences distort event occupancy

Temporal GANT evaluation from the manuscript: generated sequences occupy too many intervals overall and too few of the larger-magnitude intervals.

Assess which properties of the observed sequences are reproduced.

2026 manuscript under review · Occupied intervals, not total events · Source-case results

The article revisits what the model should learn

Empirical reformulation

Aggregate targets: occurrence, exceedance and count summaries.

The classification branch uses tree models.

A future neural formulation

Occurrence probability, followed by magnitude conditional on occurrence.

Evaluate both parts and their joint behavior.

Today we isolate occurrence in a fixed future interval as the learning target.

ETAS-informed generation remains future work. The two-part neural model is a proposal.

The practical puts the same decisions in your hands

Case Representation Output
Spatial Statistics ETAS time series and maps Intensity maximum and its macrozone
Isidora / manuscript Mainshock conditions; binned sequences Conditional sequences; aggregate targets
Today Daily catalogue channels Weekly occurrence probability

Live: LSTM and Transformer. Laboratory extensions: MLP, CNN1D, maps and graphs.

Today’s target: at least one event next week

A forecast origin separates 30 historical input days from a seven-day binary target interval.

30 past days → probability of at least one catalogue M ≥ 5 event in seven days.

One input: 30 days × 3 channels · Region: 28–34°S, 69–74°W · Input M ≥ 4.5 · Monday origins

Evaluate using the temporal order of the problem

Chronological training, validation and test periods, with crossing future target weeks excluded.

Today we compare validation probabilities using Brier score, log loss and calibration.

Train 2001–2016 · Validation 2017–2020 · Test 2021–2025 was inspected previously and is excluded from this comparison.

Experiment: how does the same history enter each model?

30 days × 3 catalogue channels → LSTM → weekly probability

The same history → Transformer → the same target

Build one example: identify the input window, its origin and its label.

Train and evaluate: read compile, fit, predict and compare probability errors.

What changes when different architectures process the same information?

Prepared temporal notebook · Seed 7 · MLP and CNN1D for the laboratory · Activity-score experiment as an extension

A Transformer combines the same history differently

Compact Transformer: the same 30 days by three catalogue channels, a 16-feature projection with learned day positions, one two-head attention block with residual normalization and a feed-forward layer, average pooling, and a weekly probability.

Same history and target. A different way to combine information across days.

Current temporal tutorial · 2,257 trainable parameters versus 1,297 for LSTM · One fixed, small configuration

Validation probabilities from one declared seed

Model Validation Brier Validation log loss
Training prevalence 0.187690 0.562855
Logistic regression 0.203664 0.614405
LSTM 0.181521 0.547512
Transformer 0.187901 0.562477

Lower is better. Compare these scores with the calibration plot and its sample counts.

These are development results from one region and one validation period.

Linear temporal tutorial · One fit per neural model, seed 7 · 208 validation origins · Logistic regression uses all 90 lagged values

What can this result tell us about the problem?

Reading the validation results

Compare each model with training prevalence.

Read Brier, log loss and calibration together.

Questions for the next experiment

Which representation and target preserve useful structure?

What information is available in the catalogue?

Would more data, capacity or a different training budget change the result?

Useful deep learning depends on how we represent the phenomenon and formulate the task.

Validation also controls early stopping. One seed does not measure initialization sensitivity.

What did we learn, and what would we test next?

  1. What information did our daily channels preserve or discard?
  2. What did the LSTM and Transformer comparison establish?
  3. Which extension would you test next: maps, graphs or an activity score?

Deep learning offers flexible tools. Their scientific value depends on the question and the evidence.

Homework: guided reading or a controlled extension · Build on today’s LSTM and Transformer comparison

Appendix · the LSTM steps

inputs = keras.Input(shape=(30, 3))
state = keras.layers.LSTM(16)(inputs)
probability = keras.layers.Dense(1, activation="sigmoid")(state)
model_lstm = keras.Model(inputs, probability)
model_lstm.compile(optimizer="adam", loss="binary_crossentropy")
history_lstm = model_lstm.fit(
    X_train, y_train, validation_data=(X_val, y_val),
    epochs=20, batch_size=64, callbacks=[stopping_lstm], shuffle=False)
p_lstm = model_lstm.predict(X_val).ravel()
brier = np.mean((p_lstm - y_val) ** 2)

Define the model · compile · fit · predict · evaluate probabilities

Appendix · quantities and scientific meaning

Quantity Question answered
Process intensity At what conditional rate do events occur?
Magnitude How large is an event?
Shaking intensity What motion or effects occur at a location?
Risk What consequences may occur given exposure and vulnerability?

Today’s output is a probability of a defined occurrence event.

Appendix · earlier activity-score extension

r_t(s)=\sum_{t-L\le t_i<s}10^{\alpha(M_i-M_0)} \left(1+\frac{s-t_i}{c}\right)^{-p},\quad s\le t.

Illustrative: α=1, c=1 day, p=1.1, L=30 days.

Earlier feature experiment: both LSTMs use the same available history.

Representation: log1p(score) is an additional daily channel.

Optional extension from the earlier practical. The illustrative index is not a fitted ETAS rate.

Appendix · original generative examples

Original manuscript figure comparing observed and generated sequences for cGAN, WGAN-GP and GANT in two illustrative cases.

2026 manuscript under review · Original PI annotations use a mean-reference denominator · Illustrative cases

Appendix · evidence and availability

Check Meaning
Chronological, purged targets Future target intervals do not cross partitions.
Training-only normalization Later inputs do not set scaling parameters.
Frozen revised catalogue Data bytes and provenance are reproducible.
Historical catalogue vintages Not reconstructed.

Reproducibility and historical information availability are separate requirements.

Appendix · sources and further reading