Validation indices
Statistical indices for model validation
This page describes in detail the indices used in validation to compare modelled values with those observed by the monitoring stations. Each index measures a different aspect of the model-observation agreement (systematic bias, error, correlation, ability to reproduce variability, ability to detect threshold exceedances): for a robust assessment they should be read together, not in isolation.
Notation. For each pair of time-aligned values: \( o_i \) = observed value, \( m_i \) = modelled value, with \( i = 1 \dots n \) (number of valid pairs). Means are denoted \( \bar o \) and \( \bar m \), standard deviations \( \sigma_o \) and \( \sigma_m \). The fractional and normalized indices (NMB, NME, MFB, MFE) are meaningful only for non-negative quantities (concentrations); for sign-changing quantities use the "absolute" indices (MB, MAE, RMSE, R, KGE, NSE).
1. Bias (systematic error)
Measure the average tendency of the model to overestimate (positive value) or underestimate (negative value) the observations. Optimal value: 0.
MB — Mean Bias
Mean bias in the physical units of the phenomenon. It is the most straightforward to interpret but scale-dependent: most useful at low concentrations, where percentage indices become unstable.
NMB — Normalized Mean Bias
Bias normalized by the sum of observations: expresses the overall over/under-estimation as a percentage. It is the typical reference for regulatory assessment (e.g. annual mean of NO2).
MFB — Mean Fractional Bias
Symmetric and bounded fractional bias: it weights overestimation and underestimation equally and is robust to outliers and skewed distributions. Recommended by the EPA guidelines (Boylan & Russell, 2006) for particulate matter, with an acceptability threshold of \( |\mathrm{MFB}| < 60\% \).
Interactive example — bias
2. Error (accuracy)
Measure the magnitude of the deviations, regardless of sign. Optimal value: 0. RMSE penalizes large errors (peaks) more heavily.
MAE — Mean Absolute Error
Mean absolute error: the "typical" error in physical units. More robust than RMSE to anomalous values.
RMSE — Root Mean Square Error
Root mean square error: due to the squaring it weights large deviations heavily, so it is sensitive to the model's ability to reproduce peaks. Always \( \mathrm{RMSE} \ge \mathrm{MAE} \).
CRMSE — Centered RMSE
"Centered" RMSE, i.e. with the mean bias removed: it isolates the error of shape/amplitude (phase and variability) from the absolute level error. It is the quantity shown on Taylor diagrams.
NME — Normalized Mean Error
Mean absolute error normalized by the observations: the MAE expressed as a percentage, comparable across pollutants of different scale.
MFE — Mean Fractional Error
Symmetric fractional error, the "absolute-value" counterpart of MFB. Recommended by the EPA for PM with a threshold of \( \mathrm{MFE} < 75\% \).
Interactive example — RMSE vs MAE
3. Correlation (association)
Measure how much model and observations "move together" over time, regardless of bias and scale. Optimal value: 1.
R — Pearson correlation coefficient
Measures the strength of the linear relationship. It sees neither the bias nor the difference in amplitude: a model can have \( r=1 \) while being shifted or scaled. It must therefore always be paired with a bias index and a variability index.
R² — Coefficient of determination
Square of Pearson's r: fraction of the variance of the observations linearly explained by the model (e.g. \( R^2 = 0.8 \) → 80% of the variability). It is redundant with R, so it is normally not weighted in the score, but remains selectable.
Spearman — Rank correlation
It is Pearson's coefficient computed on the data ranks (\( R_{o_i} \) and \( R_{m_i} \)): it measures monotonic association (not just linear) and is robust to outliers and skewed distributions.
Only with no tied ranks (no repeated values) does it reduce to the compact form \( \rho = 1 - \dfrac{6\sum_i d_i^{\,2}}{n\left(n^2-1\right)} \), with \( d_i = R_{o_i}-R_{m_i} \). The actual computation uses the general definition, corrected for tied ranks (ties), not the approximation. Useful where the relationship is non-linear (e.g. SO2, O3, NOx).
MI — Mutual Information
Measures the general dependence between observed and model, including non-linear and non-monotonic relationships: how much uncertainty about one variable is reduced by knowing the other. Unlike R and Spearman, it captures any form of relationship.
Estimated via a 2D histogram of the two series: it is 0 when observed and model are independent and grows with dependence. Note: it measures dependence, not agreement — it stays high even for a strongly biased or anti-correlated model (it ignores bias and scale) and is not normalized. For this reason it has weight 0 in UISH: it is intended as a non-linear dependence diagnostic for other projects.
Interactive example — correlation (R, R², Spearman, MI)
4. Efficiency (overall skill scores)
Synthetic indices that combine several aspects into a single score. Optimal value: 1.
IOA — Index of Agreement (Willmott)
Measures the overall agreement by normalizing the squared error against the maximum possible difference. \( d=1 \) perfect agreement, \( d=0 \) no agreement.
KGE — Kling-Gupta Efficiency
It explicitly decomposes performance into three components — correlation \( r \), bias \( \beta \) and variability \( \gamma \) — and combines them. It is a very informative skill score: a low value immediately indicates which of the three components is lacking.
NSE — Nash-Sutcliffe Efficiency
Compares the model error with the variability of the observations. \( \mathrm{NSE}=1 \) perfect model; \( \mathrm{NSE}=0 \) the model is as good as simply using the observed mean; \( \mathrm{NSE}<0 \) the model is worse than the mean.
Interactive example — KGE / NSE
5. Regression and variability (diagnostics)
Supporting quantities: they normally do not enter the weighted score but help interpret the other indices (they are used in the scatter plot and the Taylor diagram).
Slope and intercept (linear regression)
Regression line of the model on the observations. The slope \( a \) indicates a proportional (scale/amplitude) error, the intercept \( b \) a constant error (offset). Ideally \( a=1 \) and \( b=0 \).
Standard deviations and their ratio
Measure the variability of observations and model. The ratio of the standard deviations tells whether the model reproduces the amplitude of the fluctuations: >1 the model is too "noisy", <1 too flat. It is one of the coordinates of the Taylor diagram.
6. Categorical (threshold) indices
They do not assess the continuous value but the model's ability to detect events, i.e. exceedances of a threshold \( \tau \) (e.g. a regulatory limit). They are built from the contingency table that classifies each instant according to observed and modelled threshold:
| Observed \( \ge \tau \) | Observed \( < \tau \) | |
|---|---|---|
| Model \( \ge \tau \) | \( a \) — hit | \( b \) — false alarm |
| Model \( < \tau \) | \( c \) — miss | \( d \) — correct negative |
POD — Probability of Detection
Fraction of real events correctly detected by the model (detection rate). It does not penalize false alarms: read it together with FAR.
FAR — False Alarm Ratio
Fraction of model alarms that turn out to be false. On its own it can be "gamed" by never raising an alarm: complementary to POD.
CSI — Critical Success Index
Combines hits, false alarms and misses into a single score (it ignores the correct negatives \( d \), often very numerous). A good summary of event detection quality.
HSS — Heidke Skill Score
Measures the model's skill relative to a random forecast: \( \mathrm{HSS}=1 \) perfect detection, \( \mathrm{HSS}=0 \) equivalent to chance, \( \mathrm{HSS}<0 \) worse than chance. It also accounts for the correct negatives.
Interactive example — categorical (threshold) indices
7. Exposure indices and error decomposition
Derived indicators used in the most recent charts of the validation page: the regulatory ozone-exposure indices (compared model vs measurements) and the decomposition of the RMSE into systematic and stochastic parts.
SOMO35 — Sum of Ozone Means Over 35 ppb
Annual sum of the excesses of the daily maximum of the 8-hour running mean (\( \mathrm{MDA8} \)) above 70 µg/m³ (≈ 35 ppb). It is the reference health indicator for chronic ozone exposure. Being an index accumulated on the upper tails of the distribution, it amplifies the model's systematic biases on the peaks.
AOT40 — Accumulated Ozone exposure over Threshold 40 ppb
Sum of the hourly excesses above 80 µg/m³ (≈ 40 ppb) during the daylight hours of the growing season: it is the vegetation-protection indicator (Dir. 2008/50/EC). The model-observed comparison tells whether the model is reliable for regulatory purposes.
Error decomposition (Willmott)
From the linear regression \( \hat m_i = a + b\,o_i \), the systematic part \( \mathrm{RMSE}_s \) measures the error explained by the regression line (reducible with a calibration/bias correction), the stochastic part \( \mathrm{RMSE}_u \) the residual scatter around it (tied to model physics and resolution). Knowing which one dominates points to the most effective improvement path.
Episodes (stratification)
Besides the temporal partitions (season, hour, month…), the metrics are
also computed on the peak episodes of the observations only:
day_p90 (hours of the days whose daily mean is > P90),
hour_p90 (hours with value > hourly P90) and
exceed (exceedances of the regulatory threshold). The
comparison with the all-data metrics shows how much the model quality
degrades precisely in the episodes that matter for health and for the
legal limits.
Keywords: validation, metrics, statistics, bias, RMSE, correlation, KGE, NSE, Pearson, Spearman, indices