Skip to content
AdvancedObservation, analysis and ongoing review

Forecast Verification: Matched Cases and Reference Skill

Match real forecasts to compatible observations, inspect scalar, wind and probability errors, and compare reference skill on common cases.

Forecast learnersWeather-data analysts

Workflow

  1. Define the question before collecting cases

    Use the actual observation or review interval

    Choose one variable, unit, height, location footprint and valid interval per run. Record the forecast model/version, issue time, lead time and a reference forecast that would have been available then. Define any weather-regime grouping before reviewing performance so the verification question does not shift to favor one model.

  2. Match forecast and observation identities

    Use the actual observation or review interval

    Join by station or extraction location, valid time, level and accumulation interval. Keep issue time separate from valid time and use explicit UTC. Retain missing values and quality flags. When several forecast cycles verify at the same time, they must use the same observation; these repeated cases are not independent samples.

  3. Inspect scalar error and reference coverage

    Use the actual observation or review interval

    Calculate forecast-minus-observation bias, MAE and RMSE together with sample count. Review group and exact-lead summaries rather than comparing unlike lead distributions. Pearson correlation is undefined for a constant series or fewer than two pairs. Reference skill must compare model and baseline on the identical common subset.

  4. Use circular errors for wind direction

    Use the actual observation or review interval

    Match wind direction references and averaging periods. Compute the shortest angular difference across north instead of subtracting 350° from 10° as a 340° miss. Exclude cases where either wind is calm and retain missing counts. A zero circular resultant has no unique mean error.

  5. Verify probabilities against a defined event

    Use the actual observation or review interval

    State the threshold, footprint and time interval before coding observed events as 0 or 1. Use the original probabilities for Brier score and a documented reference for Brier skill. Reliability bins show observed event frequency and count; they must not replace exact probabilities with bin centers in the score.

  6. Review limitations and retain a reproducible result

    Use the actual observation or review interval

    Save input cases, missing-data rules, model/reference provenance and grouping choices. Compare additional periods or stations when the question requires them, and assess uncertainty with an appropriate statistical method outside these descriptive scores. Do not invent universal model MAE targets or claim a small sample proves one model superior.

Tools Used

Checklist

0 / 6 completed

Loading your checklist…

Preparation

Review

Reference Materials

Verification definitionsStandard

MET documents continuous error scores, correlation, MSE skill, Brier score and reference skill definitions.

Observation provenanceStandard

Verification depends on actual observations with compatible sampling and quality context.

  • Inspect a reference score of zero

    Skill relative to a perfect reference has a zero denominator and must remain undefined.

  • Keep dependent cases identifiable

    Multiple leads or forecast cycles can share one observed event; a large row count is not automatically a large independent sample.

Safety Notes

  • Descriptive scores do not establish statistical significance or operational reliability by themselves.
  • A retrospective score must not use information unavailable at the forecast issue time as an undisclosed reference.