Key messages from “The Jive Verification System and Its Transformative Impact on Weather Forecasting Operations,” by Nicholas Loveday (Bureau of Meteorology), Deryn Griffiths, Tennessee Leeuwenburg, Robert J. Taggart, Thomas C. Pagano, George Cheng, Kevin Plastow, Elizabeth E. Ebert, Cassandra Templeton, Maree Carroll, Mohammadreza Khanarmuei, and Isha Nagpal. Published online in BAMS, November 2024. For the full, citable article, click the link above.
As models and postprocessing techniques advance, weather forecasting agencies face several questions: What is the role of the operational meteorologist in making forecasts? Where, when, and how can meteorologists improve upon automated guidance? What parts of the forecast production process could be automated? What are the optimal forecasting strategies for meteorologists? The Australian Bureau of Meteorology has grappled with these questions for over a decade. Answering them quantitatively required comprehensive verification, which only became possible with the development of the Jive verification system. Starting as a research project in 2015 and becoming operational in 2022, Jive was designed to evaluate the Bureau’s public weather forecasts.
Like the U.S. National Weather Service, operational meteorologists in Australia produce weather forecasts for the public by “editing” grids of automated forecasts where subject matter expertise allows. These edited gridded forecasts populate the Bureau’s website and mobile app. They are also provided to downstream users such as fire agencies. Jive aims to make this process more efficient and quantify the value added by human forecasters compared to automated systems. Insights from Jive have informed automation decisions, led to more accurate forecasts, and clarified the evolving role of meteorologists.

Jive is a comprehensive verification environment featuring dashboards for visualizing daily forecast performance and comparing different forecast systems (e.g., automated vs. human-edited) over longer periods. A key component is a Jupyter Notebook server equipped with the Jive Python library, allowing meteorologists and scientists to use a web browser to conduct detailed case studies, develop new verification methods, and experiment with forecasting strategies. The system includes a wide range of verification metrics, including standard measures as well as innovative, user-focused metrics developed specifically to answer questions that the Bureau faced. These include methods for measuring how well forecasts predict extremes and metrics assessing the economic value of forecasts. Built using Python and libraries like xarray, the system allows for flexible analysis of various data types (point, grid).
Evidence Targeted Automation dashboards in Jive compared the performance and value of purely automated postprocessed forecasts and those edited by meteorologists. This analysis revealed where human intervention helped accuracy and where it inadvertently degraded it. This evidence helped foster meteorologists’ confidence in trusting automated systems in more situations, enabling a more streamlined forecast process. Exploring verification results in Jive allowed us to understand where meteorologists consistently made better predictions than the automated guidance (e.g., for wind speeds on high mountain peaks). These insights were then used to improve the automated systems themselves. Now, automation handles routine situations for many parameters more effectively, freeing up meteorologists to focus on high-impact weather events and decision-support activities.
One surprising aspect of the work was that the verification highlighted ambiguities in how the forecasts should be defined, such as differing interpretations of wind speed forecasts. Automated systems might target a median on-the-hour value, while forecasters focused on predicting peak winds, believing this was more relevant for warnings. Jive’s analysis spurred discussions that led to clearer definitions and the creation of distinct forecast products for different user needs (e.g., “wind-on-hour” vs. “max-wind-in-hour”). Clarifying how forecasts are defined removed the tensions between the goals of automated guidance and meteorologists. This helped remove barriers to meteorologists adopting and relying on automated guidance.
Access to dedicated dashboards and Jupyter Notebooks allowed operational staff to directly engage with verification data, giving them the ability to review their recent performance, identify systematic biases (like underforecasting the daily chance of precipitation), develop targeted forecasting strategies, and verify their effectiveness, sometimes achieving substantial skill improvements.
The Jive project shows the significant value of investing in verification science and technology and embedding these tools and practices within the daily operational workflow. It shows how verification results can guide the decisions to automate the forecast process. Across a range of forecast parameters, informed use of Jive has typically added around 1-2 lead days of skill to weather forecasts in Australia.
Future plans include extending Jive’s capabilities to verify weather warnings and developing tools to assist meteorologists in advising specific industries and emergency services in making optimal decisions. Additionally, with the rise of data-driven (AI) weather prediction models, there will be the need to understand how they can most effectively be used in the forecast production process. The verification core of Jive has been migrated to the open-source python package called scores (https://scores.readthedocs.io/) where new methods continue to be implemented. As forecasting continues to evolve, verification systems like Jive will remain essential drivers of continuous improvement, ensuring weather information effectively serves the community’s needs.
METADATA
AMS: What would you like readers to learn from this article?
Nicholas Loveday (Bureau of Meteorology): I would like readers to understand how much of an impact forecast verification can have in uplifting the quality of forecasts provided by meteorological agencies. Well-designed verification systems with appropriate methods and tools are essential for providing the evidence for an organization to make big decisions around automation, while at the same time enabling weather forecasters to be able to produce the best forecasts possible.
AMS: How did you become interested in the topic of this article?
NL: I started off as an operational weather forecaster around a decade ago, but was disappointed to find that there was no way to easily evaluate my forecasts to see if I was having a positive impact. As soon as I found out that a research project had started to build a verification system to solve this problem, I jumped straight across into that research team.
AMS: What got you initially interested in meteorology or the related field you are in?
NL: My Gran was always interested in weather, and she took daily observations on her farm for several decades. When I was a child, she shared her love of weather with me, often showing me weather documentaries. When I started university, atmospheric science sounded like the most interesting subject to major in, plus it meant potential for a job that would have a huge positive impact on society.
AMS: What surprised you the most about the work you document in your article?
NL: It surprised me that weather forecasters were extremely skeptical of any verification that showed that the automated forecasts did better. At the same time, they loved seeing results that showed that they performed better than the automated guidance. It took a while for them to accept that there were some parameters that the automated forecast provided great forecasts for, and that they were actually degrading the quality of those forecasts by adjusting them. Now they try to only adjust the automated guidance whenever they have clear evidence (usually in the form of verification) that it will lead to a better forecast.
AMS: What was the biggest challenge you encountered while doing this work?
NL: One thing that really challenged us was finding a fair way to evaluate how well the weather forecasts predicted extremes, as there were no suitable methods that we could find in the literature. Fortunately, Rob Taggart (one of the coauthors) developed a really nice approach for evaluating threshold weighted verification scores to evaluate how well the forecasts predicted extremes.
AMS: How will you follow up?
NL: Our research will likely shift more towards verifying warnings as well as developing methods and tools to compare how AI weather models compare to traditional physics-based NWP models. We’ll also spend time building up our new open-source verification package, which will replace the verification core of Jive. Additionally, the team has several performance upgrades planned for the Jive verification system.
