Prithvi WxC: Foundation Model for Weather and Climate

Johannes Schmude,Sujit Roy,Will Trojak,Johannes Jakubik,Daniel Salles Civitarese,Shraddha Singh,Julian Kuehnert,Kumar Ankur,Aman Gupta,Christopher E Phillips,Romeo Kienzler,Daniela Szwarcman,Vishal Gaur,Rajat Shinde,Rohit Lal,Arlindo Da Silva,Jorge Luis Guevara Diaz,Anne Jones,Simon Pfreundschuh,Amy Lin,Aditi Sheshadri,Udaysankar Nair,Valentine Anantharaj,Hendrik Hamann,Campbell Watson,Manil Maskey,Tsengdar J Lee,Juan Bernabe Moreno,Rahul Ramachandran
2024-09-20
Abstract:Triggered by the realization that AI emulators can rival the performance of traditional numerical weather prediction models running on HPC systems, there is now an increasing number of large AI models that address use cases such as forecasting, downscaling, or nowcasting. While the parallel developments in the AI literature focus on foundation models -- models that can be effectively tuned to address multiple, different use cases -- the developments on the weather and climate side largely focus on single-use cases with particular emphasis on mid-range forecasting. We close this gap by introducing Prithvi WxC, a 2.3 billion parameter foundation model developed using 160 variables from the Modern-Era Retrospective Analysis for Research and Applications, Version 2 (MERRA-2). Prithvi WxC employs an encoder-decoder-based architecture, incorporating concepts from various recent transformer models to effectively capture both regional and global dependencies in the input data. The model has been designed to accommodate large token counts to model weather phenomena in different topologies at fine resolutions. Furthermore, it is trained with a mixed objective that combines the paradigms of masked reconstruction with forecasting. We test the model on a set of challenging downstream tasks namely: Autoregressive rollout forecasting, Downscaling, Gravity wave flux parameterization, and Extreme events estimation. The pretrained model with 2.3 billion parameters, along with the associated fine-tuning workflows, has been publicly released as an open-source contribution via Hugging Face.
Machine Learning,Atmospheric and Oceanic Physics
What problem does this paper attempt to address?
### What problems does this paper attempt to solve? This paper aims to solve the problem that existing deep - learning models in the field of meteorological and climate prediction are too specialized. Specifically, most of the existing deep - learning models mainly focus on specific tasks, such as short - term forecasting, downscaling or nowcasting, and these models are usually not well - adapted to other related tasks. To fill this gap, the author introduced Prithvi WxC, a foundation model with 2.3 billion parameters. This model is trained on the MERRA - 2 data set of 160 variables and adopts an encoder - decoder architecture, combining multiple latest Transformer model concepts. Prithvi WxC is designed to be able to effectively handle regional and global dependencies under different topologies and support a large number of tokens to simulate weather phenomena at a fine resolution. In addition, Prithvi WxC is trained using a hybrid objective function, combining the two paradigms of mask reconstruction and prediction. This enables the model to be tested in multiple challenging downstream tasks, including autoregressive rolling prediction, downscaling, gravity - wave flux parameterization, and extreme - event estimation. In this way, Prithvi WxC can not only perform well in a variety of meteorological and climate applications, but also provides an extensible and flexible framework that can be fine - tuned for different specific tasks. This versatility is lacking in existing specialized models, so Prithvi WxC is expected to become a powerful tool in the field of meteorological and climate science. #### Specific problem summary: 1. **Multi - task adaptability**: Most of the existing deep - learning models focus on a single task, while Prithvi WxC aims to be a foundation model that can adapt to multiple tasks. 2. **Capturing complex dependencies**: Prithvi WxC can better capture regional and global dependencies in the input data through its encoder - decoder architecture and Transformer model concepts. 3. **Large - scale data processing ability**: Prithvi WxC is designed to be able to process a large number of tokens, thereby modeling weather phenomena in different spatial contexts. 4. **Hybrid - objective training**: By combining the objective functions of mask reconstruction and prediction, Prithvi WxC can simultaneously learn different features of the data during the training process. 5. **Downstream - task verification**: Prithvi WxC has been verified in multiple downstream tasks to ensure its effectiveness in practical applications. I hope this summary can help you understand the core problems of this paper and their solutions. If you have more questions or need further information, please feel free to let me know!