With complex, chaotic systems, predictions are notoriously unreliable, especially once several variables start interacting. That was precisely the puzzle at the heart of three days in Bonn, as early-career researchers from around the world came together to bring machine learning and Earth system science into the same room. Could the outcome of this hackathon have been predicted? Hardly — and that is exactly what makes events like this so exciting.
A hackathon for the Earth system
From 19 to 21 August, the Institute of Computer Science at the University of Bonn hosted the second Hackathon on Machine Learning for the Earth System, held directly after the workshop of the same name. A hackathon brings people together for a limited period of time to code, experiment and present results in teams, working on concrete, real-world questions.
And the questions on the table here could hardly be more topical: how do we improve numerical weather predictions with ML? How do we build climate emulators capable of capturing complex Earth system processes? And what will the foundation models of the future look like for the Earth system sciences?


Stronger together: a project of many partners
This event was only possible thanks to the collaboration of several institutions: the Center for Earth System Observation and Computational Analysis (CESOC), the University of Bonn, the University of Cologne and Forschungszentrum Jülich joined forces with the Transdisciplinary Research Area Modelling (TRA) to make it happen. It is precisely this blend of Earth system science, computer science and high-performance computing that makes the hackathon so distinctive — bringing together people who rarely sit at the same table.
Five challenges, five routes into the Earth system
At the heart of the hackathon were five challenges, on which teams worked over the course of three days. Each was led by tutors from different research institutions — a genuine cross-institutional effort.


Challenge 1 – Integrating SST into weather and climate models How can sea surface temperature (SST) be built into a modern, data-driven weather model to better capture long-term warming trends? Teams worked with the WeatherGenerator model, compared two different integration strategies, and developed their own evaluation routines to test them. Tutors: Savvas Melidonis, Jifeng Wang, Jehangir Awan and Ankit Patnala (Forschungszentrum Jülich, JSC)
Challenge 2 – Quantisation techniques Earth observation data is vast — a challenge even for powerful HPC systems. This challenge explored vector quantisation (VQ), Finite Scalar Quantisation (FSQ) and Lookup Free Quantisation (LFQ), techniques that encode data compactly while keeping it meaningful. Tutors: Marieke Wesselkamp, Hakam Shams, Vitus Benson; co-creator: Sebastian Hoffmann (MPI for Biochemistry, Jena)
Challenge 3 – From Lorenz to AI weather models A model can perform brilliantly in the short term and still drift away from realistic behaviour over longer rollouts. Using the famous Lorenz system as a minimal testbed, teams investigated when and why neural emulators behave like chaotic dynamical systems — and when they don’t. Tutor: Dwaipayan Chatterjee (IMKTRO, KIT)
Challenge 4 – AI-based ocean and sea-ice modelling While AI-based atmospheric weather models are advancing rapidly, extending these capabilities to the ocean and cryosphere is still in its early stages. Using GLORYS data and frameworks such as geoarches and anemoi, teams tackled topics from land-masking to the coupling of ocean, sea-ice and atmosphere, and high-resolution simulations. Tutors: Kacper Nowak (AWI), Nils Hutter & Janika Rhein (GEOMAR)
Challenge 5 – Temporal downscaling for extreme weather events Much weather and climate data is only available at coarse temporal resolution — too coarse to properly capture events like heavy rainfall or heatwaves. Focusing on atmospheric rivers in the North Pacific, teams trained models to downscale ERA5 data from 6-hour to 1-hour resolution, using the ExtremeWeatherBench tool for evaluation. Tutors: Sorcha Owens (UKMO) and Manvendra Janmaijaya (Turing Institute); co-creators: Stephen Haddad, Karina Bett-Williams (UKMO)




What really carried these challenges, in the end, was the tutors themselves: they didn’t just guide their topics with real expertise, they also formed the teams and built a genuine sense of team spirit. That was clearly visible in the final presentations — as teams walked through their challenge, their results and their reflections, it was obvious how much shared effort and enthusiasm had gone into the three days.
Input from the best: the lectures
The hackathon was accompanied by a series of lectures:
- Gunjan Joshi (HZDR) – “Earth Observation Foundation Models: From Pixels to Planetary”
- Martin Schultz (Forschungszentrum Jülich, JSC) – “Fundamentals of conventional ML-based weather modelling”
- Matthias Karlbauer (ECMWF) – “Extended-range forecasting with deep learning models”
A mix of foundational knowledge and current research — exactly the right basis for three days of hands-on coding and data work.


International, diverse, engaged
Participants came from all over the world and brought a wide range of backgrounds, from Earth system science to computer science. Particularly pleasing: with a 45% share of women, the hackathon was noticeably more diverse than many comparable tech events.
The backbone of the hackathon was, for the most part, female – organised by Sorcha Owens (UKMO), Florentine Weber (JSC/FZJ), Marieke Wesselkamp (MPI) and Martin Schultz (JSC/FZJ).
And the output?
Not everything went to plan: HPC availability caused some bottlenecks during the event, which occasionally made things harder for the teams. But even the best weather forecasting models — built on huge amounts of data and sophisticated algorithms — are of little use without the computing power to run them. What stays with us, though, is the way of thinking we develop when discussing and designing such models and their architecture — that was the heart of our MLESM26 hackathon.
The teams’ final presentations were correspondingly impressive, and in the end two teams were awarded prizes after a very close race for team spirit, originality and completeness.


Feedback from tutors and participants alike was consistently positive — most said they would take part again in a heartbeat. As organisers, we are delighted by that, and we take it as a clear signal.
Could we have predicted the output at the start? Probably not — but that’s rather the point with chaotic systems, whether in the atmosphere or in a conference room.
See you in 2027 — let’s see what we come up with next.

