NASA-IBM Lunar Foundation Model Opens 17 Years of Moon Data

NASA-IBM Lunar Foundation Model Opens 17 Years of Moon Data

NASA and IBM Research have released an open-source AI model built for studying the Moon. The NASA-IBM Lunar Foundation Model was developed with several academic institutions, and the two organizations call it one of the first open foundation models aimed at lunar science. It performs best at predicting polar ice deposits and detecting craters.

Kevin Murphy, NASA's chief science data officer, put the problem plainly. NASA has built a huge scientific record of the Moon over decades, he said, but gathering data is only half the work. Scientists also need to be able to use it.

Why a foundation model

Task-specific algorithms are trained for one job. A foundation model is pretrained on large amounts of unlabeled data and can then be adapted to new tasks with only a handful of labeled examples. Lunar research has plenty of observation data and very few labels, so the team sees this as the approach's biggest advantage.

The training data

The model was trained from scratch on SomBench, which the team describes as the largest co-registered multimodal lunar corpus so far. It holds nearly 2 million tile bundles covering 11 modalities at two spatial scales:

  • About 1 million high-resolution images from the Narrow Angle Camera (NAC), at roughly 1 meter per pixel
  • Just under 964,000 multispectral images from the Wide Angle Camera, at 100 meters per pixel

Most of the data comes from 17 years of observations by NASA's Lunar Reconnaissance Orbiter (LRO). According to NASA, LRO's data volume is larger than that of all its other planetary missions combined. The team added gravity data from GRAIL, hydrogen data from Lunar Prospector and mineralogy data from Kaguya/SELENE, a probe run by Japan's space agency JAXA.

The result is more than 30 spatially aligned data layers from nine instruments and four missions. To stop test data from leaking into training, the corpus was split geographically by map zone instead of shuffling individual tiles at random.

Telling the model about the light

The architecture is based on TerraMind, a multimodal Earth observation model, though the lunar version was trained from scratch rather than fine-tuned. One design choice stands out. On the Moon, lighting geometry affects how the surface looks more than the surface's real properties do. So instead of making the model work this out from raw pixels, the team gives it illumination angles, sun position and tile extent as explicit input.

The model also learns from fine and coarse imagery in one training run. A technique called FlexiViT lets the same trained model handle different image patch sizes without retraining.

Where it wins, and where it doesn't

The team tested the model on crater detection at 100-meter and 1-meter scales, on polar ice prediction, and on segmenting Irregular Mare Patches (IMPs). These are young volcanic features that complicate accepted timelines of how the Moon cooled. According to the technical report, the pretrained model matched or beat common baselines and an identical, randomly initialized control model on every task.

Ice prediction showed the largest gain. Permanently shadowed polar regions can preserve water ice for billions of years and are seen as a possible source of water, oxygen and rocket fuel. IBM says the model reduced prediction error by up to 22 percent compared with the best baseline, SwinV2-B. On coarse crater detection it beat SwinV2-B by nearly 19 percent while trained on only half the data.

The picture is less clear elsewhere. Even the untrained control model beat five of seven baselines on ice prediction. The researchers attribute part of this to architecture: each data layer gets its own processing path, while baselines stack all inputs as channels. On meter-scale craters and IMPs, results roughly tie the strongest baselines, within run-to-run variance. IBM claims a 3 percent lead on IMPs, but the numbers look comparable rather than clearly better.

LoRA, a lighter fine-tuning method that trains only a fraction of the parameters, kept up with full fine-tuning and did better on crater detection.

Known limits

The model is not suited for absolute geodetic positioning. In generation tests, latitude and longitude were off by dozens of degrees in some cases, and elevation could appear at shifted heights even when shapes were correct. The authors frame it as a base for downstream tasks, not a substitute for physical instruments. Controlled experiments isolating each design choice are still pending, and some test sets are small.

Juan Bernabé-Moreno, director of IBM Research Europe, said the model links observations across instruments and surfaces patterns that are hard to see in isolation.

The model is on Hugging Face, the code is on GitHub, and it is integrated into the open-source TerraTorch toolkit. The pretraining datasets and benchmarks are public too.

Our Take

This release fits a steady pattern. NASA and IBM have worked on foundation models under a Space Act Agreement since early 2022, starting with the Prithvi Earth observation model in 2023. Google DeepMind's AlphaEarth Foundations follows a similar idea for Earth. It also sits alongside other efforts to make AI a practical research tool for scientists, and broader interest in AI and space.

The honest reporting of limits is useful. The strongest result, ice prediction, appears to come partly from architecture rather than pretraining alone, which suggests the design lessons may matter as much as the weights. It is worth watching whether the promised ablation studies confirm which parts actually drive the gains, and whether outside lunar researchers adopt the model for real mission planning.