NASA and IBM Open-Sourced a Lunar Foundation Model Trained on 17 Years of Orbiter Data

Seventeen years of looking at the Moon had produced more data than every other NASA planetary mission combined. Making it usable for machine learning was the harder problem.
On October 5, NASA and IBM Research, working with several academic institutions, released the NASA-IBM Lunar Foundation Model. The model is open source, published on Hugging Face with code on GitHub, and integrated into the open-source TerraTorch toolkit. The pretraining datasets and benchmark collections went out with it.
The model cuts polar ice-deposit prediction error by up to 22 percent compared with the best baseline, SwinV2-B, and improves coarse-scale crater detection by nearly 19 percent while using only half the labeled data.
Why a foundation model instead of a task-specific one
Lunar science has an odd data problem. Observations are plentiful. Labels are scarce. Building a model for crater detection at one scale, another for ice stability, and another for surface segmentation means building from scratch each time, on small labeled sets, with no shared representation underneath.
A foundation model flips that. It pretrains on large volumes of unlabeled data, then adapts to a specific task with only a few labeled examples. Kevin Murphy, NASA's chief science data officer, framed the release in exactly those terms: NASA has spent decades building a scientific record of the Moon, and collecting data is only part of the job. The rest is making it easier for scientists to use.
The training corpus is SomBench, described as the largest co-registered multimodal lunar dataset assembled so far: nearly two million tile bundles across eleven modalities and two spatial scales. About one million high-resolution images come from the Narrow Angle Camera at roughly one meter per pixel. Nearly 964,000 multispectral images come from the Wide Angle Camera at 100 meters per pixel. The bulk traces back to the Lunar Reconnaissance Orbiter, supplemented by data from GRAIL, Lunar Prospector, and JAXA's Kaguya probe. Over 30 spatially aligned data layers from nine instruments across four missions.

The design choice that did most of the work
The architecture adapts TerraMind, a multimodal Earth-observation model, but the team trained the lunar version from scratch rather than fine-tuning an existing one. The Moon has no atmosphere and no weather, so assumptions inherited from an Earth model turn worse than useless: actively misleading. On the Moon, how something looks is largely a function of how it is lit.
That observation drove the second key decision. The model receives imaging geometry as explicit context for each tile: illumination angles, sun position, tile extent. Instead of inferring lighting from raw pixels, it is told. One of the more telling findings is that an architecturally identical control model with random initialization still performed competitively on ice prediction, which the researchers took as evidence that the data handling method itself carried a lot of the gain.
FlexiViT handles the patch-size variation, letting the trained model adapt to tasks at different image scales without retraining.
Benchmarks and the limits the authors named
The team tested the model on crater detection at 100-meter and 1-meter scales, polar ice deposit prediction, and segmentation of irregular mare patches. The pretrained model matched or beat common baselines and the random-initialization control. Ice prediction saw the largest gains.
The technical report is explicit about what the model cannot do. It is not suited for absolute geodetic positioning. In generation tests, latitude and longitude were off by dozens of degrees in some cases, and elevation structures appeared with shifted absolute height values even when shape reconstruction was correct. Juan Bernabé-Moreno of IBM Research Europe described the model as a reusable foundation for downstream lunar research rather than a substitute for physical measurements.
That distinction carries through the whole release. The model accelerates analysis: it estimates where water ice could remain stable in permanently shadowed regions, maps craters to estimate surface ages, and flags irregular mare patches for volcanic study. It does not confirm that exploitable resources exist or guarantee that a site is safe to land on.
What it cannot do
The report's own limitations section is worth reading carefully, because it defines the boundary between a research tool and an instrument. The model is not suitable for absolute geodetic positioning: in generation tests, latitude and longitude were off by dozens of degrees in some cases, and elevation structures appeared with shifted absolute height values even when the reconstructed shapes were correct.
Translated into practice, that means the model is useful for finding and characterizing features, and unreliable for saying exactly where they are to a high precision. A mission planner could use it to narrow down candidate regions for polar ice, then confirm with targeted measurements. A team that treated its output as survey-grade coordinates would be misusing it.
The authors frame the release the same way, as a reusable foundation for downstream research rather than a substitute for physical measurement. That honesty about scope is what makes the release usable, because it tells downstream teams which tasks to trust the model with and which to keep doing the hard way.
Why the Earth-model adaptation matters
TerraMind was built for Earth observation, where imagery is shaped by atmosphere, vegetation cycles, and human activity. Adapting its architecture while training from scratch is a deliberate middle path. The team kept what generalizes, which is the machinery for handling multimodal, spatially aligned data, and discarded what does not, which is any assumption about what a surface should look like.
The illumination-as-input decision shows why that distinction matters. On Earth, a satellite image carries enough visual redundancy that a model can often infer lighting from texture and shadow. On the Moon, the surface is largely uniform and the shadow is the signal. Handing the model sun angle and illumination geometry directly, rather than asking it to reconstruct them, let it spend its capacity on the science tasks instead of on re-deriving physics it was already given.
That is a pattern that transfers. Any domain where a physical measurement is easy to compute but expensive for a model to infer is a candidate for explicit conditioning: depth from stereo, atmospheric correction, sensor calibration. The Moon happened to make the case cleanly because it has so few confounding variables.
What open release means here
Open-sourcing a model and its datasets changes who can participate. Research teams outside NASA and IBM can reproduce the published results, fine-tune the model on their own problems, and build applications for future missions without negotiating data access. When labeled lunar data is expensive to produce, a pretrained model that needs only a few examples is the difference between a study that happens and one that does not.
The release is also a template. The Moon was indexed by a model built for a specific kind of imagery, fed the physics that determines what that imagery looks like, and trained on data that had been sitting in archives for seventeen years waiting to be processed. The same pattern applies to any domain with abundant unlabeled observations and scarce labels, which is most of scientific remote sensing.
Related articles
BOSSFIGHT Ran Frontier Models as Coffee Shop Owners for 24 Weeks. Most Lost to Doing Nothing.
Resisting a bribe on a quiz and refusing to bend in a live quarter are two different skills, and current evaluations tend to measure the easier one.
Four Steps Instead of Forty: How Distillation Is Squeezing Open Image Models Onto Consumer GPUs
Low-step distillation is a trade, not a free lunch. Text fidelity, editing precision, and multi-reference consistency are the first things to suffer.
Agents Started Directing Short Films This Week. The Bottleneck Moved to Quality Control.
Generation is cheap, orchestration is the product, and the thing that decides whether a pipeline is useful is whether it can tell when it has failed.
Reka Rho-1 Puts Understanding and Generation in the Same KV Cache
The same weights that predict a camera image also drive robot movement, because both live in the same representational space.