This is an abbreviated version of an article that first appeared in First Break magazine of Feb 2026
Identifying the location of the Earth’s natural resources is critical to energy security and to provide the resources the energy transition requires. Resources such as hydrocarbons, copper, lithium, hydrogen and geothermal energy are often explored for using suites of geological and geophysical data. Machine learning (ML) is transforming the way geological interpretations are conducted by enabling more efficient, accurate, and scalable analysis, extracting patterns and interpretations from combinations of datatypes. Its uses have been demonstrated in geological mapping, seismic interpretation, mineral exploration, as well as environmental and engineering applications (see e.g. Cracknell & Reading (2014), Li et al., (2019), Rodriguez-Galiano et al., (2015), Baghbani et al., (2022) respectively).
One type of ML is forest-based regression, a technique that has been successfully used by Getech to map global surface heat flow in areas away from well and probe measurements. By understanding the geological and geophysical factors that contribute to surface heat flow we can generate a database of variables that might contribute to surface heat flow. Large crustal scale and smaller local-scale variations in heat flow make it challenging to correlate individual variables with heat flow; therefore, a multivariate, ML based approach is sensible.
Another type of ML algorithms helps indicates whether the conditions beneficial for the presence of a particular geological resource at a known location are replicated elsewhere. For this, we supply the model with a set of known locations where a certain occurrence is present, and a broader suite of potentially influential explanatory variables (these are datasets that cover the wider area and are likely to show certain characteristics around the locations of mineral presence). The algorithm then identifies the key relationships between inputs that help predict where similar variable combinations — and thus potentially undiscovered occurrences — may exist. This article focuses on one such approach: Presence-only prediction.
Presence-only prediction
Presence-only prediction is an ML technique that uses a maximum entropy method to predict the probability an occurrence being present at a certain location. The maximum entropy method in machine learning is a way of estimating the most unbiased probability distribution possible, given only the information we actually know. The method was originally developed to model ecological species distribution (Phillips et al., 2006). In that instance there are known geographical locations where species can be observed as being present, but it is not possible to state that there are absence locations with 100% certainty just because species have not been observed yet. The fact that it is not required to provide the model with absence locations (i.e. locations where the occurrence is definitely known not to exist) is a strength of the approach. This makes it ideal for applying to scenarios such as mineral exploration, where ground truth information may be limited in under-explored areas.
The Presence-only prediction modelling is based on training data (hereafter ‘occurrence locations’), which details locations where the presence has already been proven (such as proven mineral deposits), and a series of explanatory variables. The relation between the explanatory variables at the occurrence location is used to devise the model which is subsequently used to predict the probability of the presence at locations away from the original training points.
The study area is discretised into a series of background points, where the possibility of presence is possible, but unknown. Absence is not assumed in any location across the study area, and the method compares the values of the explanatory variables at each occurrence location. The various datasets undergo data preparation, explanatory variable transformation, output data preparation, and model validation.

Figure 1 – Example of a ML-driven workflows using Getech’s Globe data to identify favourable locations for a specific target.
The statistical success of the prediction is related to the threshold applied, which relates to the amount of the unknown area that will be flagged as having the potential for presence. Care must be taken to limit the amount of the unknown areas that are flagged as having the potential for an occurrence, so that the result remains realistic and ranked based on the most likely locations for success. In the example model review shown below, a threshold of 0.5 is declared as the potential for presence, resulting in a model that can successfully predict the locations of the training points for over 80% of the occurrence locations. Lowering the threshold would increase this percentage of the occurrence locations being successfully identified by the model, but at the expense of more of the unknown areas being flagged as having potential for presence.

Figure 2 – The model aims to correctly classify training sites as Presence while minimising false positives among Background points.
The influence of the explanatory variables on the final model can be analysed by looking at the response curves. Each curve shows the effect that changing the values in each explanatory variable has on the presence probability, while keeping all other factors the same. Response curves that show a flat line have little predictive power to indicate the likelihood of a presence occurring. All other curve types show a relationship between that variable and the presence of the known features.

Figure 3 – Example response curves for 4 explanatory variables in a mineral predictive model. The ASTER index for Fe2Ox, MVI susceptibility at 1500m depth and the distance to mapped structures are aiding this example model predict presence away from the known mineral deposits. In this case, the crustal thickness is having little effect.
Mineral exploration potential
Getech has utilised Presence-only prediction methodology for client mineral exploration studies, notably in areas of magmatic arcs, and fold and thrust belts. In these projects a range of mineral systems were analysed using training data from known discoveries, and a set of explanatory datasets was developed that included remote-sensing data alongside Getech’s Globe product and potential fields datasets.
In one mining exploration example study Advanced Spaceborne Thermal Emission and Reflection Radiometer (ASTER) data were processed and levelled for each of the 14 wavelength bands, and 27 derived mineral indices. Sub-surface data were enhanced by conducting voxel-based inversion including 3D magnetic vector inversion (MVI) to model the amplitude and direction of magnetisation at depth. In total ninety-seven explanatory variables were generated.

Figure 4 – Additional explanatory variables for mineral exploration, including Landsat (left), ASTER (centre) and radiometric data (right).

Figure 5 – MVI model showing 3D distribution of magnetisation. Depth slices from these models were important sub-surface explanatory variables.
Across the area of interest, three models were produced, one for each target type (porphyry copper, magmatic minerals and sedimentary exhalative type deposits). In each model different training data was applied to the same area and variables, with the ML placing different weightings on each explanatory variable depending on its use for detecting favourable conditions for each target type. This powerful technique provided insights into geographical areas where multiple data show correlations similar to those observed at identified mineral systems, but also gives insight into which variables are the main influences on the prediction allowing a better understanding of the key datasets supporting the discovery of future mineral deposits. The results of the studies aligned well with areas that had independently been identified for future exploration potential, as well as highlighting previously unconsidered locations that may have high potential.

Figure 6 – Example outputs for the presence-only prediction mapping the locations of likely occurrences of porphyry copper (purple), magmatic minerals (orange) and SedEx type deposits (green).
Summary
Machine learning is becoming a powerful tool in a range of applications, including geoscience. The ability to integrate multiple datatypes that cannot be directly linked through a physics-based approach, and to optimise the recognition of patterns between them without user bias has a wide range of applications, particularly in geological resource identification.
At Getech we have integrated remote-sensing and subsurface datasets along with our Globe datasets from regional and local-scales to successfully provide machine learning based predictions for mineral prospectivity. This approach complements traditional play-based exploration methodologies, and the incorporation of expert validation should not be overlooked (Davies et al., 2025). In addition, this ML approach also provides feedback on the importance of different data types to identify targets, which in turn guides future data collection and exploration strategies.
If you would like to find out more, contact us.
References
Baghbani, A., Choudhury, T., Costa, S., Reiner, J. 2022. Application of artificial intelligence in geotechnical engineering: A state-of-the-art review. Earth-Science Reviews. 228, 103991
Cracknell, M.J., Reading, A.M. 2014. Geological mapping using remote sensing data: A comparison of five machine learning algorithms, their response to variations in the spatial distribution of training data and the use of explicit spatial information. Computers & Geosciences, 63, 22-33
Davies, R.S., Trott, M., Georgi, J., Farrar, A. 2025. Artificial intelligence and machine learning to enhance critical mineral deposit discovery. Geosystems and Geoenvironment, 4(2), 100361
Li, D., Peng, S., Lu, Y., Guo, Y., Cui, X. 2019. Seismic structure interpretation based on machine learning: A case study in coal mining. Interpretation, 7(3) SE69-79.
Phillips, S. J., Anderson, R.P., Schapire, R. E. 2006. Maximum entropy modeling of species geographic distributions. Ecological Modelling, 190, 231-259.
Rodriguez-Galiano, V., Sanchez-Castillo, M., Chica-Olmo, M., Chica-Rivas, M. 2015. Machine learning predictive models for mineral prospectivity: An evaluation of neural networks, random forest, regression trees and support vector machines. Ore Geology Reviews, 71, 804-818
Webb, P., Cheyney, S., Masterton, S., Green, C. 2024. Predicting Heat Flow for Resource Exploration Using Random Forest Regression. Artificial Intelligence for Geological Modelling and Mapping conference, University of Exeter, UK, May 22-23 2024





