Nested sampling algorithm

Template:Bayesian statistics The nested sampling algorithm is a computational approach to the Bayesian statistics problems of comparing models and generating samples from posterior distributions. It was developed in 2004 by physicist John Skilling.^[1]

Background

Bayes' theorem can be applied to a pair of competing models $M_{1}$ and $M_{2}$ for data $D$ , one of which may be true (though which one is unknown) but which both cannot be true simultaneously. The posterior probability for $M_{1}$ may be calculated as:

\begin{matrix} P (M_{1} ∣ D) & = \frac{P (D ∣ M_{1}) P (M_{1})}{P (D)} \\ = \frac{P (D ∣ M_{1}) P (M_{1})}{P (D ∣ M_{1}) P (M_{1}) + P (D ∣ M_{2}) P (M_{2})} \\ = \frac{1}{1 + \frac{P (D ∣ M_{2})}{P (D ∣ M_{1})} \frac{P (M_{2})}{P (M_{1})}} \end{matrix}

The prior probabilities $M_{1}$ and $M_{2}$ are already known, as they are chosen by the researcher ahead of time. However, the remaining Bayes factor $P (D ∣ M_{2}) / P (D ∣ M_{1})$ is not so easy to evaluate, since in general it requires marginalizing nuisance parameters. Generally, $M_{1}$ has a set of parameters that can be grouped together and called $θ$ , and $M_{2}$ has its own vector of parameters that may be of different dimensionality, but is still termed $θ$ . The marginalization for $M_{1}$ is

P (D ∣ M_{1}) = \int d θ P (D ∣ θ, M_{1}) P (θ ∣ M_{1})

and likewise for $M_{2}$ . This integral is often analytically intractable, and in these cases it is necessary to employ a numerical algorithm to find an approximation. The nested sampling algorithm was developed by John Skilling specifically to approximate these marginalization integrals, and it has the added benefit of generating samples from the posterior distribution $P (θ ∣ D, M_{1})$ .^[2] It is an alternative to methods from the Bayesian literature^[3] such as bridge sampling and defensive importance sampling.

Here is a simple version of the nested sampling algorithm, followed by a description of how it computes the marginal probability density $Z = P (D ∣ M)$ where $M$ is $M_{1}$ or $M_{2}$ :

Start with  $N$  points  $θ_{1}, \dots, θ_{N}$  sampled from prior.
for  $i = 1$  to  $j$  do        % The number of iterations j is chosen by guesswork.
     $L_{i} : = \min ($ current likelihood values of the points $)$ ;
     $X_{i} : = \exp (- i / N);$ 
     $w_{i} : = X_{i - 1} - X_{i}$ 
     $Z : = Z + L_{i} \cdot w_{i};$ 
    Save the point with least likelihood as a sample point with weight  $w_{i}$ .
    Update the point with least likelihood with some Markov chain Monte Carlo steps according to the prior, accepting only steps that
    keep the likelihood above  $L_{i}$ .
end
return  $Z$ ;

At each iteration, $X_{i}$ is an estimate of the amount of prior mass covered by the hypervolume in parameter space of all points with likelihood greater than $θ_{i}$ . The weight factor $w_{i}$ is an estimate of the amount of prior mass that lies between two nested hypersurfaces ${θ ∣ P (D ∣ θ, M) = P (D ∣ θ_{i - 1}, M)}$ and ${θ ∣ P (D ∣ θ, M) = P (D ∣ θ_{i}, M)}$ . The update step $Z : = Z + L_{i} w_{i}$ computes the sum over $i$ of $L_{i} w_{i}$ to numerically approximate the integral

\begin{matrix} P (D ∣ M) & = \int P (D ∣ θ, M) P (θ ∣ M) d θ \\ = \int P (D ∣ θ, M) d P (θ ∣ M) \end{matrix}

In the limit $j \to \infty$ , this estimator has a positive bias of order $1 / N$ ^[4] which can be removed by using $(1 - 1 / N)$ instead of the $\exp (- 1 / N)$ in the above algorithm.

The idea is to subdivide the range of $f (θ) = P (D ∣ θ, M)$ and estimate, for each interval $[f (θ_{i - 1}), f (θ_{i})]$ , how likely it is a priori that a randomly chosen $θ$ would map to this interval. This can be thought of as a Bayesian's way to numerically implement Lebesgue integration.^[5]

Choice of MCMC algorithm

The original procedure outlined by Skilling (given above in pseudocode) does not specify what specific Markov chain Monte Carlo algorithm should be used to choose new points with better likelihood.

Skilling's own code examples (such as one in Sivia and Skilling (2006),^[6] available on Skilling's website) chooses a random existing point and selects a nearby point chosen by a random distance from the existing point; if the likelihood is better, then the point is accepted, else it is rejected and the process repeated. Mukherjee et al. (2006)^[7] found higher acceptance rates by selecting points randomly within an ellipsoid drawn around the existing points; this idea was refined into the MultiNest algorithm^[8] which handles multimodal posteriors better by grouping points into likelihood contours and drawing an ellipsoid for each contour.

Implementations

Example implementations demonstrating the nested sampling algorithm are publicly available for download, written in several programming languages.

Simple examples in C, R, or Python are on John Skilling's website.
A Haskell port of the above simple codes is on Hackage.
An example in R originally designed for fitting spectra is described on Bojan Nikolic's website and is available on GitHub.
A NestedSampler is part of the Python toolbox BayesicFitting^[9] for generic model fitting and evidence calculation. It is available on GitHub.
An implementation in C++, named DIAMONDS, is on GitHub.
A highly modular Python parallel example for statistical physics and condensed matter physics uses is on GitHub.
pymatnest is a package designed for exploring the energy landscape of different materials, calculating thermodynamic variables at arbitrary temperatures and locating phase transitions is on GitHub
The MultiNest software package is capable of performing nested sampling on multi-modal posterior distributions.^[8]^[10] It has interfaces for C++, Fortran and Python inputs, and is available on GitHub.
PolyChord is another nested sampling software package available on GitHub. PolyChord's computational efficiency scales better with an increase in the number of parameters than MultiNest, meaning PolyChord can be more efficient for high dimensional problems.^[11] It has interfaces to likelihood functions written in Python, Fortran, C, or C++.
NestedSamplers.jl, a Julia package for implementing single- and multi-ellipsoidal nested sampling algorithms is on GitHub.
Korali is a high-performance framework for uncertainty quantification, optimization, and deep reinforcement learning, which also implements nested sampling.

Applications

Since nested sampling was proposed in 2004, it has been used in many aspects of the field of astronomy. One paper suggested using nested sampling for cosmological model selection and object detection, as it "uniquely combines accuracy, general applicability and computational feasibility."^[7] A refinement of the algorithm to handle multimodal posteriors has been suggested as a means to detect astronomical objects in extant datasets.^[10] Other applications of nested sampling are in the field of finite element updating where the algorithm is used to choose an optimal finite element model, and this was applied to structural dynamics.^[12] This sampling method has also been used in the field of materials modeling. It can be used to learn the partition function from statistical mechanics and derive thermodynamic properties.^[13]

Dynamic nested sampling

Dynamic nested sampling is a generalisation of the nested sampling algorithm in which the number of samples taken in different regions of the parameter space is dynamically adjusted to maximise calculation accuracy.^[14] This can lead to large improvements in accuracy and computational efficiency when compared to the original nested sampling algorithm, in which the allocation of samples cannot be changed and often many samples are taken in regions which have little effect on calculation accuracy.

Publicly available dynamic nested sampling software packages include:

Template:Proper name - a Python implementation of dynamic nested sampling which can be downloaded from GitHub.^[15]
dyPolyChord: a software package which can be used with Python, C++ and Fortran likelihood and prior distributions.^[16] dyPolyChord is available on GitHub.

Dynamic nested sampling has been applied to a variety of scientific problems, including analysis of gravitational waves,^[17] mapping distances in space^[18] and exoplanet detection.^[19]

References

Template:Reflist

[skilling1-1] Template:Cite journal

[skilling2-2] Template:Cite journal

[chen-3] Template:Cite book

[4] Template:Cite journal

[Jasa-5] Template:Cite journal

[6] Template:Cite book

[mukherjee-7] 7.0 ^7.1 Template:Cite journal

[multinest-8] 8.0 ^8.1 Template:Cite journal

[kester-9] Template:Cite journal

[feroz-10] 10.0 ^10.1 Template:Cite journal

[11] Template:Cite journal

[12] Template:Cite journal

[partay-13] Template:Cite journal

[14] Template:Cite journal

[15] Template:Cite journal

[16] Template:Cite journal

[17] Template:Cite journal

[18] Template:Cite journal

[19] Template:Cite journal

[1]

[2]

[3]

[4]

[5]

[6]

[7]

[8]

[9]

[10]

[11]

[12]

[13]

[14]

[15]

[16]

[17]

[18]

[19]

Nested sampling algorithm

Contents

Background

Choice of MCMC algorithm

Implementations

Applications

Dynamic nested sampling

See also

References

Navigation menu

Nested sampling algorithm

Background

Choice of MCMC algorithm

Implementations

Applications

Dynamic nested sampling

See also

References

Navigation menu

Search