<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">AMT</journal-id><journal-title-group>
    <journal-title>Atmospheric Measurement Techniques</journal-title>
    <abbrev-journal-title abbrev-type="publisher">AMT</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Atmos. Meas. Tech.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1867-8548</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/amt-11-4627-2018</article-id><title-group><article-title>A neural network approach to estimating a posteriori distributions <?xmltex \hack{\break}?> of Bayesian retrieval problems</article-title><alt-title>Bayesian retrievals using neural
networks</alt-title>
      </title-group><?xmltex \runningtitle{Bayesian retrievals using neural
networks}?><?xmltex \runningauthor{S.~Pfreundschuh et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Pfreundschuh</surname><given-names>Simon</given-names></name>
          <email>simon.pfreundschuh@chalmers.se</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Eriksson</surname><given-names>Patrick</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-8475-0479</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Duncan</surname><given-names>David</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-1955-4391</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Rydberg</surname><given-names>Bengt</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Håkansson</surname><given-names>Nina</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-2929-1617</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Thoss</surname><given-names>Anke</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Department of Space, Earth and Environment, Chalmers University of Technology, Gothenburg, Sweden</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Möller Data Workflow Systems AB, Gothenburg, Sweden</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Swedish Meteorological and Hydrological Institute (SMHI), Norrköping, Sweden</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Simon Pfreundschuh (simon.pfreundschuh@chalmers.se)</corresp></author-notes><pub-date><day>9</day><month>August</month><year>2018</year></pub-date>
      
      <volume>11</volume>
      <issue>8</issue>
      <fpage>4627</fpage><lpage>4643</lpage>
      <history>
        <date date-type="received"><day>26</day><month>March</month><year>2018</year></date>
           <date date-type="rev-request"><day>29</day><month>March</month><year>2018</year></date>
           <date date-type="rev-recd"><day>26</day><month>June</month><year>2018</year></date>
           <date date-type="accepted"><day>28</day><month>June</month><year>2018</year></date>
      </history>
      <permissions>
        
        
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018.html">This article is available from https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018.html</self-uri><self-uri xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018.pdf">The full text article is available as a PDF file from https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018.pdf</self-uri>
      <abstract>
    <p id="d1e141">A neural-network-based method, quantile regression neural networks (QRNNs), is
proposed as a novel approach to estimating the a posteriori distribution of
Bayesian remote sensing  retrievals. The advantage of QRNNs over conventional
neural network retrievals is that they learn to predict not only a single
retrieval value but also the associated, case-specific uncertainties. In this
study, the retrieval performance of QRNNs is characterized and compared to
that of other state-of-the-art retrieval methods. A synthetic retrieval
scenario is presented and used as a validation case for the application of
QRNNs to Bayesian retrieval problems. The QRNN retrieval performance is
evaluated against Markov chain Monte Carlo simulation and another Bayesian
method based on Monte Carlo integration over a retrieval database. The
scenario is also used to investigate how different hyperparameter
configurations and training set sizes affect the retrieval performance. In the
second part of the study, QRNNs are applied to the retrieval of cloud top
pressure from observations by the Moderate Resolution Imaging
Spectroradiometer (MODIS). It is shown that QRNNs are not only capable of
achieving similar accuracy to standard neural network retrievals but also
provide statistically consistent uncertainty estimates for non-Gaussian
retrieval errors. The results presented in this work show that QRNNs are able
to combine the flexibility and computational efficiency of the machine
learning approach with the theoretically sound handling of uncertainties of
the Bayesian framework. Together with this article, a Python implementation of
QRNNs is released through a public repository to make the method available to
the scientific community.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <title>Introduction</title>
      <p id="d1e151">The retrieval of atmospheric quantities from remote sensing measurements
constitutes an inverse problem that generally does not admit a unique, exact
solution. Measurement and modeling errors, as well as limited sensitivity of the
observation system, preclude the assignment of a single, discrete solution to a
given observation. A meaningful retrieval should thus consist of a retrieved
value and an estimate of uncertainty describing a range of values that are
likely to produce a measurement similar to the one observed. However, even
if a retrieval method allows for explicit modeling of retrieval uncertainties,
their computation and representation are often possible only in an approximate
manner.</p>
      <?pagebreak page4628?><p id="d1e154">The Bayesian framework provides a formal way of handling the ill-posedness of
the retrieval problem and its associated uncertainties. In the Bayesian
formulation <xref ref-type="bibr" rid="bib1.bibx32" id="paren.1"/>, the solution of the inverse problem is given by
the a posteriori distribution <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, i.e., the conditional
distribution of the retrieval quantity <inline-formula><mml:math id="M2" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> given the observation <inline-formula><mml:math id="M3" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula>.
Under the modeling assumptions, the posterior distribution represents all
available knowledge about the retrieval quantity <inline-formula><mml:math id="M4" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> after the measurement,
accounting for all considered retrieval uncertainties. Bayes' theorem states
that the a posteriori distribution is proportional to the product <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of the a priori distribution <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and the conditional probability
of the observed measurement <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The a priori distribution
<inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents knowledge about the quantity <inline-formula><mml:math id="M9" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> that is available before
the measurement and can be used to aid the retrieval with supplementary
information. <?xmltex \hack{\newpage}?></p>
      <p id="d1e280">For a given retrieval, the a posteriori
distribution can generally not be expressed in closed form, and different
methods have been developed to compute approximations to it. In cases that
allow a sufficiently precise and efficient simulation of the measurement, a
forward model can be used to guide the solution of the inverse problem. If
such a forward model is available, the most general technique to compute the
a posteriori distribution is Markov chain Monte Carlo (MCMC) simulation. MCMC
denotes a set of methods that iteratively generate a sequence of samples,
whose sampling distribution approximates the true a posteriori distribution.
MCMC simulations have the advantage of allowing the estimation of the a
posteriori distribution without requiring any simplifying assumptions on a
priori knowledge, measurement error or the forward model. The
disadvantage of MCMC simulation is that each retrieval requires a
high number of forward-model evaluations, which in many cases makes the
method computationally too demanding to be practical. For
remote sensing retrievals, the
method is therefore of interest rather for testing and validation
<xref ref-type="bibr" rid="bib1.bibx35" id="paren.2"/>, such as in the retrieval algorithm developed by
<xref ref-type="bibr" rid="bib1.bibx8" id="text.3"/>.</p>
      <p id="d1e289">A method that avoids costly forward-model evaluations during the retrieval
has been proposed by <xref ref-type="bibr" rid="bib1.bibx21" id="text.4"/>. The method is based on Monte Carlo
integration of importance-weighted samples in a retrieval database
<inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, which consists of pairs of observations
<inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and corresponding values <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of the retrieval quantity. The
method will be referred to in the following as Bayesian Monte Carlo
integration (BMCI). Even though the method is less computationally demanding
than methods involving forward-model calculations during the retrieval, it
may require the traversal of a potentially large retrieval database.
Furthermore, the incorporation of ancillary data to aid the retrieval
requires careful stratification of the retrieval database, as it is performed
in the Goddard profiling algorithm <xref ref-type="bibr" rid="bib1.bibx22" id="paren.5"/> for the retrieval of
precipitation profiles. Further applications of the method can be found for
example in the work by <xref ref-type="bibr" rid="bib1.bibx33" id="text.6"/> or <xref ref-type="bibr" rid="bib1.bibx8" id="text.7"/>.</p>
      <p id="d1e364">The optimal estimation method <xref ref-type="bibr" rid="bib1.bibx32" id="paren.8"/>, in short OEM (also 1D-Var, for
one-dimensional variational retrieval), simplifies the Bayesian retrieval
problem assuming that a priori knowledge and measurement uncertainty both follow
Gaussian distributions and that the forward model is only moderately nonlinear.
Under these assumptions, the a posteriori distribution is approximately
Gaussian. The retrieved values in this case are the mean and maximum of the a
posteriori distribution, which coincide for a Gaussian distribution, together
with the covariance matrix describing the width of the a posteriori
distribution. In cases where an efficient forward model for the computation of
simulated measurements and corresponding Jacobians is available, the OEM has
become the quasi-standard method for Bayesian retrievals. Nonetheless, even
neglecting the validity of the assumptions of Gaussian a priori and measurement
errors as well as linearity of the forward model, the method is unsuitable for
retrievals that involve complex radiative processes. In particular, since the
OEM requires the computation of the Jacobian of the forward model,
processes such as surface or cloud scattering become too expensive to model
online during the retrieval.</p>
      <p id="d1e370">Compared to the Bayesian retrieval methods discussed above, machine learning
provides a more flexible approach to learning computationally efficient retrieval
mappings directly from data. Large amounts of data available from simulations,
collocated observations or in situ measurements, as well as increasing computational
power to speed up the training, have made machine learning techniques
an attractive alternative to approaches based on (Bayesian) inverse modeling.
Numerous applications of machine learning regression methods to retrieval
problems can be found in the recent literature <xref ref-type="bibr" rid="bib1.bibx17 bib1.bibx15 bib1.bibx34 bib1.bibx38 bib1.bibx14 bib1.bibx3" id="paren.9"/>.
All of these examples, however, neglect the probabilistic character of the
inverse problem and provide only a scalar estimate of the retrieval. Uncertainty
estimates in these retrievals are provided in the form of mean errors computed
on independent test data, which is a clear drawback compared to Bayesian
methods. A notable exception is the work by <xref ref-type="bibr" rid="bib1.bibx1" id="text.10"/>,
which applies the Bayesian framework to estimate errors in the retrieved
quantities due to uncertainties on the learned neural network parameters.
However, the only difference to the approaches listed above is that the
retrieval errors, estimated from the error covariance matrix observed on the
training data, are corrected for uncertainties in the network parameters. With
respect to the intrinsic retrieval uncertainties, the approach is thus afflicted
with the same limitations. Furthermore, the complexity of the required numerical
operations makes it suitable only for small training sets and simple networks.</p>
      <p id="d1e379">In this article, quantile regression neural networks (QRNNs) are proposed as a
method to use neural networks to estimate the a posteriori distribution of
remote sensing retrievals. Originally proposed by <xref ref-type="bibr" rid="bib1.bibx20" id="text.11"/>,
quantile regression is a method for fitting statistical models to quantile
functions of conditional probability distributions. Applications of quantile
regression using neural networks <xref ref-type="bibr" rid="bib1.bibx5" id="paren.12"/> and other machine learning
methods <xref ref-type="bibr" rid="bib1.bibx24" id="paren.13"/> exist, but to the best knowledge of the authors this
is the first application of QRNNs to remote sensing retrievals. The aim of this
work is to combine the flexibility and computational efficiency of the machine
learning approach with the theoretically sound handling of uncertainties in the
Bayesian framework.</p>
      <p id="d1e391">A formal description of QRNNs and the retrieval methods against which they
will be evaluated is provided in Sect. <xref ref-type="sec" rid="Ch1.S2"/>. A simulated
retrieval scenario is used to validate the approach against BMCI and MCMC in
Sect. <xref ref-type="sec" rid="Ch1.S3"/>. Section <xref ref-type="sec" rid="Ch1.S4"/> presents the application of
QRNNs to the retrieval of cloud top pressure and associated uncertainties
from satellite observations in the<?pagebreak page4629?> visible and infrared. Finally, the
conclusions from this work are presented in Sect. <xref ref-type="sec" rid="Ch1.S5"/>.</p>
</sec>
<sec id="Ch1.S2">
  <title>Methods</title>
      <p id="d1e408">This section introduces the Bayesian retrieval formulation and the retrieval
methods used in the subsequent experiments. Two Bayesian methods, Markov chain
Monte Carlo simulation and Bayesian Monte Carlo integration, are presented.
Quantile regression neural networks are introduced as a machine learning
approach to estimating the a posteriori distribution of Bayesian retrieval
problems. The section closes with a discussion of the statistical metrics that
are used to compare the methods.</p>
<sec id="Ch1.S2.SS1">
  <title>The retrieval problem</title>
      <p id="d1e416">The general problem considered here is the retrieval of a scalar quantity <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>∈</mml:mo><mml:mi mathvariant="normal">R</mml:mi></mml:mrow></mml:math></inline-formula> from an indirect measurement given in the form of an observation
vector <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="normal">R</mml:mi><mml:mi>m</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>. In the Bayesian framework, the retrieval
problem is formulated as finding the posterior distribution <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of
the quantity <inline-formula><mml:math id="M16" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> given the measurement <inline-formula><mml:math id="M17" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula>. Formally, this solution can be
obtained by application of Bayes' theorem:
            <disp-formula id="Ch1.E1" content-type="numbered"><mml:math id="M18" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>∫</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          The a priori distribution <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the knowledge about the quantity <inline-formula><mml:math id="M20" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>
that is available prior to the measurement. The a priori knowledge introduced
into the retrieval formulation regularizes the ill-posed inverse problem and
ensures that the retrieval solution is physically meaningful. The a posteriori
distribution of a scalar retrieval quantity <inline-formula><mml:math id="M21" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> can be represented by the
corresponding cumulative distribution function (CDF) <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>,
which is defined as
            <disp-formula id="Ch1.E2" content-type="numbered"><mml:math id="M23" display="block"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mi>x</mml:mi></mml:munderover><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mspace width="0.33em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S2.SS2">
  <title>Bayesian retrieval methods</title>
      <p id="d1e664">Bayesian retrieval methods are methods that use the expression for the a
posteriori distribution in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) to compute a solution to the
retrieval problem. Since the a posteriori distribution can generally not be
computed or sampled directly, these methods approximate the posterior
distribution to varying degrees of accuracy.</p>
<sec id="Ch1.S2.SS2.SSS1">
  <title>Markov chain Monte Carlo</title>
      <p id="d1e674">MCMC simulation denotes a set of methods for the
generation of samples from arbitrary posterior distributions <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
The general principle is to compute samples from an approximate distribution and
refine them in a way such that their distribution converges to the true a
posteriori distribution <xref ref-type="bibr" rid="bib1.bibx10" id="paren.14"/>. In this study, the Metropolis algorithm is
used to implement MCMC. The Metropolis algorithm iteratively generates a
sequence of states <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi></mml:mrow></mml:math></inline-formula> using a symmetric proposal
distribution <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In each step of the algorithm, a
proposal <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> for the next step is generated by sampling from
<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The proposed state <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is accepted as
the next simulation step <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with probability <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:mtext>min</mml:mtext><mml:mfenced close="}" open="{"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>. Otherwise
<inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is rejected and the current simulation step <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is kept
for <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. If the proposal distribution <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mi>J</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is
symmetric and samples generated from it satisfy the Markov chain property with a
unique stationary distribution, the Metropolis algorithm is guaranteed to
produce a distribution of samples which converges to the true a posteriori
distribution.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <title>Bayesian Monte Carlo integration</title>
      <p id="d1e945">The BMCI method is based on the use of importance sampling to approximate
integrals over the a posteriori distribution of a given retrieval case. Consider an
integral of the form
              <disp-formula id="Ch1.E3" content-type="numbered"><mml:math id="M36" display="block"><mml:mrow><mml:mo movablelimits="false">∫</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.33em"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
            Applying Bayes' theorem, the integral can be written as
              <disp-formula id="Ch1.Ex1"><mml:math id="M37" display="block"><mml:mrow><mml:mo movablelimits="false">∫</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.33em"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mo movablelimits="false">∫</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>∫</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup><mml:mo>)</mml:mo><mml:mspace width="0.33em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.33em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
            The last integral can be approximated by a sum over an observation
database <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> that is distributed according
to the a priori distribution <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>:
              <disp-formula id="Ch1.Ex2"><mml:math id="M40" display="block"><mml:mrow><mml:mo movablelimits="false">∫</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mspace linebreak="nobreak" width="0.33em"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>≈</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>C</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            with the normalization factor <inline-formula><mml:math id="M41" display="inline"><mml:mi>C</mml:mi></mml:math></inline-formula> given by <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></inline-formula>
The weights <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are given by  the probability <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
of the observed measurement <inline-formula><mml:math id="M45" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> conditional on the database
measurement <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which is usually assumed to be multivariate
Gaussian with covariance matrix <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">S</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>:
              <disp-formula id="Ch1.Ex3"><mml:math id="M48" display="block"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>∝</mml:mo><mml:mi>exp⁡</mml:mi><mml:mfenced open="{" close="}"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:msubsup><mml:mi mathvariant="bold">S</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
            By approximating integrals of Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>), it is possible
to estimate the expectation value and variance of the a posteriori distribution by
choosing <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="script">E</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, respectively.
While this is suitable to represent Gaussian distributions, a more general
representation of the a posteriori distribution can be obtained by
estimating the corresponding CDF (cf. Eq. <xref ref-type="disp-formula" rid="Ch1.E2"/>) using
              <disp-formula id="Ch1.E4" content-type="numbered"><mml:math id="M51" display="block"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>≈</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>C</mml:mi></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<?pagebreak page4630?><sec id="Ch1.S2.SS3">
  <title>Machine learning</title>
      <p id="d1e1569">Neglecting uncertainties, the retrieval of a quantity <inline-formula><mml:math id="M52" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> from a measurement
vector <inline-formula><mml:math id="M53" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> may be viewed as a simple multiple regression task. In
machine learning, regression problems are typically approached by training a
parametrized model <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>:</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>↦</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:math></inline-formula> to predict a desired output
<inline-formula><mml:math id="M55" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> from given input <inline-formula><mml:math id="M56" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula>. Unfortunately, the use of the variables
<inline-formula><mml:math id="M57" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M58" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> in machine learning is directly opposite to their use
in inverse theory. For the remainder of this section the variables <inline-formula><mml:math id="M59" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula>
and <inline-formula><mml:math id="M60" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> will be used to denote, respectively, the input and output to
the machine learning model to ensure consistency with the common notation in
the field of machine learning. The reader must keep in mind that the method
is applied in the later sections to predict a retrieval quantity <inline-formula><mml:math id="M61" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> from a
measurement <inline-formula><mml:math id="M62" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula>.</p>
<sec id="Ch1.S2.SS3.SSS1">
  <title>Supervised learning and loss functions</title>
      <p id="d1e1664">Machine learning regression models are trained using supervised training, in
which the model <inline-formula><mml:math id="M63" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> learns the regression mapping from a training set
<inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> with input values <inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and expected
output values <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The training is performed by finding model parameters
that minimize the mean of a given loss function <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi mathvariant="script">L</mml:mi><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> on the training set. The most common loss function for regression
tasks is the squared error loss
              <disp-formula id="Ch1.E5" content-type="numbered"><mml:math id="M68" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="normal">se</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            which trains the model <inline-formula><mml:math id="M69" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> to minimize the mean squared distance of the
neural network prediction <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> from the expected output <inline-formula><mml:math id="M71" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> on
the training set. If the estimand
<inline-formula><mml:math id="M72" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> is a random vector drawn from a conditional probability
distribution <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, a regressor trained using a squared
error loss function learns to predict the conditional expectation value of
the distribution <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> <xref ref-type="bibr" rid="bib1.bibx2" id="paren.15"/>. Depending on the
choice of the loss function, the regressor can also learn to predict other
statistics of the distribution <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> from the training data.</p>
</sec>
<sec id="Ch1.S2.SS3.SSS2">
  <title>Quantile regression</title>
      <p id="d1e1920">Given the cumulative distribution function <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of a probability distribution
<inline-formula><mml:math id="M77" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>, its <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> quantile <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is defined as

                  <disp-formula id="Ch1.E6" content-type="numbered"><mml:math id="M80" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mo movablelimits="false">inf⁡</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mi>x</mml:mi><mml:mspace linebreak="nobreak" width="0.33em"/><mml:mo>:</mml:mo><mml:mspace width="0.33em" linebreak="nobreak"/><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>≥</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo mathvariant="italic">}</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            i.e., the greatest lower bound of all values of <inline-formula><mml:math id="M81" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> for which <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>≥</mml:mo><mml:mi mathvariant="italic">τ</mml:mi></mml:mrow></mml:math></inline-formula>. As shown by <xref ref-type="bibr" rid="bib1.bibx19" id="text.16"/>, the <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>th</mml:mtext></mml:mrow></mml:math></inline-formula> quantile <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of
<inline-formula><mml:math id="M85" display="inline"><mml:mi>F</mml:mi></mml:math></inline-formula> minimizes the expectation value <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">E</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:msubsup><mml:mo>∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mi mathvariant="normal">∞</mml:mi></mml:msubsup><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mspace width="0.33em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> of the function

                  <disp-formula id="Ch1.E7" content-type="numbered"><mml:math id="M87" display="block"><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mfenced open="{" close=""><mml:mtable class="cases" rowspacing="0.2ex" columnspacing="1em" columnalign="left left" framespacing="0em"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>&lt;</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>)</mml:mo><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>otherwise</mml:mtext><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e2245">By training a machine learning regressor <inline-formula><mml:math id="M88" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> to minimize the mean of the quantile loss
function <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> over a training set <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, the regressor learns to predict the quantiles of the
conditional distribution <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. This can be extended to obtain an
approximation of the cumulative distribution function of <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> by
training the network to estimate multiple quantiles of <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S2.SS3.SSS3">
  <title>Neural networks</title>
      <p id="d1e2379">A neural network computes a vector of output activations <inline-formula><mml:math id="M94" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> from a vector
of input activations <inline-formula><mml:math id="M95" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula>. Feed-forward artificial neural networks (ANNs)
compute the vector <inline-formula><mml:math id="M96" display="inline"><mml:mi mathvariant="bold-italic">y</mml:mi></mml:math></inline-formula> by application of a given number of subsequent,
learnable transformations to the input activations <inline-formula><mml:math id="M97" display="inline"><mml:mi mathvariant="bold-italic">x</mml:mi></mml:math></inline-formula>:

                  <disp-formula specific-use="align"><mml:math id="M98" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold">W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

              The activation functions <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> as well as the number and sizes of the hidden
layers <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are prescribed, structural
parameters of a neural network model, generally referred to as
hyperparameters. The learnable parameters of the model are the weight
matrices <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">W</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and bias vectors <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of each layer.
Neural networks can be efficiently trained in a supervised manner by using
gradient-based minimization methods to find suitable weights <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">W</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>
and bias vectors <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. By using the mean of the quantile loss
function <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> as the training criterion, a neural network can
be trained to predict the quantiles of the distribution <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, thus turning the network into a quantile regression neural
network.</p>
</sec>
<sec id="Ch1.S2.SS3.SSS4">
  <title>Adversarial training</title>
      <p id="d1e2617">Adversarial training is a data augmentation technique that has been proposed
to increase the robustness of neural networks to perturbations in the input
data <xref ref-type="bibr" rid="bib1.bibx13" id="paren.17"/>. It has been shown to be effective also as a method
to improve the calibration of probabilistic predictions from neural networks
<xref ref-type="bibr" rid="bib1.bibx23" id="paren.18"/>. The basic principle of adversarial training is to
augment the training data with perturbed samples that are likely to
yield a large change in the network prediction. The method used here to
implement adversarial training  is the fast gradient sign method
proposed by <xref ref-type="bibr" rid="bib1.bibx13" id="text.19"/>. For a training sample <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
consisting of input <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="normal">R</mml:mi><mml:mi>n</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> and expected output
<inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="normal">R</mml:mi><mml:mi>m</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, the corresponding adversarial sample
<inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is chosen to be

                  <disp-formula id="Ch1.E8" content-type="numbered"><mml:math id="M111" display="block"><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>⋅</mml:mo><mml:mtext>sign</mml:mtext><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="script">L</mml:mi><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>;</mml:mo></mml:mrow></mml:math></disp-formula>

            i.e., the direction of the perturbation is chosen in such a way that
it maximizes the absolute change in the loss function <inline-formula><mml:math id="M112" display="inline"><mml:mi mathvariant="script">L</mml:mi></mml:math></inline-formula> due
to an infinitesimal change in the input parameters. The adversarial
perturbation factor <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> determines the strength of the perturbation
and becomes an additional hyperparameter of the neural network model.</p>
</sec>
</sec>
<?pagebreak page4631?><sec id="Ch1.S2.SS4">
  <title>Evaluating probabilistic predictions</title>
      <p id="d1e2812">A problem that remains is how to compare two estimates <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mo>′</mml:mo><mml:mo>′</mml:mo></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> of a given a posteriori distribution against a single observed sample
<inline-formula><mml:math id="M115" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> from the true distribution <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. A good
probabilistic prediction for the value <inline-formula><mml:math id="M117" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> should be sharp, i.e., concentrated in
the vicinity of <inline-formula><mml:math id="M118" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>, but at the same time well calibrated, i.e., predicting
probabilities that truthfully reflect observed frequencies <xref ref-type="bibr" rid="bib1.bibx12" id="paren.20"/>.
Summary measures for the evaluation of predicted conditional distributions are
called scoring rules <xref ref-type="bibr" rid="bib1.bibx11" id="paren.21"/>. An important property of scoring rules
is propriety, which formalizes the concept of the scoring rule rewarding both
sharpness and calibration of the prediction. Besides providing reliable
measures for the comparison of probabilistic predictions, proper scoring rules
can be used as loss functions in supervised learning to incentivize
statistically consistent predictions.</p>
      <p id="d1e2903">The quantile loss function given in Eq. (<xref ref-type="disp-formula" rid="Ch1.E7"/>) is a proper
scoring rule for quantile estimation and can thus be used to compare the
skill of different methods for quantile estimation <xref ref-type="bibr" rid="bib1.bibx11" id="paren.22"/>. Another
proper scoring rule for the evaluation of an estimated cumulative
distribution function <inline-formula><mml:math id="M119" display="inline"><mml:mi>F</mml:mi></mml:math></inline-formula> against an observed value <inline-formula><mml:math id="M120" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> is the continuous
ranked probability score (CRPS):

                <disp-formula id="Ch1.E9" content-type="numbered"><mml:math id="M121" display="block"><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>CRPS</mml:mtext><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mi mathvariant="normal">∞</mml:mi></mml:munderover><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mi>F</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>≤</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mspace width="0.33em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          Here, <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>≤</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the indicator function that is equal to 1 when
the condition <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>≤</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is true and <inline-formula><mml:math id="M124" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> otherwise.
For the methods used in this article the integral can only be evaluated
approximately. The exact way in which this is done for each method is
described in detail in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1.SSS3"/> and
<xref ref-type="sec" rid="Ch1.S3.SS1.SSS4"/>.</p>
      <p id="d1e3045">The scoring rules presented above evaluate probabilistic predictions against a
single observed value. However, since MCMC simulations can be used to
approximate the true a posteriori distribution to an arbitrary degree of
accuracy, the probabilistic predictions obtained from BMCI and QRNN can be
compared directly to the a posteriori distributions obtained using MCMC. In
the idealized case where the modeling assumptions underlying the MCMC
simulations are true, the sampling distribution obtained from MCMC will
converge to the true posterior and can be used as a ground truth to assess the
predictions obtained from the other methods.</p>
<sec id="Ch1.S2.SS4.SSS1">
  <title>Calibration plots</title>
      <p id="d1e3053">Calibration plots are a graphical method for assessing the calibration of
prediction intervals derived from probabilistic predictions. For a set of
prediction intervals with probabilities <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the fraction
of cases for which the true value did lie within the bounds of the interval
is plotted against the value <inline-formula><mml:math id="M126" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>. If the predictions are well calibrated, the
probabilities <inline-formula><mml:math id="M127" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> match the observed frequencies and the calibration curve is
close to the diagonal <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:math></inline-formula>. An example of a calibration plot for three
different predictors is given in Fig. <xref ref-type="fig" rid="Ch1.F1"/>.
Compared to the scoring rules described above, the advantage of the
calibration curves is that they indicate whether the predicted intervals are
too narrow or too wide. Predictions that overestimate the uncertainty yield
intervals that are too wide and result in a calibration curve that lies above
the diagonal, whereas observations underestimating the uncertainty will yield
a calibration curve that lies below the diagonal.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><caption><p id="d1e3112">Example of a calibration plot displaying calibration curves for
overly confident predictions (dark gray), well-calibrated predictions (red)
and overly cautious predictions (blue).</p></caption>
            <?xmltex \igopts{width=227.622047pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f01.pdf"/>

          </fig>

</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <title>Application to a synthetic retrieval case</title>
      <p id="d1e3129">In this section, a simulated retrieval of column water vapor (CWV) from passive
microwave observations is used to benchmark the performance of BMCI and QRNN
against MCMC simulation. The retrieval case has been set up to provide an
idealized but realistic scenario in which the true a posteriori distribution can
be approximated using MCMC simulation. The MCMC results can therefore be used as
the reference to investigate the retrieval performance of QRNNs and BMCI. Furthermore
the influence of different hyperparameters on the performance of the QRNN, as
well as how the size of the training set and retrieval database
impact the performance of QRNNs and BMCI, is investigated.</p>
<sec id="Ch1.S3.SS1">
  <title>The retrieval</title>
      <p id="d1e3137">For this experiment, the retrieval of CWV from passive
microwave observations over the ocean is considered. The state<?pagebreak page4632?> of the
atmosphere is represented by profiles of temperature and water vapor
concentrations on 15 pressure levels between <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and 10 <inline-formula><mml:math id="M130" display="inline"><mml:mi mathvariant="normal">hPa</mml:mi></mml:math></inline-formula>. The
variability of these quantities has been estimated based on ECMWF ERA-Interim
data <xref ref-type="bibr" rid="bib1.bibx7" id="paren.23"/> from the year 2016, restricted to latitudes
between <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:mn mathvariant="normal">23</mml:mn><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:mn mathvariant="normal">66</mml:mn><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> N. Parametrizations of the multivariate
distributions of temperature and water vapor were obtained by fitting a joint
multivariate normal distribution to the temperature and the logarithm of
water vapor concentrations. The fitted distribution represents the a priori
knowledge on which the simulations are based.</p>
<sec id="Ch1.S3.SS1.SSS1">
  <title>Forward-model simulations</title>
      <p id="d1e3190">The Atmospheric Radiative Transfer Simulator (ARTS; <xref ref-type="bibr" rid="bib1.bibx4" id="altparen.24"/>) is used to
simulate satellite observations of the atmospheric states sampled from the a
priori distribution. The observations consist of simulated brightness
temperatures from five channels around 23, 88, 165 and 183 GHz
(cf. Table <xref ref-type="table" rid="Ch1.T1"/>) of the ATMS sensor.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1"><caption><p id="d1e3201">Observation channels used for the synthetic retrieval of column
water vapor.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Channel</oasis:entry>
         <oasis:entry colname="col2">Center frequency</oasis:entry>
         <oasis:entry colname="col3">Offset</oasis:entry>
         <oasis:entry colname="col4">Bandwidth</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M133" display="inline"><mml:mn mathvariant="normal">23.8</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M134" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M135" display="inline"><mml:mn mathvariant="normal">270</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M136" display="inline"><mml:mi mathvariant="normal">MHz</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M137" display="inline"><mml:mn mathvariant="normal">88.2</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M138" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M139" display="inline"><mml:mn mathvariant="normal">500</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M140" display="inline"><mml:mi mathvariant="normal">MHz</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M141" display="inline"><mml:mn mathvariant="normal">165.5</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M142" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M143" display="inline"><mml:mn mathvariant="normal">300</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M144" display="inline"><mml:mi mathvariant="normal">MHz</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M145" display="inline"><mml:mn mathvariant="normal">183.3</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M146" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">7 <inline-formula><mml:math id="M147" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M148" display="inline"><mml:mn mathvariant="normal">2000</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M149" display="inline"><mml:mi mathvariant="normal">MHz</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M150" display="inline"><mml:mn mathvariant="normal">183.3</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M151" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">3 <inline-formula><mml:math id="M152" display="inline"><mml:mi mathvariant="normal">GHz</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M153" display="inline"><mml:mn mathvariant="normal">1000</mml:mn></mml:math></inline-formula> <inline-formula><mml:math id="M154" display="inline"><mml:mi mathvariant="normal">MHz</mml:mi></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e3446">The simulations take into account only absorption and emission from water
vapor. Ocean surface emissivities are computed using the FASTEM-6 model
<xref ref-type="bibr" rid="bib1.bibx18" id="paren.25"/> with an assumed surface wind speed of zero. The sea surface
temperature is assumed to be equal to the temperature at the pressure level
closest to the surface but no lower than 270 <inline-formula><mml:math id="M155" display="inline"><mml:mi mathvariant="normal">K</mml:mi></mml:math></inline-formula>. Sensor
characteristics and absorption lines are taken from the ATMS sensor
descriptions that are provided within the ARTS XML data package. Simulations
are performed for a nadir-looking sensor and neglecting polarization. The
observation uncertainty is assumed to be independent Gaussian noise with a
standard deviation of 1 <inline-formula><mml:math id="M156" display="inline"><mml:mi mathvariant="normal">K</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S3.SS1.SSS2">
  <title>MCMC implementation</title>
      <p id="d1e3472">The MCMC retrieval is based on a Python implementation of the Metropolis
algorithm <xref ref-type="bibr" rid="bib1.bibx10" id="paren.26"><named-content content-type="post">chap. 12</named-content></xref> that has been developed within the context of
this study. It is released as part of the <italic>typhon: tools for atmospheric research</italic> software package <xref ref-type="bibr" rid="bib1.bibx36" id="paren.27"/>.</p>
      <p id="d1e3486">The MCMC retrieval is performed in the space of atmospheric states described
by the profiles of temperature and the logarithm of water vapor
concentrations. The multivariate Gaussian distribution that has been obtained
by fit to the ERA-Interim data is taken as the a priori distribution. A random
walk is used as the proposal distribution, with its covariance matrix taken as
the a priori covariance matrix. A single MCMC retrieval consists of eight
independent runs, initialized with different random states sampled from the a
priori distribution. Each run starts with a warm-up phase followed by an
adaptive phase during which the covariance matrix of the proposal distribution
is scaled adaptively to keep the acceptance rate of proposed states close to
the optimal 21 % <xref ref-type="bibr" rid="bib1.bibx10" id="paren.28"/>. This is followed by a production phase during
which 5000 samples of the a posteriori distribution are generated. Only 1 out
of 20 generated samples is kept in order to decrease the correlation between
the resulting states. Convergence of each simulation is checked by computing
the scale reduction factor <inline-formula><mml:math id="M157" display="inline"><mml:mover accent="true"><mml:mi>R</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> and the effective number of independent
samples. The retrieval is accepted only if the scale reduction factor is
smaller than 1.1 and the effective sample size larger than 100. Each MCMC
retrieval generates a sequence of atmospheric states from which the column
water vapor is obtained by integration of the water vapor concentration
profile. The distribution of observed CWV values is then taken as the
retrieved a posteriori distribution.</p>
</sec>
<sec id="Ch1.S3.SS1.SSS3">
  <title>QRNN implementation</title>
      <p id="d1e3508">The implementation of quantile regression neural networks is based on the
Keras Python package for deep learning <xref ref-type="bibr" rid="bib1.bibx6" id="paren.29"/>. It is also released
as part of the typhon package.</p>
      <p id="d1e3514">For the training of quantile regression neural networks, the quantile loss
function <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> has been implemented so that it can be
used as a training loss function within the Keras framework. The function can
be initialized with a sequence of quantile fractions <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> allowing the neural network to learn to predict the corresponding
quantiles <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e3593">Custom data generators have been added to the implementation  to incorporate
information on measurement uncertainty into the training
process. If the training data are noise free, the data generator can be used to
add noise to each training batch according to the assumptions on measurement
uncertainty. The noise is added immediately  before the data are passed to the
neural network, keeping the original training data noise free. This ensures
that the network does not see the same noisy training sample twice during
training, thus counteracting overfitting.</p>
      <p id="d1e3596">An adaptive form of stochastic batch gradient descent is used for the neural
network training. During the training, loss is monitored on a validation set.
When the loss on the validation set has not decreased for a certain number of
epochs, the training rate is reduced by a given reduction factor. The training
stops when a predefined minimum learning rate is reached.</p>
      <p id="d1e3600">The reconstruction of the CDF from the estimated quantiles is obtained
by using the quantiles as nodes of a piecewise linear approximation and
extending the first and last<?pagebreak page4633?> segments out to 0 and 1, respectively.
This approximation is also used to compute the CRPS on the test
data.</p>
</sec>
<sec id="Ch1.S3.SS1.SSS4">
  <title>BMCI implementation</title>
      <p id="d1e3610">The BMCI method has likewise been implemented in Python and added to the
typhon package. In addition to retrieving the first two moments of the
posterior distribution, the implementation provides functionality to
retrieve the posterior CDF using Eq. (<xref ref-type="disp-formula" rid="Ch1.E4"/>). Approximate posterior
quantiles are computed by interpolating the inverse CDF at the desired quantile
values. To compute the CRPS for a given retrieval, the trapezoidal rule
is used to perform the integral over the values <inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in the retrieval database
<inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS2">
  <title>QRNN model selection</title>
      <p id="d1e3665">Just as with common neural networks, QRNNs have several hyperparameters that cannot
be learned directly from the data but need to be tuned independently. For this
study the dependence of the QRNN performance on its hyperparameters has been
investigated. The results are included here as they may be a helpful reference
for future applications of QRNNs.</p>
      <p id="d1e3668">For this analysis, hyperparameters describing the structure of the QRNN model
are investigated separately from training parameters. The hyperparameters
describing the structure of the QRNN are
<list list-type="order"><list-item>
      <p id="d1e3673">the number of hidden layers,</p></list-item><list-item>
      <p id="d1e3677">the number of neurons per layer,</p></list-item><list-item>
      <p id="d1e3681">the type of activation function.</p></list-item></list>
The training method  described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1.SSS3"/> is
defined by the following training parameters:
<list list-type="custom"><list-item><label>4.</label>
      <p id="d1e3689">the batch size used for stochastic batch gradient descent,</p></list-item><list-item><label>5.</label>
      <p id="d1e3693">the minimum learning rate at which the training is stopped,</p></list-item><list-item><label>6.</label>
      <p id="d1e3697">the learning rate decay factor,</p></list-item><list-item><label>7.</label>
      <p id="d1e3701">the number of training epochs without progress on the validation set
before the learning rate is reduced.</p></list-item></list></p>
<sec id="Ch1.S3.SS2.SSS1">
  <title>Structural parameters</title>
      <p id="d1e3709">To investigate the influence of hyperparameters 1–3 on the performance of
the QRNN, 10-fold cross validation on the training set consisting of <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>
samples has been used to estimate the performance of different hyperparameter
configurations. As a performance metric, the mean quantile loss on the validation
set averaged over all predicted quantiles for <inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.95</mml:mn></mml:mrow></mml:math></inline-formula> is used. A grid search over a subspace of the configuration space
was performed to find optimal parameters. The results of the analysis are
displayed in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. For the configurations considered,
the layer width has the most significant effect on the performance.
Nevertheless, only small performance gains are obtained by increasing the
layer width to values above 64 neurons. Another general observation is that
networks with three hidden layers generally outperform networks with fewer
hidden layers. Networks using rectified linear unit (ReLU) activation
functions not only achieve slightly better performance than networks using
tanh or sigmoid activation functions but also show significantly lower
variability. Based on these results, a neural network with three hidden
layers, 128 neurons in each layer and ReLU activation functions has been
selected for the comparison to BMCI.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><caption><p id="d1e3760">Mean validation set loss (solid lines) and standard deviation (shading)
of different hyperparameter configurations with respect to layer width (number of neurons).
Different lines display the results for different numbers of hidden layers <inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi mathvariant="normal">h</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.
The three panels show the results for ReLU, tanh and sigmoid activation functions.</p></caption>
            <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f02.pdf"/>

          </fig>

</sec>
<sec id="Ch1.S3.SS2.SSS2">
  <title>Training parameters</title>
      <p id="d1e3786">For the optimization of training parameters 4–7, a very coarse grid
search was performed, using only three different values for each parameter.
In general, the training parameters showed only little effect (<inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> % for
the combinations considered here) on the QRNN performance compared to the
structural parameters. The best cross-correlation performance was obtained
for slow training with a small learning rate reduction factor of <inline-formula><mml:math id="M167" display="inline"><mml:mn mathvariant="normal">1.5</mml:mn></mml:math></inline-formula> and
decreasing the learning rate only after <inline-formula><mml:math id="M168" display="inline"><mml:mn mathvariant="normal">10</mml:mn></mml:math></inline-formula> training epochs without
reduction of the validation loss. No significant increase in performance
could be observed for values of the learning rate minimum below <inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>.
With respect to the batch size, the best results were obtained for a batch
size of 128 samples.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS3">
  <title>Comparison against MCMC</title>
      <p id="d1e3834">In this section, the performance of a single QRNN and an ensemble of 10 QRNNs
is analyzed. The predictions from the ensemble are obtained by averaging the
predictions from each network in the ensemble. All tests in this subsection are
performed for a single QRNN, the ensemble of QRNNs and BMCI. The retrieval database used
for BMCI and the training of the QRNNs in this experiment consists of <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> entries.</p>
      <p id="d1e3848">Figure <xref ref-type="fig" rid="Ch1.F3"/> displays retrieval results for eight example cases. The
choice of the cases is based on the Kolmogorov–Smirnov (KS) statistic, which
corresponds to the maximum absolute deviation of the predicted CDF from the
reference CDF obtained by MCMC simulation. A small KS value indicates a good
prediction of the true CDF, while a high value is obtained for large deviations
between predicted and reference CDF. The cases shown correspond to the 10th, 50th,
90th and 99th percentile of the distribution of KS values obtained using BMCI
or a single QRNN. In this way they provide a qualitative overview of the performance
of the methods.</p>
      <p id="d1e3853">In the displayed cases, both methods are generally successful in predicting the
a posteriori distribution. Only for the <inline-formula><mml:math id="M171" display="inline"><mml:mn mathvariant="normal">99</mml:mn></mml:math></inline-formula>th percentile of the KS value distribution
does the BMCI prediction show significant deviations from the reference<?pagebreak page4634?> distribution.
The jumps in the estimated a posteriori CDF indicate that the deviations are due to
undersampling of the input space in the retrieval database. This results in
excessively high weights attributed to the few entries close to the
observation. For this specific case the QRNN provides a better estimate of the a
posteriori CDF even though both predictions are based on the same data.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><caption><p id="d1e3865">Retrieved a posteriori CDFs obtained using MCMC (gray), BMCI
(blue), a single QRNN (red line) and an ensemble of QRNNs (red marker). Cases
displayed in the first row correspond to the 1st, 50th, 90th and 99th
percentiles of the distribution of the Kolmogorov–Smirnov statistic of BMCI
compared to the MCMC reference. The second row displays the same percentiles of
the distribution of the Kolmogorov–Smirnov statistic of the single-QRNN
predictions compared to MCMC.</p></caption>
          <?xmltex \igopts{width=469.470472pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f03.pdf"/>

        </fig>

      <p id="d1e3875">Another way of displaying the estimated a posteriori distribution is by means
of its probability density function (PDF), which is defined as the derivative
of its CDF. For the QRNN, the PDF is approximated by simply deriving
the piecewise linear approximation to the CDF and setting the boundary values
to zero. For BMCI, the a posteriori PDF can be approximated using a histogram of the
CWV values in the database weighted by the corresponding weights <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
The PDFs for the cases corresponding to the CDFs show in
Fig. <xref ref-type="fig" rid="Ch1.F3"/> are shown in Fig. <xref ref-type="fig" rid="Ch1.F4"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><caption><p id="d1e3901">Retrieved a posteriori PDFs corresponding to the CDFs displayed
in Fig. <xref ref-type="fig" rid="Ch1.F3"/> obtained using MCMC (gray), BMCI (blue), a single
QRNN (red line) and an ensemble of QRNNs (red marker).</p></caption>
          <?xmltex \igopts{width=483.69685pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f04.pdf"/>

        </fig>

      <p id="d1e3912">To obtain a more comprehensive view on the performance of QRNNs and BMCI,
the predictions obtained from both methods are compared to those obtained
from MCMC for 6500 test cases. For the comparison, let the <italic>effective quantile fraction</italic> <inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>eff</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> be defined as the fraction of MCMC
samples that are less than or equal to the predicted quantile
<inline-formula><mml:math id="M174" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="true" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> obtained from QRNN or BMCI. In general, the predicted
quantile <inline-formula><mml:math id="M175" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:math></inline-formula> will not correspond exactly to the true quantile
<inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> but rather to an effective quantile <inline-formula><mml:math id="M177" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>eff</mml:mtext></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, defined by
the fraction <inline-formula><mml:math id="M178" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>eff</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> of the samples of the distribution that are
smaller than or equal to the predicted value <inline-formula><mml:math id="M179" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal" stretchy="true">^</mml:mo></mml:mover></mml:math></inline-formula>. The
resulting distributions of the effective quantile fractions for BMCI and
QRNNs are displayed in Fig. <xref ref-type="fig" rid="Ch1.F5"/> for the estimated
quantiles for <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn></mml:mrow></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><caption><p id="d1e4037">Distribution of effective quantile fractions <inline-formula><mml:math id="M181" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>eff</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> achieved by
QRNN and BMCI on the test data. Panel <bold>(a)</bold> displays the performance of a
single QRNN compared to BMCI; panel <bold>(b)</bold> displays the performance of the ensemble.</p></caption>
          <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f05.pdf"/>

        </fig>

      <p id="d1e4063">For an ideal estimator of the quantile <inline-formula><mml:math id="M182" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the resulting distribution
would be a delta function centered at <inline-formula><mml:math id="M183" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>. Due to the estimation error,
however, the <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">τ</mml:mi><mml:mtext>eff</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> values are distributed around the true quantile
fraction <inline-formula><mml:math id="M185" display="inline"><mml:mi mathvariant="italic">τ</mml:mi></mml:math></inline-formula>. The results show that both BMCI and QRNN provide fairly
accurate estimates of the quantiles of the a posterior distribution. Furthermore,
all methods  yield equally good predictions, making the distributions virtually
identical.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <title>Training set size impact</title>
      <p id="d1e4109">Finally, we investigate how the size of the training data set used in the training
of the QRNN (or as a retrieval database for BMCI) affects the performance of the
retrieval method. This has been done by randomly generating training subsets
from the original training data with sizes logarithmically spaced between <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">3</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>
and <inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> samples. For each size, five random training subsets have been
generated and used to retrieve the test data with a single QRNN and BMCI. As test data,
a separate test set consisting of <inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">5</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> simulated observation vectors and
corresponding CWV values is used.</p>
      <p id="d1e4145">Figure <xref ref-type="fig" rid="Ch1.F6"/> displays the means of the mean absolute percentage
error (MAPE, panel a) and the mean CRPS
(panel b) achieved by both methods on the differently sized training sets. For
the computation of the MAPE, the CWV prediction is taken as the median of the
estimated a posteriori distribution obtained using QRNNs or BMCI. This value
is compared to the true CWV value corresponding to the atmospheric state that
has been used in the simulation. As expected, the performance of both methods
improves with the size of the training set. With respect to the MAPE, both
methods perform equally well for a training set size of <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>, but the QRNN
outperforms BMCI for all smaller training set sizes. With respect to CRPS, a
similar behavior is observed. These are reassuring results, as they indicate
that not only the accuracy of the predictions (measured by the MAPE and CRPS)
improves as the amount of training data increases, but also their calibration
(measured only by the CRPS).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><caption><p id="d1e4163">MAPE <bold>(a)</bold> and CRPS <bold>(b)</bold> achieved by QRNN (red) and BMCI (blue)
on the test set using differently sized training sets and retrieval
databases. For each size, five random subsets of the original training data were
generated. The lines display the means of the observed values. The shading
indicates the range of <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula> around the mean.</p></caption>
          <?xmltex \igopts{width=327.206693pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f06.pdf"/>

        </fig>

      <p id="d1e4188">Finally, the mean of the quantile loss <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">L</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> on the test set for
<inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, 0.5 and 0.9 has been considered (Fig. <xref ref-type="fig" rid="Ch1.F7"/>).
Qualitatively, the results are similar to the ones obtained<?pagebreak page4635?> using MAPE and
CRPS. The QRNN outperforms BMCI for smaller training set sizes but converges
to similar
values for training set sizes of <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><caption><p id="d1e4230">Mean quantile loss for different training set sizes <inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mtext>train</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, 0.5 and 0.9.</p></caption>
          <?xmltex \igopts{width=441.017717pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f07.pdf"/>

        </fig>

      <p id="d1e4262">The results presented in this section indicate that QRNNs can, at least
under idealized conditions, be used to estimate the a posteriori distribution of
Bayesian retrieval problems. Moreover, they were shown to work equally well
as BMCI for large data sets. What is interesting is that, for smaller data sets,
QRNNs even provide better estimates of the a posteriori distribution than BMCI.
This indicates that QRNNs provide a better representation of the functional
dependency of the a posteriori distribution on the observation data, thus
achieving better interpolation in the case of scarce training data. Nonetheless,
it remains to be investigated if this advantage can also be observed for
real-world data.</p>
      <p id="d1e4265">A possible approach to handling scarce retrieval databases with BMCI is to
artificially increase the assumed measurement uncertainty. This has not been
performed for the BMCI results presented here and may improve the performance of
the method. The difficulty with this approach is that the method formulation
is based on the assumption of a sufficiently large database and thus can,
at least formally, not handle scarce training data. Finding a suitable way to
increase the measurement uncertainty would thus require either additional
methodological development or invention of a heuristic approach, both of which
are outside the scope of this study.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <title>Retrieving cloud top pressure from MODIS using QRNNs</title>
      <p id="d1e4275">In this section, QRNNs are applied to retrieve cloud top pressure (CTP) using
observations from the Moderate Resolution Imaging Spectroradiometer (MODIS;
<xref ref-type="bibr" rid="bib1.bibx30" id="altparen.30"/>). The experiment is based on the work by <xref ref-type="bibr" rid="bib1.bibx14" id="text.31"/>,
who developed the NN-CTTH algorithm, a neural-network-based retrieval of cloud top pressure. A QRNN-based CTP
retrieval is compared to the NN-CTTH algorithm, and how QRNNs can be used to
estimate the retrieval uncertainty is investigated.</p>
<sec id="Ch1.S4.SS1">
  <title>Data</title>
      <p id="d1e4289">The QRNN uses the same data for training as the reference NN-CTTH algorithm.
The data set consists of MODIS Level 1B data <xref ref-type="bibr" rid="bib1.bibx25 bib1.bibx26" id="paren.32"/>
collocated with cloud properties obtained from CALIOP (Cloud-Aerosol Lidar with Orthogonal Polarization; <xref ref-type="bibr" rid="bib1.bibx39" id="altparen.33"/>). The
<italic>top layer pressure</italic> variable from the CALIOP data is used as a
retrieval target. The data were taken from all orbits from 24 days (the 1st
and 14th of every month) from the year 2010. In <xref ref-type="bibr" rid="bib1.bibx14" id="text.34"/> multiple
neural networks are trained using varying combinations of input features
derived from different MODIS channels and ancillary NWP data in order to
compare retrieval performance for different inputs. Of the different neural
network configurations presented in <xref ref-type="bibr" rid="bib1.bibx14" id="text.35"/>, the version denoted by
NN-AVHRR (development version of the NN-CTTH algorithm
which uses only channels available from the Advanced Very High Resolution Radiometer
(AVHRR)) is used for comparison against the QRNN. This version uses only the
11 and 12 <inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> channels from MODIS. In addition to single pixel
input, the input features comprise structural information in the form of
various statistics computed on a <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> neighborhood around the center
pixel. The ancillary numerical weather prediction (NWP) data provided to the
network consist of surface pressure and temperature,<?pagebreak page4637?> temperatures at five
pressure levels and column-integrated water vapor. These are also the input
features that are used for the training of the QRNN. The training data used
for the QRNN are the <italic>training</italic> and <italic>during-training validation set</italic> from <xref ref-type="bibr" rid="bib1.bibx14" id="text.36"/>. The comparison to the NN-AVHRR version of the
NN-CTTH algorithm uses the data set for <italic>testing under development</italic>
from <xref ref-type="bibr" rid="bib1.bibx14" id="text.37"/>.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <title>Training</title>
      <p id="d1e4352">The same training scheme as described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS1.SSS3"/>
is used for the training of the QRNNs. The training progress, based on which
the learning rate is reduced or training aborted, is monitored using the
during-training validation data set from <xref ref-type="bibr" rid="bib1.bibx14" id="text.38"/>. After
performing a grid search (results not shown) over width, depth and minibatch
size, the best performance on the validation set was obtained for networks
with four layers with 64 neurons each, ReLU activation functions and a batch
size of 128 samples.</p>
      <p id="d1e4360">The main difference in the training process compared to the experiment from
the previous section is how measurement uncertainties are incorporated. For
the simulated retrieval, the training data was noise free, so measurement
uncertainties could be realistically represented by adding noise according to
the sensor characteristics. This is not the case for MODIS observations;
instead, adversarial training is used here to ensure well-calibrated
predictions. For the tuning of the perturbation parameter
<inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> (cf. Sect. <xref ref-type="sec" rid="Ch1.S2.SS3.SSS4"/>), the
calibration on the during-training validation set was monitored using a
calibration plot. Ideally, it would be desirable to use a separate data set
to tune this parameter, but this was sufficient in this case to achieve good
results on the test data. The calibration curves obtained using different
values of <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> are displayed in
Fig. <xref ref-type="fig" rid="Ch1.F8"/>. It can be seen from the plot that
without adversarial training (<inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) the predictions
obtained from the QRNN are overly confident, leading to prediction intervals
that underrepresent the uncertainty in the retrieval. Since adversarial
training may be viewed as a way of representing<?pagebreak page4638?> observation uncertainty in
the training data, larger values of <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> lead to less
confident predictions. Based on these results, <inline-formula><mml:math id="M202" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> is
chosen for the training.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8"><caption><p id="d1e4433">Calibration of the QRNN prediction intervals on the validation set
used during training. The curves display the results for no adversarial training
(<inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) and adversarial training with perturbation
factor <inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">δ</mml:mi><mml:mtext>adv</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.01</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=213.395669pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f08.pdf"/>

        </fig>

      <p id="d1e4480">Except for the use of adversarial training, the structure of the underlying
network and the training process of the QRNN are fairly similar to what is used
for the NN-CTTH retrieval. The QRNN uses four instead of two hidden layers with
64 neurons in each of them instead of 30 in the first and 15 in second layer.
While this makes the neural network used in the QRNN slightly more complex, this
should not be a major drawback since computational performance is generally not
critical for neural network retrievals.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <title>Prediction accuracy</title>
      <p id="d1e4489">Most data analysis will likely require a single predicted value for the cloud top
pressure. To derive a point value from the QRNN prediction, the median of the
estimated a posteriori distribution is used.</p>
      <p id="d1e4492">The distributions of the resulting median pressure values on the
<italic>testing-during-development</italic> data set are displayed in
Fig. <xref ref-type="fig" rid="Ch1.F9"/> together with the retrieved pressure values
from the NN-CTTH algorithm. The distributions are displayed separately for
low, medium and high clouds (as classified by the CALIOP feature
classification flag) as well as the complete data set. From these results it
can be seen that the values predicted by the QRNN have stronger peaks low in
the atmosphere for low clouds and high in the atmosphere for high clouds. For
medium clouds the peak is more spread out and has heavier tails low and high
in the atmosphere than the values retrieved by the NN-CTTH algorithm.</p>
      <p id="d1e4500">Figure <xref ref-type="fig" rid="Ch1.F10"/> displays the error distributions of the predicted
CTP values on the testing-during-development data set, again separated by
cloud type as well as for the complete data set. Both the simple QRNN and the ensemble
of QRNNs perform slightly better than the NN-CTTH algorithm for low and high clouds.
For medium clouds, no significant difference in the performance of the methods can
be observed. The ensemble of QRNNs seems to slightly improve upon the prediction
accuracy of a single QRNN, but the difference is likely negligible. Compared to
the QRNN results, the CTP predicted by NN-CTTH is biased low for low clouds and
biased high for high clouds.</p>
      <p id="d1e4505">Even though both the QRNN and the NN-CTTH retrieval use the same input and
training data, the predictions from both retrievals differ considerably. Using
the Bayesian framework, this can likely be explained by the fact that the two
retrievals estimate different statistics of the a posteriori distribution. The
NN-CTTH algorithm has been trained using a squared error loss function which
will lead the algorithm to predict the mean of the a posteriori distribution.
The QRNN retrieval, on the other hand, predicts the median of the a posteriori
distribution. Since the median minimizes the expected absolute error, it is
expected that the CTP values predicted by the QRNN yield overall smaller errors.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><caption><p id="d1e4511">Distributions of predicted CTP values <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext>CTP</mml:mtext><mml:mtext>pred</mml:mtext></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
for high clouds <bold>(a)</bold>, medium clouds <bold>(b)</bold>, low
clouds <bold>(c)</bold> and the complete test set <bold>(d)</bold>.</p></caption>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f09.pdf"/>

        </fig>

      <?xmltex \floatpos{p}?><fig id="Ch1.F10" specific-use="star"><caption><p id="d1e4549">Error distributions of predicted CTP values <inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mtext>CTP</mml:mtext><mml:mtext>pred</mml:mtext></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>
with respect to CTP from CALIOP (<inline-formula><mml:math id="M207" display="inline"><mml:mrow><mml:msub><mml:mtext>CTP</mml:mtext><mml:mtext>ref</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) for high clouds <bold>(a)</bold>,
medium clouds <bold>(b)</bold>, low clouds <bold>(c)</bold> and the
complete test set <bold>(d)</bold>.</p></caption>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f10.pdf"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS4">
  <title>Uncertainty estimation</title>
      <p id="d1e4604">The NN-CTTH algorithm retrieves CTP but does not provide case-specific
uncertainty estimates. Instead, an estimate of uncertainty is provided in the
form of the observed mean absolute error (MAE) on the test set. In order to compare
these uncertainty estimates with those obtained using QRNNs, Gaussian error
distributions are fitted to the observed error based on the observed MAE and mean squared error (MSE). A Gaussian error model
is chosen here as it is arguably the most common distribution used to represent
random errors.</p>
      <p id="d1e4607">A plot of the errors observed on the testing-during-development data
set and the fitted Gaussian error distributions is displayed in panel a of
Fig. <xref ref-type="fig" rid="Ch1.F11"/>. The fitted error curves correspond to the Gaussian
probability density functions with the same MAE and MSE as observed on the
test data. Panel b displays the observed error together with the predicted
error obtained from a single QRNN. The predicted error is computed as the
deviation of a random sample of the estimated a posteriori distribution from
its median. The fitted Gaussian error distributions clearly do not provide a
good fit to the observed error. On the other hand, the predicted errors
obtained from the QRNN a posteriori distributions yield good agreement with
the observed error. This indicates that the QRNN successfully learned to
predict retrieval uncertainties. Furthermore, the results show that the
ensemble of QRNNs actually provides a slightly worse fit to the observed
error than a single QRNN. An ensemble of QRNNs thus does not necessarily
improve the calibration of the predictions.</p>
      <?pagebreak page4639?><p id="d1e4612">The Gaussian error model based on the MAE fit has also been used to produce
prediction intervals for the CTP values obtained from the NN-CTTH algorithm.
Figure <xref ref-type="fig" rid="Ch1.F12"/> displays the resulting calibration curves for the
NN-CTTH algorithm, a simple QRNN and an ensemble of QRNNs. The results support
the finding that a single QRNN is able to provide well-calibrated probabilistic
predictions of the a posteriori distribution. The calibration curve for the
ensemble predictions is virtually identical to that for the single network. The
NN-CTTH predictions using a Gaussian fit are not as well calibrated and tend to
provide prediction intervals that are too wide for <inline-formula><mml:math id="M208" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, 0.3, 0.5 and 0.7 but
overly narrow intervals for <inline-formula><mml:math id="M209" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.9</mml:mn></mml:mrow></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{p}?><fig id="Ch1.F11" specific-use="star"><caption><p id="d1e4643">Predicted and observed error distributions. Panel <bold>(a)</bold>
displays the observed error for the NN-CTTH retrieval as well as the
Gaussian error distributions that have been fitted to the observed
error distribution based on the MAE and MSE. Panel <bold>(b)</bold>
displays the observed test set error for a single QRNN as well as the
predicted error obtained as the deviation of a random sample of the
predicted a posteriori distribution from the median. Panel <bold>(c)</bold> displays
the same for the ensemble of QRNNs.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f11.pdf"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12"><caption><p id="d1e4664">Calibration plot for prediction intervals derived from the Gaussian
error model for the NN-CTTH algorithm (blue), the single QRNN (dark gray) and
the ensemble of QRNNs (red).</p></caption>
          <?xmltex \igopts{width=227.622047pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f12.pdf"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS5">
  <title>Sensitivity to a priori distribution</title>
      <p id="d1e4679">As shown above, the predictions obtained from the QRNN are statistically
consistent in the sense that they predict probabilities that match observed
frequencies when applied to test data. This, however, requires that the test data
are statistically consistent with the training data. Statistically consistent
here means that both data sets come from the same generating distribution or, in
more Bayesian terms, the same a priori distribution. What happens when this is
not the case can be seen when the calibration with respect to different cloud
types is computed. Figure <xref ref-type="fig" rid="Ch1.F13"/> displays calibration
curves computed separately for low, medium and high clouds. As can be seen from the plot, the QRNN
predictions are no longer equally well calibrated. Viewed from the Bayesian
perspective, this is not very surprising as CTP values for median clouds have a
significantly different a priori distribution compared to CTP values for all
cloud types, thus giving different a posteriori distributions.</p>
      <p id="d1e4684">For the NN-CTTH algorithm, the results look different. While for low clouds
the calibration deteriorates, the calibration is even slightly improved for
high clouds. This is not surprising as the Gaussian fit may be more
appropriate on different subsets of the test data.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <title>Conclusions</title>
      <?pagebreak page4641?><p id="d1e4695">In this article, quantile regression neural networks have been proposed as a
method to estimate a posteriori distributions of Bayesian remote sensing retrievals.
They have been applied to two retrievals of scalar atmospheric variables. It has
been demonstrated that QRNNs are capable of providing accurate and
well-calibrated probabilistic predictions in agreement with the Bayesian
formulation of the retrieval problem.</p>
      <p id="d1e4698">The synthetic retrieval case presented in Sect. <xref ref-type="sec" rid="Ch1.S3"/> shows that
the conditional distribution learned by the QRNN is the same as the Bayesian a
posteriori distribution obtained from methods that are directly based on the
Bayesian formulation. This in itself seems worthwhile to note, as it reveals the
importance of the training set statistics that implicitly represent the a priori
knowledge. On the synthetic data set, QRNNs compare well to BMCI and even perform
better for small data sets. This indicates that they are able to handle the
“curse of dimensionality” <xref ref-type="bibr" rid="bib1.bibx9" id="paren.39"/> better than BMCI, which would make them more suitable
for the application to retrieval problems with high-dimensional measurement
spaces.</p>
      <p id="d1e4706">While the optimization of computational performance of the BMCI method has
not been investigated in this work, at least compared to a naive
implementation of BMCI, QRNNs allow for retrievals that are at least 1 order of magnitude
faster. QRNN retrievals can be
easily parallelized, and hardware optimized implementations are available for
all modern computing architectures, thus providing very good performance out
of the box.</p>
      <p id="d1e4709">Based on these very promising results, the next step in this line of research
should be to compare QRNNs and BMCI on a real retrieval case to investigate if
the findings from the simulations carry over to the real world. If this is the
case, significant reductions in the computational cost of operational retrievals
and maybe even better retrieval performance could be achieved using QRNNs.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F13"><caption><p id="d1e4715">Calibration of the prediction intervals obtained from NN-CTTH (blue) and
a single QRNN (red) with respect to specific cloud types.</p></caption>
        <?xmltex \igopts{width=227.622047pt}?><graphic xlink:href="https://amt.copernicus.org/articles/11/4627/2018/amt-11-4627-2018-f13.pdf"/>

      </fig>

      <p id="d1e4724">In the second retrieval application presented in this article, QRNNs have been
used to retrieve cloud top pressure from MODIS observations. The results show
that not only are QRNNs able to improve upon state-of-the-art retrieval
accuracy, but they can also learn to predict retrieval uncertainty. The ability of QRNNs
to provide statistically consistent, case-specific uncertainty estimates should
make them a very interesting alternative to non-probabilistic neural network
retrievals. Nonetheless, the sensitivity of the QRNN approach to a priori
assumptions has also been demonstrated. The posterior distribution learned by the
QRNN depends on the validity of the a priori assumptions encoded in the training
data. In particular, accurate uncertainty estimates can only be expected if the
retrieved observations follow the same distribution as the training data. This,
however, is a limitation inherent to all empirical methods.</p>
      <p id="d1e4727">The second application case presented here demonstrated the ability of QRNNs
to represent non-Gaussian retrieval errors. While, as shown in this study,
this is also the case for BMCI (Eq. <xref ref-type="disp-formula" rid="Ch1.E4"/>), it is common in
practice to estimate only mean and standard deviation of the a posteriori
distribution. Furthermore, implementations usually assume Gaussian
measurement errors, which is an unlikely assumption if the observations in
the retrieval database contain modeling errors. By requiring no assumptions
whatsoever on the involved uncertainties, QRNNs may provide a more suitable
way of representing (non-Gaussian) retrieval uncertainties.</p>
      <?pagebreak page4642?><p id="d1e4732">The application of the Bayesian framework to neural network retrievals opens the
door to a number of interesting applications that could be pursued in future
research. It would for example be interesting to investigate if the a priori
information can be separated from the information contained in the retrieved
measurement. This would make it possible to remove the dependency of the
probabilistic predictions on the a priori assumptions, which can currently be
considered a limitation of the approach. Furthermore, estimated a posteriori
distributions obtained from QRNNs could be used to estimate the information
content in a retrieval following the methods outlined by <xref ref-type="bibr" rid="bib1.bibx32" id="text.40"/>.</p>
      <p id="d1e4738">In this study only the retrieval of scalar quantities was considered. Another
aspect of the application of QRNNs to remote sensing retrievals that remains to
be investigated is how they can be used to retrieve vector-valued retrieval
quantities, such as concentration profiles of atmospheric gases or
particles. While the generalization to marginal, multivariate quantiles should
be straightforward, it is unclear whether a better approximation of the
quantile contours of the joint a posteriori distribution can be obtained using
QRNNs.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability">

      <p id="d1e4745">The implementation of the retrieval methods that were used
in this article has been published as parts of the <italic>typhon: tools for atmospheric research</italic>, <ext-link xlink:href="https://doi.org/10.5281/zenodo.1300319" ext-link-type="DOI">10.5281/zenodo.1300319</ext-link> <xref ref-type="bibr" rid="bib1.bibx36" id="paren.41"/> software
package. The source code for the calculations presented in
Sects. <xref ref-type="sec" rid="Ch1.S3"/> (<ext-link xlink:href="https://doi.org/10.5281/zenodo.1207351" ext-link-type="DOI">10.5281/zenodo.1207351</ext-link>) and <xref ref-type="sec" rid="Ch1.S4"/>
(<ext-link xlink:href="https://doi.org/10.5281/zenodo.1207349" ext-link-type="DOI">10.5281/zenodo.1207349</ext-link>) is accessible from public repositories
<xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx29" id="paren.42"/>.</p>
  </notes><notes notes-type="authorcontribution">

      <p id="d1e4774">All authors contributed to the study through discussion and
feedback. PE and BR proposed the application of QRNNs to remote sensing
retrievals. The study was designed and implemented by SP, who also prepared
the manuscript including figures, text and tables. AT and NH provided the
training data for the cloud top pressure retrieval.</p>
  </notes><notes notes-type="competinginterests">

      <p id="d1e4780">The authors declare that they have no conflict of
interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e4786">The scientists at Chalmers University of Technology were funded by the Swedish National Space Board.</p><p id="d1e4788">The authors would like to acknowledge the work of Ronald Scheirer and Sara
Hörnquist, who were involved in the creation of the collocation data set that was
used as training and test data for the cloud top pressure retrieval.</p><p id="d1e4790">Numerous free software packages were used to perform the numerical
experiments presented in this article and visualize their results. The
authors would like to acknowledge the work of all the developers who
contributed to making these tools freely available to the scientific
community, in particular the work by <xref ref-type="bibr" rid="bib1.bibx16" id="text.43"/>, <xref ref-type="bibr" rid="bib1.bibx27" id="text.44"/>,
<xref ref-type="bibr" rid="bib1.bibx37" id="text.45"/> and the Python Software Foundation (<xref ref-type="bibr" rid="bib1.bibx31" id="year.46"/>).<?xmltex \hack{\newline}?><?xmltex \hack{\newline}?>Edited by: Andrew
Sayer <?xmltex \hack{\newline}?> Reviewed by: Christian Kummerow and one anonymous
referee</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Aires et al.(2004)Aires, Prigent, and Rossow</label><mixed-citation>Aires, F., Prigent, C., and Rossow, W. B.: Neural network uncertainty
assessment using Bayesian statistics with application to remote sensing: 2.
Output errors, J. Geophys. Res., 109,  d10304, <ext-link xlink:href="https://doi.org/10.1029/2003JD004174" ext-link-type="DOI">10.1029/2003JD004174</ext-link>,
2004.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Bishop(2006)</label><mixed-citation>
Bishop, C. M.: Pattern Recognition and Machine Learning, Springer-Verlag New
York, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Brath et al.(2018)Brath, Fox, Eriksson, Harlow, Burgdorf, and
Buehler</label><mixed-citation>Brath, M., Fox, S., Eriksson, P., Harlow, R. C., Burgdorf, M., and Buehler,
S. A.: Retrieval of an ice water path over the ocean from ISMAR and MARSS
millimeter and submillimeter brightness temperatures, Atmos. Meas. Tech., 11,
611–632, <ext-link xlink:href="https://doi.org/10.5194/amt-11-611-2018" ext-link-type="DOI">10.5194/amt-11-611-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Buehler et al.(2018)Buehler, Mendrok, Eriksson, Perrin, Larsson, and
Lemke</label><mixed-citation>Buehler, S. A., Mendrok, J., Eriksson, P., Perrin, A., Larsson, R., and
Lemke, O.: ARTS, the Atmospheric Radiative Transfer Simulator – version 2.2,
the planetary toolbox edition, Geosci. Model Dev., 11, 1537–1556,
<ext-link xlink:href="https://doi.org/10.5194/gmd-11-1537-2018" ext-link-type="DOI">10.5194/gmd-11-1537-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Cannon(2011)</label><mixed-citation>Cannon, A. J.: Quantile regression neural networks: Implementation in R and
application to precipitation downscaling, Comput. Geosci., 37, 1277–1284,
<ext-link xlink:href="https://doi.org/10.1016/j.cageo.2010.07.005" ext-link-type="DOI">10.1016/j.cageo.2010.07.005</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Chollet et al.(2015)</label><mixed-citation>Chollet, F. et al.: Keras, available at:
<uri>https://github.com/fchollet/keras</uri> (last access: 30 March 2018), 2015.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Dee et al.(2011)</label><mixed-citation>Dee, D. P., Uppala, S. M., Simmons, A. J., Berrisford, P., Poli, P.,
Kobayashi,
S., Andrae, U., Balmaseda, M. A., Balsamo, G., Bauer, P., Bechtold, P.,
Beljaars, A. C. M., van de Berg, L., Bidlot, J., Bormann, N., Delsol, C.,
Dragani, R., Fuentes, M., Geer, A. J., Haimberger, L., Healy, S. B.,
Hersbach, H., Hólm, E. V., Isaksen, L., Kållberg, P., Köhler, M.,
Matricardi, M., McNally, A. P., Monge-Sanz, B. M., Morcrette, J.-J., Park,
B.-K., Peubey, C., de Rosnay, P., Tavolato, C., Thépaut, J.-N., and Vitart,
F.: The ERA-Interim reanalysis: configuration and performance of the data
assimilation system, Q. J. Roy. Meteor. Soc., 137, 553–597,
<ext-link xlink:href="https://doi.org/10.1002/qj.828" ext-link-type="DOI">10.1002/qj.828</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Evans et al.(2012)Evans, Wang, O'C Starr, Heymsfield, Li, Tian,
Lawson, Heymsfield, and Bansemer</label><mixed-citation>Evans, K. F., Wang, J. R., O'C Starr, D., Heymsfield, G., Li, L., Tian, L.,
Lawson, R. P., Heymsfield, A. J., and Bansemer, A.: Ice hydrometeor profile
retrieval algorithm for high-frequency microwave radiometers: application to
the CoSSIR instrument during TC4, Atmos. Meas. Tech., 5, 2277–2306,
<ext-link xlink:href="https://doi.org/10.5194/amt-5-2277-2012" ext-link-type="DOI">10.5194/amt-5-2277-2012</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Friedman et al.(2001)Friedman, Hastie, and Tibshirani</label><mixed-citation>
Friedman, J., Hastie, T., and Tibshirani, R.: The elements of statistical
learning, vol. 1, Springer series in statistics, New York, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Gelman et al.(2013)Gelman, Carlin, Stern, Dunson, Vehtari, and
Rubin</label><mixed-citation>
Gelman, A., Carlin, J., Stern, H., Dunson, D., Vehtari, A., and Rubin, D.:
Bayesian Data Analysis, 3rd Edn., Chapman &amp; Hall/CRC Texts in
Statistical Science, Taylor &amp; Francis, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Gneiting and Raftery(2007)</label><mixed-citation>Gneiting, T. and Raftery, A. E.: Strictly Proper Scoring Rules, Prediction,
and
Estimation, J. Atmos. Sci., 102, 359–378, <ext-link xlink:href="https://doi.org/10.1198/016214506000001437" ext-link-type="DOI">10.1198/016214506000001437</ext-link>,
2007.</mixed-citation></ref>
      <?pagebreak page4643?><ref id="bib1.bibx12"><label>Gneiting et al.(2005)Gneiting, E., Westveld III, and
Goldman</label><mixed-citation>Gneiting, T., E., R. A., Westveld III, A. H. A. H., and Goldman, T.:
Calibrated
Probabilistic Forecasting Using Ensemble Model Output Statistics and Minimum
CRPS Estimation, Mon. Weather Rev., 133, 1098–1118, <ext-link xlink:href="https://doi.org/10.1175/MWR2904.1" ext-link-type="DOI">10.1175/MWR2904.1</ext-link>,
2005.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Goodfellow et al.(2014)Goodfellow, Shlens, and
Szegedy</label><mixed-citation>
Goodfellow, I. J., Shlens, J., and Szegedy, C.: Explaining and harnessing
adversarial examples, arXiv preprint arXiv:1412.6572, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{H{\aa}kansson et~al.(2018)H{\aa}kansson, Adok, Thoss, Scheirer, and
H\"{o}rnquist}}?><label>Håkansson et al.(2018)Håkansson, Adok, Thoss, Scheirer, and
Hörnquist</label><mixed-citation>Håkansson, N., Adok, C., Thoss, A., Scheirer, R., and Hörnquist, S.:
Neural network cloud top pressure and height for MODIS, Atmos. Meas. Tech.,
11, 3177–3196, <ext-link xlink:href="https://doi.org/10.5194/amt-11-3177-2018" ext-link-type="DOI">10.5194/amt-11-3177-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Holl et al.(2014)Holl, Eliasson, Mendrok, and Buehler</label><mixed-citation>Holl, G., Eliasson, S., Mendrok, J., and Buehler, S. A.: SPARE-ICE:
Synergistic
ice water path from passive operational sensors, J. Geophys.
Res.-Atmos., 119, 1504–1523, <ext-link xlink:href="https://doi.org/10.1002/2013JD020759" ext-link-type="DOI">10.1002/2013JD020759</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Hunter(2007)</label><mixed-citation>Hunter, J. D.: Matplotlib: A 2D graphics environment, Comput. Sci. Eng., 9,
90–95, <ext-link xlink:href="https://doi.org/10.1109/MCSE.2007.55" ext-link-type="DOI">10.1109/MCSE.2007.55</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Jim\'{e}nez et~al.(2003)}}?><label>Jiménez et al.(2003)</label><mixed-citation>Jiménez, C., Eriksson, P., and Murtagh, D.: Inversion of Odin limb sounding
submillimeter observations by a neural network technique, Radio Sci., 38,
8062,  <ext-link xlink:href="https://doi.org/10.1029/2002RS002644" ext-link-type="DOI">10.1029/2002RS002644</ext-link>,  2003.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Kazumori and English(2015)</label><mixed-citation>Kazumori, M. and English, S. J.: Use of the ocean surface wind direction
signal
in microwave radiance assimilation, Q. J. Roy. Meteor. Soc., 141, 1354–1375,
<ext-link xlink:href="https://doi.org/10.1002/qj.2445" ext-link-type="DOI">10.1002/qj.2445</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Koenker(2005)</label><mixed-citation>
Koenker, R.: Quantile Regression, Econometric Society Monographs, Cambridge
University Press, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Koenker and Bassett Jr.(1978)</label><mixed-citation>
Koenker, R. and Bassett Jr., G.: Regression quantiles, Econometrica, 46,
33–50,
1978.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Kummerow et al.(1996)Kummerow, Olson, and Giglio</label><mixed-citation>Kummerow, C., Olson, W. S., and Giglio, L.: A simplified scheme for obtaining
precipitation and vertical hydrometeor profiles from passive microwave
sensors, IEEE Geosci. Remote S., 34, 1213–1232, <ext-link xlink:href="https://doi.org/10.1109/36.536538" ext-link-type="DOI">10.1109/36.536538</ext-link>,
1996.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Kummerow et al.(2015)Kummerow, Randel, Kulie, Wang,
Ferraro, Joseph Munchak, and Petkovic</label><mixed-citation>Kummerow, C. D., Randel, D. L., Kulie, M., Wang, N.-Y., Ferraro,
R.,
Joseph Munchak, S., and Petkovic, V.: The Evolution of the Goddard
Profiling Algorithm to a Fully Parametric Scheme, J. Atmos. Ocean. Tech.,
32, 2265–2280, <ext-link xlink:href="https://doi.org/10.1175/JTECH-D-15-0039.1" ext-link-type="DOI">10.1175/JTECH-D-15-0039.1</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Lakshminarayanan et al.(2016)Lakshminarayanan, Pritzel, and
Blundell</label><mixed-citation>
Lakshminarayanan, B., Pritzel, A., and Blundell, C.: Simple and
Scalable Predictive Uncertainty Estimation using Deep Ensembles, ArXiv
e-prints, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Meinshausen(2006)</label><mixed-citation>
Meinshausen, N.: Quantile Regression Forests, J. Mach. Learn. Res., 7,
983–999,
2006.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{{MODIS Characterization Support Team}(2015{\natexlab{a}})}}?><label>MODIS Characterization Support Team(2015a)</label><mixed-citation>MODIS Characterization Support Team: MODIS/Aqua Calibrated Radiances 5-Min
L1B Swath 1km, <ext-link xlink:href="https://doi.org/10.5067/MODIS/MYD021KM.006" ext-link-type="DOI">10.5067/MODIS/MYD021KM.006</ext-link>, 2015a.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{{MODIS Characterization Support Team}(2015{\natexlab{b}})}}?><label>MODIS Characterization Support Team(2015b)</label><mixed-citation>MODIS Characterization Support Team: MODIS/Aqua Geolocation Fields 5Min L1A
Swath 1km, <ext-link xlink:href="https://doi.org/10.5067/MODIS/MYD03.NRT.006" ext-link-type="DOI">10.5067/MODIS/MYD03.NRT.006</ext-link>, 2015b.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{P{\'{e}}rez and Granger(2007)}}?><label>Pérez and Granger(2007)</label><mixed-citation>Pérez, F. and Granger, B. E.: IPython: a system for interactive
scientific
computing, Comput. Sci. Eng., 9, 21–29, <ext-link xlink:href="https://doi.org/10.1109/MCSE.2007.53" ext-link-type="DOI">10.1109/MCSE.2007.53</ext-link>, 2007.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Pfreundschuh(2018{\natexlab{a}})}}?><label>Pfreundschuh(2018a)</label><mixed-citation>Pfreundschuh, S.: Predicting retrieval uncertainties using neural networks,
available at: <uri>https://github.com/simonpf/predictive_uncertainty</uri> (last
access: 20 March 2018), <ext-link xlink:href="https://doi.org/10.5281/zenodo.1207351" ext-link-type="DOI">10.5281/zenodo.1207351</ext-link>, 2018a.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Pfreundschuh(2018{\natexlab{b}})}}?><label>Pfreundschuh(2018b)</label><mixed-citation>Pfreundschuh, S.: A cloud top pressure retrieval using QRNNs, available at:
<uri>https://github.com/simonpf/ctp_qrnn</uri> (last access: 20 March 2018), <ext-link xlink:href="https://doi.org/10.5281/zenodo.1207349" ext-link-type="DOI">10.5281/zenodo.1207349</ext-link>,
2018b.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Platnick et~al.(2003)Platnick, King, Ackerman, Menzel, Baum,
Ri{\'{e}}di, and Frey}}?><label>Platnick et al.(2003)Platnick, King, Ackerman, Menzel, Baum,
Riédi, and Frey</label><mixed-citation>
Platnick, S., King, M. D., Ackerman, S. A., Menzel, W. P., Baum, B. A.,
Riédi, J. C., and Frey, R. A.: The MODIS cloud products: Algorithms and
examples from Terra, IEEE T. Geosci.  Remote, 41,
459–473, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>The Python Software Foundation(2018)</label><mixed-citation>The Python Software Foundation: The Python Language Reference, available
at: <uri>https://docs.python.org/3/reference/index.html</uri> (last access:
20 March 2018), 2018.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Rodgers(2000)</label><mixed-citation>
Rodgers, C. D.: Inverse Methods For Atmospheric Sounding: Theory And
Practice,
Series On Atmospheric, Oceanic And Planetary Physics, World Scientific
Publishing Company, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Rydberg et al.(2009)Rydberg, Eriksson, Buehler, and
Murtagh</label><mixed-citation>Rydberg, B., Eriksson, P., Buehler, S. A., and Murtagh, D. P.: Non-Gaussian
Bayesian retrieval of tropical upper tropospheric cloud ice and water vapour
from Odin-SMR measurements, Atmos. Meas. Tech., 2, 621–637,
<ext-link xlink:href="https://doi.org/10.5194/amt-2-621-2009" ext-link-type="DOI">10.5194/amt-2-621-2009</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{Strandgren et~al.(2017)Strandgren, Bugliaro, Sehnke, and
Schr\"{o}der}}?><label>Strandgren et al.(2017)Strandgren, Bugliaro, Sehnke, and
Schröder</label><mixed-citation>Strandgren, J., Bugliaro, L., Sehnke, F., and Schröder, L.: Cirrus cloud
retrieval with MSG/SEVIRI using artificial neural networks, Atmos. Meas.
Tech., 10, 3547–3573, <ext-link xlink:href="https://doi.org/10.5194/amt-10-3547-2017" ext-link-type="DOI">10.5194/amt-10-3547-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Tamminen and Kyr{\"{o}}l{\"{a}}(2001)}}?><label>Tamminen and Kyrölä(2001)</label><mixed-citation>
Tamminen, J. and Kyrölä, E.: Bayesian solution for nonlinear and
non-Gaussian inverse problems by Markov chain Monte Carlo method, J. Geophys.
Res., 106, 14377–14390, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>The typhon authors(2018)</label><mixed-citation>The typhon authors: typhon – Tools for atmospheric research, available at:
<uri>https://github.com/atmtools/typhon</uri>, last access: 20 March 2018.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Walt et al.(2011)Walt, Colbert, and Varoquaux</label><mixed-citation>
Walt, S. V. D., Colbert, S. C., and Varoquaux, G.: The NumPy array: a
structure
for efficient numerical computation, Comput. Sci. Eng., 13, 22–30, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Wang et al.(2017)Wang, Prigent, Aires, and Jimenez</label><mixed-citation>
Wang, D., Prigent, C., Aires, F., and Jimenez, C.: A statistical retrieval of
cloud parameters for the millimeter wave Ice Cloud Imager on board MetOp-SG,
IEEE Access, 5, 4057–4076, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Winker et al.(2009)Winker, Vaughan, Omar, Hu, Powell, Liu, Hunt, and
Young</label><mixed-citation>
Winker, D. M., Vaughan, M. A., Omar, A., Hu, Y., Powell, K. A., Liu, Z.,
Hunt,
W. H., and Young, S. A.: Overview of the CALIPSO mission and CALIOP data
processing algorithms, J. Atmos. Ocean. Tech., 26, 2310–2323, 2009.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>A neural network approach to estimating a posteriori distributions  of Bayesian retrieval problems</article-title-html>
<abstract-html><p>A neural-network-based method, quantile regression neural networks (QRNNs), is
proposed as a novel approach to estimating the a posteriori distribution of
Bayesian remote sensing  retrievals. The advantage of QRNNs over conventional
neural network retrievals is that they learn to predict not only a single
retrieval value but also the associated, case-specific uncertainties. In this
study, the retrieval performance of QRNNs is characterized and compared to
that of other state-of-the-art retrieval methods. A synthetic retrieval
scenario is presented and used as a validation case for the application of
QRNNs to Bayesian retrieval problems. The QRNN retrieval performance is
evaluated against Markov chain Monte Carlo simulation and another Bayesian
method based on Monte Carlo integration over a retrieval database. The
scenario is also used to investigate how different hyperparameter
configurations and training set sizes affect the retrieval performance. In the
second part of the study, QRNNs are applied to the retrieval of cloud top
pressure from observations by the Moderate Resolution Imaging
Spectroradiometer (MODIS). It is shown that QRNNs are not only capable of
achieving similar accuracy to standard neural network retrievals but also
provide statistically consistent uncertainty estimates for non-Gaussian
retrieval errors. The results presented in this work show that QRNNs are able
to combine the flexibility and computational efficiency of the machine
learning approach with the theoretically sound handling of uncertainties of
the Bayesian framework. Together with this article, a Python implementation of
QRNNs is released through a public repository to make the method available to
the scientific community.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Aires et al.(2004)Aires, Prigent, and Rossow</label><mixed-citation>
Aires, F., Prigent, C., and Rossow, W. B.: Neural network uncertainty
assessment using Bayesian statistics with application to remote sensing: 2.
Output errors, J. Geophys. Res., 109,  d10304, <a href="https://doi.org/10.1029/2003JD004174" target="_blank">https://doi.org/10.1029/2003JD004174</a>,
2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Bishop(2006)</label><mixed-citation>
Bishop, C. M.: Pattern Recognition and Machine Learning, Springer-Verlag New
York, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Brath et al.(2018)Brath, Fox, Eriksson, Harlow, Burgdorf, and
Buehler</label><mixed-citation>
Brath, M., Fox, S., Eriksson, P., Harlow, R. C., Burgdorf, M., and Buehler,
S. A.: Retrieval of an ice water path over the ocean from ISMAR and MARSS
millimeter and submillimeter brightness temperatures, Atmos. Meas. Tech., 11,
611–632, <a href="https://doi.org/10.5194/amt-11-611-2018" target="_blank">https://doi.org/10.5194/amt-11-611-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Buehler et al.(2018)Buehler, Mendrok, Eriksson, Perrin, Larsson, and
Lemke</label><mixed-citation>
Buehler, S. A., Mendrok, J., Eriksson, P., Perrin, A., Larsson, R., and
Lemke, O.: ARTS, the Atmospheric Radiative Transfer Simulator – version 2.2,
the planetary toolbox edition, Geosci. Model Dev., 11, 1537–1556,
<a href="https://doi.org/10.5194/gmd-11-1537-2018" target="_blank">https://doi.org/10.5194/gmd-11-1537-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Cannon(2011)</label><mixed-citation>
Cannon, A. J.: Quantile regression neural networks: Implementation in R and
application to precipitation downscaling, Comput. Geosci., 37, 1277–1284,
<a href="https://doi.org/10.1016/j.cageo.2010.07.005" target="_blank">https://doi.org/10.1016/j.cageo.2010.07.005</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Chollet et al.(2015)</label><mixed-citation>
Chollet, F. et al.: Keras, available at:
<a href="https://github.com/fchollet/keras" target="_blank">https://github.com/fchollet/keras</a> (last access: 30 March 2018), 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Dee et al.(2011)</label><mixed-citation>
Dee, D. P., Uppala, S. M., Simmons, A. J., Berrisford, P., Poli, P.,
Kobayashi,
S., Andrae, U., Balmaseda, M. A., Balsamo, G., Bauer, P., Bechtold, P.,
Beljaars, A. C. M., van de Berg, L., Bidlot, J., Bormann, N., Delsol, C.,
Dragani, R., Fuentes, M., Geer, A. J., Haimberger, L., Healy, S. B.,
Hersbach, H., Hólm, E. V., Isaksen, L., Kållberg, P., Köhler, M.,
Matricardi, M., McNally, A. P., Monge-Sanz, B. M., Morcrette, J.-J., Park,
B.-K., Peubey, C., de Rosnay, P., Tavolato, C., Thépaut, J.-N., and Vitart,
F.: The ERA-Interim reanalysis: configuration and performance of the data
assimilation system, Q. J. Roy. Meteor. Soc., 137, 553–597,
<a href="https://doi.org/10.1002/qj.828" target="_blank">https://doi.org/10.1002/qj.828</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Evans et al.(2012)Evans, Wang, O'C Starr, Heymsfield, Li, Tian,
Lawson, Heymsfield, and Bansemer</label><mixed-citation>
Evans, K. F., Wang, J. R., O'C Starr, D., Heymsfield, G., Li, L., Tian, L.,
Lawson, R. P., Heymsfield, A. J., and Bansemer, A.: Ice hydrometeor profile
retrieval algorithm for high-frequency microwave radiometers: application to
the CoSSIR instrument during TC4, Atmos. Meas. Tech., 5, 2277–2306,
<a href="https://doi.org/10.5194/amt-5-2277-2012" target="_blank">https://doi.org/10.5194/amt-5-2277-2012</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Friedman et al.(2001)Friedman, Hastie, and Tibshirani</label><mixed-citation>
Friedman, J., Hastie, T., and Tibshirani, R.: The elements of statistical
learning, vol. 1, Springer series in statistics, New York, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Gelman et al.(2013)Gelman, Carlin, Stern, Dunson, Vehtari, and
Rubin</label><mixed-citation>
Gelman, A., Carlin, J., Stern, H., Dunson, D., Vehtari, A., and Rubin, D.:
Bayesian Data Analysis, 3rd Edn., Chapman &amp; Hall/CRC Texts in
Statistical Science, Taylor &amp; Francis, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Gneiting and Raftery(2007)</label><mixed-citation>
Gneiting, T. and Raftery, A. E.: Strictly Proper Scoring Rules, Prediction,
and
Estimation, J. Atmos. Sci., 102, 359–378, <a href="https://doi.org/10.1198/016214506000001437" target="_blank">https://doi.org/10.1198/016214506000001437</a>,
2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Gneiting et al.(2005)Gneiting, E., Westveld III, and
Goldman</label><mixed-citation>
Gneiting, T., E., R. A., Westveld III, A. H. A. H., and Goldman, T.:
Calibrated
Probabilistic Forecasting Using Ensemble Model Output Statistics and Minimum
CRPS Estimation, Mon. Weather Rev., 133, 1098–1118, <a href="https://doi.org/10.1175/MWR2904.1" target="_blank">https://doi.org/10.1175/MWR2904.1</a>,
2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Goodfellow et al.(2014)Goodfellow, Shlens, and
Szegedy</label><mixed-citation>
Goodfellow, I. J., Shlens, J., and Szegedy, C.: Explaining and harnessing
adversarial examples, arXiv preprint arXiv:1412.6572, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Håkansson et al.(2018)Håkansson, Adok, Thoss, Scheirer, and
Hörnquist</label><mixed-citation>
Håkansson, N., Adok, C., Thoss, A., Scheirer, R., and Hörnquist, S.:
Neural network cloud top pressure and height for MODIS, Atmos. Meas. Tech.,
11, 3177–3196, <a href="https://doi.org/10.5194/amt-11-3177-2018" target="_blank">https://doi.org/10.5194/amt-11-3177-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Holl et al.(2014)Holl, Eliasson, Mendrok, and Buehler</label><mixed-citation>
Holl, G., Eliasson, S., Mendrok, J., and Buehler, S. A.: SPARE-ICE:
Synergistic
ice water path from passive operational sensors, J. Geophys.
Res.-Atmos., 119, 1504–1523, <a href="https://doi.org/10.1002/2013JD020759" target="_blank">https://doi.org/10.1002/2013JD020759</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Hunter(2007)</label><mixed-citation>
Hunter, J. D.: Matplotlib: A 2D graphics environment, Comput. Sci. Eng., 9,
90–95, <a href="https://doi.org/10.1109/MCSE.2007.55" target="_blank">https://doi.org/10.1109/MCSE.2007.55</a>, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Jiménez et al.(2003)</label><mixed-citation>
Jiménez, C., Eriksson, P., and Murtagh, D.: Inversion of Odin limb sounding
submillimeter observations by a neural network technique, Radio Sci., 38,
8062,  <a href="https://doi.org/10.1029/2002RS002644" target="_blank">https://doi.org/10.1029/2002RS002644</a>,  2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Kazumori and English(2015)</label><mixed-citation>
Kazumori, M. and English, S. J.: Use of the ocean surface wind direction
signal
in microwave radiance assimilation, Q. J. Roy. Meteor. Soc., 141, 1354–1375,
<a href="https://doi.org/10.1002/qj.2445" target="_blank">https://doi.org/10.1002/qj.2445</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Koenker(2005)</label><mixed-citation>
Koenker, R.: Quantile Regression, Econometric Society Monographs, Cambridge
University Press, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Koenker and Bassett Jr.(1978)</label><mixed-citation>
Koenker, R. and Bassett Jr., G.: Regression quantiles, Econometrica, 46,
33–50,
1978.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Kummerow et al.(1996)Kummerow, Olson, and Giglio</label><mixed-citation>
Kummerow, C., Olson, W. S., and Giglio, L.: A simplified scheme for obtaining
precipitation and vertical hydrometeor profiles from passive microwave
sensors, IEEE Geosci. Remote S., 34, 1213–1232, <a href="https://doi.org/10.1109/36.536538" target="_blank">https://doi.org/10.1109/36.536538</a>,
1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Kummerow et al.(2015)Kummerow, Randel, Kulie, Wang,
Ferraro, Joseph Munchak, and Petkovic</label><mixed-citation>
Kummerow, C. D., Randel, D. L., Kulie, M., Wang, N.-Y., Ferraro,
R.,
Joseph Munchak, S., and Petkovic, V.: The Evolution of the Goddard
Profiling Algorithm to a Fully Parametric Scheme, J. Atmos. Ocean. Tech.,
32, 2265–2280, <a href="https://doi.org/10.1175/JTECH-D-15-0039.1" target="_blank">https://doi.org/10.1175/JTECH-D-15-0039.1</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Lakshminarayanan et al.(2016)Lakshminarayanan, Pritzel, and
Blundell</label><mixed-citation>
Lakshminarayanan, B., Pritzel, A., and Blundell, C.: Simple and
Scalable Predictive Uncertainty Estimation using Deep Ensembles, ArXiv
e-prints, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Meinshausen(2006)</label><mixed-citation>
Meinshausen, N.: Quantile Regression Forests, J. Mach. Learn. Res., 7,
983–999,
2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>MODIS Characterization Support Team(2015a)</label><mixed-citation>
MODIS Characterization Support Team: MODIS/Aqua Calibrated Radiances 5-Min
L1B Swath 1km, <a href="https://doi.org/10.5067/MODIS/MYD021KM.006" target="_blank">https://doi.org/10.5067/MODIS/MYD021KM.006</a>, 2015a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>MODIS Characterization Support Team(2015b)</label><mixed-citation>
MODIS Characterization Support Team: MODIS/Aqua Geolocation Fields 5Min L1A
Swath 1km, <a href="https://doi.org/10.5067/MODIS/MYD03.NRT.006" target="_blank">https://doi.org/10.5067/MODIS/MYD03.NRT.006</a>, 2015b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Pérez and Granger(2007)</label><mixed-citation>
Pérez, F. and Granger, B. E.: IPython: a system for interactive
scientific
computing, Comput. Sci. Eng., 9, 21–29, <a href="https://doi.org/10.1109/MCSE.2007.53" target="_blank">https://doi.org/10.1109/MCSE.2007.53</a>, 2007.

</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Pfreundschuh(2018a)</label><mixed-citation>
Pfreundschuh, S.: Predicting retrieval uncertainties using neural networks,
available at: <a href="https://github.com/simonpf/predictive_uncertainty" target="_blank">https://github.com/simonpf/predictive_uncertainty</a> (last
access: 20 March 2018), <a href="https://doi.org/10.5281/zenodo.1207351" target="_blank">https://doi.org/10.5281/zenodo.1207351</a>, 2018a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Pfreundschuh(2018b)</label><mixed-citation>
Pfreundschuh, S.: A cloud top pressure retrieval using QRNNs, available at:
<a href="https://github.com/simonpf/ctp_qrnn" target="_blank">https://github.com/simonpf/ctp_qrnn</a> (last access: 20 March 2018), <a href="https://doi.org/10.5281/zenodo.1207349" target="_blank">https://doi.org/10.5281/zenodo.1207349</a>,
2018b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Platnick et al.(2003)Platnick, King, Ackerman, Menzel, Baum,
Riédi, and Frey</label><mixed-citation>
Platnick, S., King, M. D., Ackerman, S. A., Menzel, W. P., Baum, B. A.,
Riédi, J. C., and Frey, R. A.: The MODIS cloud products: Algorithms and
examples from Terra, IEEE T. Geosci.  Remote, 41,
459–473, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>The Python Software Foundation(2018)</label><mixed-citation>
The Python Software Foundation: The Python Language Reference, available
at: <a href="https://docs.python.org/3/reference/index.html" target="_blank">https://docs.python.org/3/reference/index.html</a> (last access:
20 March 2018), 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Rodgers(2000)</label><mixed-citation>
Rodgers, C. D.: Inverse Methods For Atmospheric Sounding: Theory And
Practice,
Series On Atmospheric, Oceanic And Planetary Physics, World Scientific
Publishing Company, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Rydberg et al.(2009)Rydberg, Eriksson, Buehler, and
Murtagh</label><mixed-citation>
Rydberg, B., Eriksson, P., Buehler, S. A., and Murtagh, D. P.: Non-Gaussian
Bayesian retrieval of tropical upper tropospheric cloud ice and water vapour
from Odin-SMR measurements, Atmos. Meas. Tech., 2, 621–637,
<a href="https://doi.org/10.5194/amt-2-621-2009" target="_blank">https://doi.org/10.5194/amt-2-621-2009</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Strandgren et al.(2017)Strandgren, Bugliaro, Sehnke, and
Schröder</label><mixed-citation>
Strandgren, J., Bugliaro, L., Sehnke, F., and Schröder, L.: Cirrus cloud
retrieval with MSG/SEVIRI using artificial neural networks, Atmos. Meas.
Tech., 10, 3547–3573, <a href="https://doi.org/10.5194/amt-10-3547-2017" target="_blank">https://doi.org/10.5194/amt-10-3547-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Tamminen and Kyrölä(2001)</label><mixed-citation>
Tamminen, J. and Kyrölä, E.: Bayesian solution for nonlinear and
non-Gaussian inverse problems by Markov chain Monte Carlo method, J. Geophys.
Res., 106, 14377–14390, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>The typhon authors(2018)</label><mixed-citation>
The typhon authors: typhon – Tools for atmospheric research, available at:
<a href="https://github.com/atmtools/typhon" target="_blank">https://github.com/atmtools/typhon</a>, last access: 20 March 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Walt et al.(2011)Walt, Colbert, and Varoquaux</label><mixed-citation>
Walt, S. V. D., Colbert, S. C., and Varoquaux, G.: The NumPy array: a
structure
for efficient numerical computation, Comput. Sci. Eng., 13, 22–30, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Wang et al.(2017)Wang, Prigent, Aires, and Jimenez</label><mixed-citation>
Wang, D., Prigent, C., Aires, F., and Jimenez, C.: A statistical retrieval of
cloud parameters for the millimeter wave Ice Cloud Imager on board MetOp-SG,
IEEE Access, 5, 4057–4076, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Winker et al.(2009)Winker, Vaughan, Omar, Hu, Powell, Liu, Hunt, and
Young</label><mixed-citation>
Winker, D. M., Vaughan, M. A., Omar, A., Hu, Y., Powell, K. A., Liu, Z.,
Hunt,
W. H., and Young, S. A.: Overview of the CALIPSO mission and CALIOP data
processing algorithms, J. Atmos. Ocean. Tech., 26, 2310–2323, 2009.
</mixed-citation></ref-html>--></article>
