<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">AMT</journal-id><journal-title-group>
    <journal-title>Atmospheric Measurement Techniques</journal-title>
    <abbrev-journal-title abbrev-type="publisher">AMT</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Atmos. Meas. Tech.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1867-8548</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/amt-14-391-2021</article-id><title-group><article-title>Classification of lidar measurements using supervised and unsupervised machine learning methods</article-title><alt-title>Machine learning for lidar data classification</alt-title>
      </title-group><?xmltex \runningtitle{Machine learning for lidar data classification}?><?xmltex \runningauthor{G. Farhani et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Farhani</surname><given-names>Ghazal</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Sica</surname><given-names>Robert J.</given-names></name>
          <email>sica@uwo.ca</email>
        <ext-link>https://orcid.org/0000-0003-2964-1664</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Daley</surname><given-names>Mark Joseph</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Department of Physics and Astronomy, The University of Western Ontario, 1151 Richmond St., <?xmltex \hack{\break}?>London, ON, N6A 3K7, Canada</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Computer Science, The Vector Institute for Artificial Intelligence, The University of Western Ontario, <?xmltex \hack{\break}?> 1151 Richmond St., London, ON, N6A 3K7, Canada</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Robert J. Sica (sica@uwo.ca)</corresp></author-notes><pub-date><day>18</day><month>January</month><year>2021</year></pub-date>
      
      <volume>14</volume>
      <issue>1</issue>
      <fpage>391</fpage><lpage>402</lpage>
      <history>
        <date date-type="received"><day>20</day><month>December</month><year>2019</year></date>
           <date date-type="accepted"><day>13</day><month>November</month><year>2020</year></date>
           <date date-type="rev-recd"><day>13</day><month>October</month><year>2020</year></date>
           <date date-type="rev-request"><day>17</day><month>February</month><year>2020</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2021 Ghazal Farhani et al.</copyright-statement>
        <copyright-year>2021</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021.html">This article is available from https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021.html</self-uri><self-uri xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021.pdf">The full text article is available as a PDF file from https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e108">While it is relatively straightforward to automate the processing of lidar signals, it is more
difficult to choose periods of “good” measurements to process. Groups use various ad hoc
procedures involving either very simple (e.g. signal-to-noise ratio) or more complex procedures
<xref ref-type="bibr" rid="bib1.bibx24" id="paren.1"><named-content content-type="pre">e.g.</named-content></xref> to perform a task that is easy to train humans to perform but is time-consuming. Here, we use machine learning techniques to train the machine to sort the measurements
before processing. The presented method is generic and can be applied to most lidars. We test the
techniques using measurements from the Purple Crow Lidar (PCL) system located in London,
Canada. The PCL has over 200 000 raw profiles in Rayleigh and Raman channels available for
classification. We classify raw (level-0) lidar measurements as “clear” sky profiles with strong
lidar returns, “bad” profiles, and profiles which are significantly influenced by clouds or
aerosol loads.  We examined different supervised machine learning algorithms including the random
forest, the support vector machine, and the gradient boosting trees, all of which can successfully
classify profiles. The algorithms were trained using about 1500 profiles for each PCL channel,
selected randomly from different nights of measurements in different years. The success rate of identification for all the channels is above 95 %.  We also used the <inline-formula><mml:math id="M1" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-distributed stochastic embedding (<inline-formula><mml:math id="M2" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE) method, which is an unsupervised algorithm, to cluster our lidar profiles. Because the <inline-formula><mml:math id="M3" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE is a data-driven method in which no labelling of the training set is needed, it is an attractive algorithm to find anomalies in lidar profiles. The method has been tested on several nights of measurements from the PCL measurements. The <inline-formula><mml:math id="M4" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE can successfully
cluster the PCL data profiles into meaningful categories. To demonstrate the use of the technique,
we have used the algorithm to identify stratospheric aerosol layers due to wildfires.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e153">Lidar (light detection and ranging) is an active remote sensing method which uses a laser to
generate photons that are transmitted to the atmosphere and are scattered back by atmospheric
constituents. The back-scattered photons are collected using a telescope. Lidars provide both high
temporal and spatial resolution profiling and are widely used in atmospheric research. The recorded
back-scattered measurements (also known as level-0 profiles) are often co-added in time and/or in
height. Before co-adding, profiles should be checked for quality purposes to remove “bad
profiles”. Bad profiles include measurements with low-power laser, high background counts,
outliers, and profiles with distorted or unusual shapes for a wide variety of instrumental or
atmospheric reasons. Moreover, depending on the lidar system and the purpose of the measurements,
profiles with traces of clouds or aerosol might be classified separately. During a measurement,
signal quality can change for different reasons including changes in sky background, the appearance
of clouds, and laser power fluctuation. Hence, it is difficult to use traditional programming
techniques to make a robust model that works under the wide range of real cases (even with multiple
layers of exception handling).</p>
      <?pagebreak page392?><p id="d1e156">In this article we propose both supervised and unsupervised machine learning approaches for level-0
lidar data classification and clustering. ML techniques hold great promise for application to the
large data sets obtained by the current and future generation of high temporal–spatial resolution
lidars. ML has been recently used to distinguish between aerosols and clouds for the Cloud-Aerosol
Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) level-2 measurements
<xref ref-type="bibr" rid="bib1.bibx25" id="paren.2"/>. Furthermore, <xref ref-type="bibr" rid="bib1.bibx17" id="text.3"/> used a neural network algorithm to estimate
the most probable aerosol types in a set of data obtained from the European Aerosol Research Lidar
Network (EARLINET). Both <xref ref-type="bibr" rid="bib1.bibx25" id="text.4"/> and <xref ref-type="bibr" rid="bib1.bibx17" id="text.5"/> concluded that their
proposed ML algorithms can classify large sets of data and can successfully distinguish between
different types of aerosols.</p>
      <p id="d1e171">A common way of classifying profiles is to define a threshold for the signal-to-noise ratio at some
altitude: any scan that does not meet the pre-defined threshold value is flagged as bad. In this
method, bad profiles may be incorrectly flagged as good, as they might pass the threshold criteria but have the wrong shape at other altitudes. Recently, <xref ref-type="bibr" rid="bib1.bibx24" id="text.6"/> suggested that a
Mann–Whitney–Wilcoxon rank-sum metric could be used to identify bad profiles. In the
Mann–Whitney–Wilcoxon test, the null hypothesis that the two populations are the same is tested
against the alternate hypothesis that there is a significant difference between the two
populations. The main advantage of this method is that it can be used when the data distribution
does not follow a Gaussian distribution. However, Monte Carlo simulations have shown that when the
two populations have similar medians, but different variances, the Mann–Whitney–Wilcoxon can
wrongfully accept the alternative hypothesis <xref ref-type="bibr" rid="bib1.bibx19" id="paren.7"/>. Here, we propose a machine learning
(ML) approach for level-0 data classification. The classification of lidar profiles is based on
supervised ML techniques which will be discussed in detail in Sect. <xref ref-type="sec" rid="Ch1.S2"/>.</p>
      <p id="d1e182">Using an unsupervised ML approach, we also examined the capability of ML to detect anomalies
(traces of wildfire smoke in lower stratosphere). The PCL is capable of detecting the smoke injected into the lower stratosphere from wildfires <xref ref-type="bibr" rid="bib1.bibx5 bib1.bibx8" id="paren.8"/>. We are interested in whether the PCL can automatically (by using ML methods) detect aerosol loads in the
upper troposphere and lower stratosphere (UTLS) after major wildfires. Aerosols in the UTLS and stratosphere have
important impacts on the radiative budget of the atmosphere. Recently, <xref ref-type="bibr" rid="bib1.bibx4" id="text.9"/>
proposed that smoke aerosols from the forest fires, unlike the aerosols from the volcanic eruptions,
can have a net positive radiative forcing. Considering that the number of forest fires
have increased, detecting the aerosol loads from fires in the UTLS and accounting for them in
atmospheric and climate models is important.</p>
      <p id="d1e192">Section <xref ref-type="sec" rid="Ch1.S2"/> is a brief description of the characteristics of the lidars we used
and an explanation of how ML can be useful for the lidar data classification. Furthermore, the
algorithms which are used in the paper are explained in detail. In Sect. <xref ref-type="sec" rid="Ch1.S3"/> we show
classification and clustering results for the PCL system. In Sect. <xref ref-type="sec" rid="Ch1.S4"/>, a summary of
the ML approach is provided, and the future directions are discussed.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Machine learning algorithms</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Instrument description and machine learning classification of data</title>
      <p id="d1e216">The PCL is a Rayleigh–Raman lidar which has been operational since 1992. Details about PCL
instrumentation can be found in <xref ref-type="bibr" rid="bib1.bibx22" id="text.10"/>. From 1992 to 2010, the lidar was located at
the Delaware Observatory (<inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:mn mathvariant="normal">42.5</mml:mn><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> N, <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mn mathvariant="normal">81.2</mml:mn><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> W) near London, Ontario, Canada. In
2012, the lidar was moved to the Environmental Science Western Field Station (<inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mn mathvariant="normal">43.1</mml:mn><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> N,
<inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:mn mathvariant="normal">81.3</mml:mn><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> W). The PCL uses a second harmonic of an Nd:YAG solid state laser. The laser
operates at 532 <inline-formula><mml:math id="M9" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">nm</mml:mi></mml:mrow></mml:math></inline-formula> and has a repetition rate of 30 <inline-formula><mml:math id="M10" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">Hz</mml:mi></mml:mrow></mml:math></inline-formula> at 1000 <inline-formula><mml:math id="M11" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mJ</mml:mi></mml:mrow></mml:math></inline-formula>. The
receiver is a liquid mercury mirror with the diameter of 2.6 <inline-formula><mml:math id="M12" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>. The PCL currently has four
detection channels:
<list list-type="order"><list-item>
      <p id="d1e305">A high-gain Rayleigh (HR) channel that detects the back-scattered counts from 25 to
110 <inline-formula><mml:math id="M13" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> altitude (vertical resolution: 7 <inline-formula><mml:math id="M14" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>).</p></list-item><list-item>
      <p id="d1e325">A low-gain Rayleigh (LR) channel that detects the back-scattered counts from 25 to
110 <inline-formula><mml:math id="M15" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> altitude (this channel is optimized to detect counts at lower altitudes where the
high-intensity back-scattered counts can saturate the detector and cause non-linearity in the
observed signal; thus, using the low-gain channel, at lower altitudes, the signal remains linear) (vertical resolution: 7 <inline-formula><mml:math id="M16" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>).</p></list-item><list-item>
      <p id="d1e345">A nitrogen Raman channel that detects the vibrational Raman-shifted back-scattered counts
above 0.5 <inline-formula><mml:math id="M17" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> in altitude (vertical resolution: 7 <inline-formula><mml:math id="M18" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>).</p></list-item><list-item>
      <p id="d1e365">A water vapour Raman channel that detects the vibrational Raman-shifted back-scattered counts
above 0.5 <inline-formula><mml:math id="M19" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> in altitude (vertical resolution: 24 <inline-formula><mml:math id="M20" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>).</p></list-item></list>
The Rayleigh channels are used for atmospheric temperature retrievals, and the water vapour and
nitrogen channels are used to retrieve water vapour mixing ratio.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e387">Example of measurements taken by PCL Rayleigh and Raman channels. <bold>(a)</bold> Examples of
bad profiles for both Rayleigh and Raman channels. In this plot, the signals in cyan and dark red
have extremely low laser power, the purple signal has extremely high background counts, and the
signal in orange has a distorted shape and high background counts. <bold>(b)</bold> Example of a good
scan in the Rayleigh channel. <bold>(c)</bold> Example of cloudy sky in the nitrogen (Raman)
channel. At about 8 <inline-formula><mml:math id="M21" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> a layer of either cloud or aerosol occurs.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f01.png"/>

        </fig>

      <?pagebreak page393?><p id="d1e413">In our lidar scan classification using supervised learning, we have a training set in which, for
each scan, counts at each altitude are considered as an attribute, and the classification of the
scan is the output value. Formally, we are trying to learn a prediction function <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>:
<inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>→</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:math></inline-formula>, which minimizes the expectation of some loss function
<inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="normal">Σ</mml:mi><mml:mi>i</mml:mi><mml:mi>N</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mtext>true</mml:mtext></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mtext>predicted</mml:mtext></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mtext>true</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula>
is the actual value (label) of the classification for each data point;
<inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mtext>predicted</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula> is the prediction generated from the prediction function, and <inline-formula><mml:math id="M27" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the
length of the data set <xref ref-type="bibr" rid="bib1.bibx1" id="paren.11"/>. Thus, the training set is a matrix with size <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>,
and each row of the matrix presents a lidar scan in the training set. The columns of this matrix
(except the last column) are photon counts at each altitude. The last column of the matrix shows
the classifications of each scan. Examples of each scan's class for PCL measurements using Rayleigh
and nitrogen Raman digital channels are shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>.</p>
      <p id="d1e546">We also examined unsupervised learning to generate meaningful clusters. We are interested in
determining whether the lidar profiles, based on their similarities (similar features), will be clustered
together. For our clustering task, a good ML method will distinguish between high background counts,
low-laser power profiles, clouds, and high-laser profiles, and put each of these in a different
cluster. Moreover, using unsupervised learning, anomalies in profiles (a.k.a. traces of smoke in
higher altitudes) should be apparent.</p>
      <p id="d1e549">Many algorithms have been developed for both supervised and unsupervised learning. In the following
section, we introduce support vector machine (SVM), decision tree, random forest, and gradient
boosting tree methods as part of ML algorithms that we have tested for sorting lidar profiles. We
also describe the <inline-formula><mml:math id="M29" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-distributed stochastic neighbour embedding method and density-based spatial
clustering as unsupervised algorithms which were used in this study.</p>
      <p id="d1e559">Recently, deep neural networks (DNNs) have received attention in the scientific community. In the
neural network approach the loss function computes the error between the output scores and target
values. The internal parameters (weights) in the algorithm are modified such that the error becomes
smaller. The process of tuning the weights continues until the error is not decreasing anymore. A
typical deep learning algorithm can have hundreds of millions of weights, inputs and target
values. Thus, the algorithm is useful when dealing with large sets of images and text
data. Although DNNs are power full tools, they are acting as black boxes and important questions
such as what features in the input data are more important remain unknown. For this study we decided
to use the classical machine learning algorithms as they can provide a better explanation of feature
selection.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Support vector machine algorithms</title>
      <p id="d1e570">SVM algorithms are popular in the remote sensing community because they can be trained with
relatively small data sets, while producing highly accurate predictions <xref ref-type="bibr" rid="bib1.bibx15 bib1.bibx7" id="paren.12"/>. Moreover,
unlike some statistical methods such as the maximum likelihood estimation that assume the data are
normally distributed, SVM algorithms do not require this assumption. This property makes them
suitable for data sets with unknown distributions. Here, we briefly describe how SVM works. More
details on the topic can be found in <xref ref-type="bibr" rid="bib1.bibx3" id="text.13"/> and <xref ref-type="bibr" rid="bib1.bibx23" id="text.14"/>.</p>
      <?pagebreak page394?><p id="d1e582">The SVM algorithm finds an optimal hyperplane that separates the data set into a distinct predefined
number of classes <xref ref-type="bibr" rid="bib1.bibx1" id="paren.15"/>. For binary classification in a linearly separable
data set, a target class <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> is considered with a set of input data vectors
<inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The optimal solution is obtained by maximizing the margin (<inline-formula><mml:math id="M32" display="inline"><mml:mi mathvariant="bold-italic">w</mml:mi></mml:math></inline-formula>) between the
separating hyperplane and the data. It can be shown that the optimal hyperplane is the solution of
the constrained quadratic equation:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M33" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E1"><mml:mtd><mml:mtext>1</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>minimize: </mml:mtext><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>‖</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:msup><mml:mo>‖</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>constraint: </mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mi mathvariant="italic">⊺</mml:mi></mml:msup><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo><mml:mi mathvariant="italic">⩾</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            In the above equation the constraint is a linear model where <inline-formula><mml:math id="M34" display="inline"><mml:mi mathvariant="bold-italic">w</mml:mi></mml:math></inline-formula> and the intercept (<inline-formula><mml:math id="M35" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula>) are
unknowns (need to be optimized). To solve this constrained optimization problem, the Lagrange
function can be built:
            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M36" display="block"><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced open="∥" close="∥"><mml:mrow><mml:msup><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced><mml:mo>-</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:msup><mml:mi/><mml:mi mathvariant="italic">⊺</mml:mi></mml:msup><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are Lagrangian multipliers. Setting the derivatives of <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> with
respect to <inline-formula><mml:math id="M39" display="inline"><mml:mi mathvariant="bold-italic">w</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M40" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> to zero:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M41" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E4"><mml:mtd><mml:mtext>4</mml:mtext></mml:mtd><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E5"><mml:mtd><mml:mtext>5</mml:mtext></mml:mtd><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            Thus we can rewrite the Lagrangian as follows:
            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M42" display="block"><mml:mrow><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msubsup><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi><mml:mi mathvariant="italic">⊺</mml:mi></mml:msubsup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          It is clear that the optimization process only depends on the dot product of the samples.</p>
      <p id="d1e1012">Many real-world problems involve non-linear data sets in which the above methodology will fail. To
tackle the non-linearity, using a non-linear function <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mi mathvariant="normal">Φ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the feature space is mapped
into higher-dimensional feature space. The Lagrangian function can be re-written as follows:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M44" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E7"><mml:mtd><mml:mtext>7</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:munder><mml:mo movablelimits="false">∑</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="italic">α</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E8"><mml:mtd><mml:mtext>8</mml:mtext></mml:mtd><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="normal">Φ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mi mathvariant="italic">⊺</mml:mi></mml:msup><mml:mi mathvariant="normal">Φ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is known as the kernel function. Kernel functions let the feature
space be mapped into higher-dimensional space without the need to calculate the transformation
function (only the kernel is needed).  More details on SVM and kernel functions can be found in
<xref ref-type="bibr" rid="bib1.bibx1" id="text.16"/>.</p>
      <p id="d1e1216">To use SVM as a multi-class classifier, some adjustments need to be made to the simple SVM binary
model. Methods like a directed acyclic graph, one-against-all, and one-against-others are among the
most successful techniques for multi-class classification. Details about these methods can be found
in <xref ref-type="bibr" rid="bib1.bibx11" id="text.17"/>.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Decision trees algorithms</title>
      <p id="d1e1230">Decision trees are nonparametric algorithms that allow complex relations between inputs and outputs
to be modelled. Moreover, they are the foundation of both random forest and boosting methods. A
comprehensive introduction to the topic can be found in <xref ref-type="bibr" rid="bib1.bibx18" id="text.18"/>. Here, we
briefly describe how a decision tree is built.</p>
      <p id="d1e1236">A decision tree is a set of (binary) decisions represented by an acyclic graph directed outward from
a root node to each leaf. Each node has one parent (except the root) and can have two children. A
node with no children is called a leaf. Decision trees can be complex depending on the data set. A
tree can be simplified by pruning, which means leaves from the upper parts of the trees will be
cut. To grow a decision tree, the following steps are taken.
<list list-type="bullet"><list-item>
      <p id="d1e1241"><italic>Defining a set of candidate splits.</italic> We should answer a question about the value of a selected
input feature to split the data set into two groups.</p></list-item><list-item>
      <p id="d1e1247"><italic>Evaluating the splits.</italic> Using a score measure, at each node, we can decide what the best
question is to be asked and what the best feature is to be used. As the goal of splitting is to
find the purest learning subset that is in each leaf, we want the output labels to be the same;
called purifying. Shannon entropy (see below) is used to evaluate the purity of each
subgroup. Thus, a split that reduces the entropy from one node to its descendent is favourable.</p></list-item><list-item>
      <p id="d1e1253"><italic>Deciding to stop splitting.</italic> We set rules to define when the splitting should be stopped, and a
node becomes a leaf. This decision can be data-driven. For example, we can stop splitting when all
objects in a node have the same label (pure node). The decision can be defined by a user as
well. For example, we can limit the maximum depth of the tree (length of the path between root and
a leaf).</p></list-item></list></p>
      <p id="d1e1258">In a decision tree, by performing a full scan of attribute space the optimal split (at each local
node) is selected, and irrelevant attributes are discarded. This method allows us to identify the
attributes that are most important in our decision-making process.</p>
      <p id="d1e1261">The metric used to judge the quality of the tree splitting is Shannon entropy
<xref ref-type="bibr" rid="bib1.bibx21" id="paren.19"/>. Shannon entropy describes the amount of information gained with
each event and is calculated as follows:
            <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M46" display="block"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Σ</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>log⁡</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents a set of probabilities that adds up to 1. <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> means that no new
information was gained in the process of splitting, and <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>H</mml:mi><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> means that the maximum amount of
information was achieved. Ideally, the produced leaves will be pure and have low entropy (meaning
all of the objects in the leaf are the same).</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Random forests</title>
      <?pagebreak page395?><p id="d1e1356">The random forest (RF) method is based on “growing” an ensemble of decision trees that vote for
the most popular class. Typically the bagging (bootstrap aggregating) method is used to generate the
ensemble of trees <xref ref-type="bibr" rid="bib1.bibx2" id="paren.20"/>. In bagging, to grow the <inline-formula><mml:math id="M50" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>th tree, a random vector
<inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> from the training set is selected. The <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> vector is
independent of the past vectors (<inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) but has the
same distribution. Then, by selecting random features, the <inline-formula><mml:math id="M54" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>th tree is generated. Each tree is
a classifier (<inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi>h</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>) that casts a vote. During the construction of
decision trees, in each interior node, the Gini index is used to evaluate the subset of selected
features. The Gini index is the measure of impurity of data <xref ref-type="bibr" rid="bib1.bibx12 bib1.bibx13" id="paren.21"/>. Thus, it is desirable to select a feature that results in a greater
decrease in the Gini index (partitioning the data into distinct classes). For a split at node <inline-formula><mml:math id="M56" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> the
index can be calculated as <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msubsup><mml:mi mathvariant="normal">Σ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the frequency of class <inline-formula><mml:math id="M59" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>
in the node <inline-formula><mml:math id="M60" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>. Finally, the class label is determined via majority voting among all the trees
<xref ref-type="bibr" rid="bib1.bibx13" id="paren.22"/>.</p>
      <p id="d1e1515">One major problem in ML is that when the algorithm becomes too complicated and perfectly fits the
training data points, it loses its generality and performs poorly on the testing set. This problem
is known as overfitting. For RF, increasing the number of trees can help with the overfitting
problem. Other parameters that can significantly influence RFs are the tree depth and the number of trees. As a tree gets deeper it has more splits while growing more trees in a forest yields a smaller prediction error. Finding the optimal depth of each tree is a critical
parameter. While leaves in a short tree may contain heterogeneous data (the leaves are not pure),
tall trees can suffer from poor generalization (overfitting problem). Thus, the optimal depth
provides a tree with pure leaves and great generalization. Detailed discussion on the RFs can be
found in <xref ref-type="bibr" rid="bib1.bibx13" id="text.23"/>.</p>
</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Gradient boosting tree methods</title>
      <p id="d1e1530">Boosting methods are based on the idea that combining many “weak” approximation models (a learning
algorithm that is slightly more accurate than 50 %) will eventually boost the predictive
performance <xref ref-type="bibr" rid="bib1.bibx11 bib1.bibx20" id="paren.24"/>. Thus, many “local rules” are combined to
produce highly accurate models.</p>
      <p id="d1e1536">In the gradient boosting method, simple parametrized models (base models) are sequentially fitted to
current residuals (known as pseudo-residuals) at each iteration. The residuals are the gradients of
the loss function (they show the difference between the predicted value and the true value) that we
are trying to minimize. The gradient boosting tree (GBT) algorithm is a sequence of simple trees
generated such that each successive tree is grown based on the prediction residual of the preceding
tree with the goal of reducing the new residual. This “additive weighted expansion” of trees will
eventually become a strong classifier <xref ref-type="bibr" rid="bib1.bibx11" id="paren.25"/>. This method can be successfully used
even when the relation between the instances and output values are complex. Compared to the RF
model, which is based on building many independent models and combining them (using some averaging
techniques), the gradient boosting method is based on building sequential models.</p>
      <p id="d1e1542">Although the GBTs show overall high performance, they require large set of training data and the
method is quite susceptible to noise. Thus, for smaller training data set the algorithm suffers from
overfitting. As the size of our data set is large, GBT could potentially be a reliable algorithm for the
classification of lidar profiles in this case.</p><?xmltex \hack{\newpage}?>
</sec>
<sec id="Ch1.S2.SS6">
  <label>2.6</label><?xmltex \opttitle{The $t$-distributed stochastic neighbour embedding method}?><title>The <inline-formula><mml:math id="M61" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-distributed stochastic neighbour embedding method</title>
      <p id="d1e1562">A detailed description of unsupervised learning can be found in <xref ref-type="bibr" rid="bib1.bibx9" id="text.26"/>. Here, we
briefly introduce two of the unsupervised algorithms that are used in this paper. The <inline-formula><mml:math id="M62" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-distributed stochastic embedding (<inline-formula><mml:math id="M63" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE) method
is an unsupervised ML algorithm that is based on stochastic neighbour embedding (SNE). In the SNE,
the data points are placed into a low-dimensional space such that the neighbourhood identity of each
data point is preserved <xref ref-type="bibr" rid="bib1.bibx10" id="paren.27"/>. The SNE is based on finding the probability that data point
<inline-formula><mml:math id="M64" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> has data point <inline-formula><mml:math id="M65" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> as its neighbour, which can formally be written as follows:
            <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M66" display="block"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>≠</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the probability of <inline-formula><mml:math id="M68" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> selecting <inline-formula><mml:math id="M69" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> as its neighbour and <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula> is the
squared Euclidean distance between two points in the high dimensional space. This can be written as follows:
            <disp-formula id="Ch1.E11" content-type="numbered"><label>11</label><mml:math id="M71" display="block"><mml:mrow><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>‖</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msup><mml:mo>‖</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>i</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is defined so that the entropy of the distribution becomes <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mi mathvariant="italic">κ</mml:mi></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M74" display="inline"><mml:mi mathvariant="italic">κ</mml:mi></mml:math></inline-formula> is the “perplexity”, which is set by the user and determines how many neighbours will be
around a selected point.</p>
      <p id="d1e1811">The SNE tries to model each data point, <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, at the higher dimension, by a point <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> at a
lower dimension such that the similarities in <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are conserved. In this low-dimensional map,
we assume that the points follow a Gaussian distribution. Thus, the SNE tries to make the best match
between the original distribution (<inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) and the induced probability distribution
(<inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>). This match is determined by minimizing the error between the two distributions, and the
best match is developed. The induced probability is defined as follows:
            <disp-formula id="Ch1.E12" content-type="numbered"><label>12</label><mml:math id="M80" display="block"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mo>‖</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msup><mml:mo>‖</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>≠</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>exp⁡</mml:mi><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mo>‖</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msup><mml:mo>‖</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e1981">The SNE algorithm aims to find a low-dimensional data representation such that the mismatch between <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> become minimized; thus in the SNE the Kullback–Leibler divergences is defined as the cost function. Using the gradient descent method the cost function is minimized. The
cost function is written as follows:
            <disp-formula id="Ch1.E13" content-type="numbered"><label>13</label><mml:math id="M83" display="block"><mml:mrow><mml:mtext>cost</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">Σ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>‖</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">Σ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="normal">Σ</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>∣</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log⁡</mml:mi><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>∣</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>∣</mml:mo><mml:mo>|</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the conditional probability distribution of all data points given data points
<inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the conditional probability for all the data points given data points
<inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e2153">Red curve: the Gaussian distribution for data points, extending from <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula> to
<inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula>. The mean of the distribution is at 0. Blue curve: the Student's <inline-formula><mml:math id="M90" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution over the
same range. The distribution is heavy-tailed, compared to the Gaussian distribution.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f02.png"/>

        </fig>

      <p id="d1e2191">The <inline-formula><mml:math id="M91" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE uses a similar approach but assumes a lower-dimensional space, which instead of being a
Gaussian distribution follows Student's <inline-formula><mml:math id="M92" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution with a single degree<?pagebreak page396?> of freedom. Thus, since
a heavy-tailed distribution is used to measure similarities between the points in the lower
dimension, the data points that are less similar will be located further from each other. To
demonstrate the difference between the Student's <inline-formula><mml:math id="M93" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution and the Gaussian distribution, we
plot the two distributions in Fig. <xref ref-type="fig" rid="Ch1.F2"/>. Here, the <inline-formula><mml:math id="M94" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> values are within
<inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula>. The Gaussian distribution with the mean at 0 and the Student's
<inline-formula><mml:math id="M97" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution with the degree of freedom of 1 are generated. As is shown in the figure, the
<inline-formula><mml:math id="M98" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> distribution peaks at a lower value and has a more pronounced tail. The above approach gives <inline-formula><mml:math id="M99" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE
an excellent capability for visualizing data, and thus, we use this method to allow scan
classification via unsupervised learning. More details on SNE and <inline-formula><mml:math id="M100" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE can be found in
<xref ref-type="bibr" rid="bib1.bibx10" id="text.28"/> and <xref ref-type="bibr" rid="bib1.bibx14" id="text.29"/>.</p>
</sec>
<sec id="Ch1.S2.SS7">
  <label>2.7</label><title>Density-based spatial clustering of applications with noise (DBSCAN)</title>
      <p id="d1e2290">DBSCAN is an unsupervised learning method that relies on density-based clustering and is capable of
discovering any arbitrary shape from a collection of points. There are two input parameters to be set
by the user: minPts, which indicates the minimum number of points needed to make a cluster, and
<inline-formula><mml:math id="M101" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> such that the <inline-formula><mml:math id="M102" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> neighbourhood of point <inline-formula><mml:math id="M103" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> denoting as <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is
defined as follows:
            <disp-formula id="Ch1.E14" content-type="numbered"><label>14</label><mml:math id="M105" display="block"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi mathvariant="italic">ϵ</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mi>q</mml:mi><mml:mo>∈</mml:mo><mml:mi>D</mml:mi><mml:mo>∣</mml:mo><mml:mtext>dis</mml:mtext><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo><mml:mo>≤</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo mathvariant="italic">}</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M106" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M107" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> are two points in data set (<inline-formula><mml:math id="M108" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula>) and <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mtext>dist</mml:mtext><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents any distance
function. Defining the two input parameters, we can make clusters. In the clustering process data
points are classified into three groups: core points, (density) reachable points, and outliers, defined as
follows.
<list list-type="bullet"><list-item>
      <p id="d1e2423"><italic>Core point.</italic> Point A is a core point if within the distance of <inline-formula><mml:math id="M110" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> at least minPts points (including A) exist.</p></list-item><list-item>
      <p id="d1e2436"><italic>Reachable point.</italic> Point B is reachable from point A if there is a path
(<inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) from A to B (<inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>=</mml:mo></mml:mrow></mml:math></inline-formula> A). All points in the path, with the possible exception of point B, are core points.</p></list-item><list-item>
      <p id="d1e2484"><italic>Outlier point.</italic> Point C is an outlier if it is not reachable from any point.</p></list-item></list></p>
      <p id="d1e2489">In this method, an arbitrary point (that has not been visited before) is selected, and using the
above steps the neighbour points are retrieved. If the created cluster has a sufficient number of points
(larger than minPts) a cluster is started. One advantage of DBSCAN is that the method can
automatically estimate the numbers of clusters.</p>
</sec>
<sec id="Ch1.S2.SS8">
  <label>2.8</label><title>Hyper-parameter tuning</title>
      <p id="d1e2500">Machine learning methods are generally parametrized by a set of hyper-parameters, <inline-formula><mml:math id="M113" display="inline"><mml:mi mathvariant="italic">λ</mml:mi></mml:math></inline-formula>. An
optimal set <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mtext>best</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> will result in an optimal algorithm which minimizes the loss
function. This set can be formally written as follows:
            <disp-formula id="Ch1.E15" content-type="numbered"><label>15</label><mml:math id="M115" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="italic">λ</mml:mi><mml:mtext>best</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mtext>argmin</mml:mtext><mml:mo mathvariant="italic">{</mml:mo><mml:mi>L</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mtext>test</mml:mtext></mml:msub><mml:mo>;</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mi>A</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mtext>train</mml:mtext></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="italic">λ</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo><mml:mo mathvariant="italic">}</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M116" display="inline"><mml:mi>A</mml:mi></mml:math></inline-formula> is the algorithm and <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>test</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mtext>train</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> are test and training
data. Searching to find the best set of hyper-parameters is mostly done using grid search method in
which a set of values on a predefined grid is proposed. Implementing each of the proposed
hyper-parameters, the algorithm will be trained, and the prediction results will be compared. Most
algorithms have only a few hyper-parameters. Depending on the learning algorithm, the size of training, and test data sets, the grid search can be a time-consuming approach. Thus automatic hyper-parameter
optimization has gained interest; details on the topic can be found in
<xref ref-type="bibr" rid="bib1.bibx6" id="text.30"/>.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Result for supervised and unsupervised learning using the PCL system</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Supervised ML results</title>
      <p id="d1e2621">To apply supervised learning to the PCL system, we randomly chose 4500 profiles from the LR, HR, and
the nitrogen vibrational Raman channels. These measurements were taken on different nights in
different years and represent different atmospheric conditions. For the LR and HR digital Rayleigh
channels, the profiles were labelled as “bad profiles” and “good profiles”. For the nitrogen
channel we added one more label that represents profiles with traces of clouds or aerosol layers,
called “cloudy” profiles. Here, by “cloud” we mean a substantial increase in scattering relative
to a clean atmosphere, which could be caused by clouds or<?pagebreak page397?> aerosol layers. The HR and LR channels
are seldom affected by clouds or aerosols as the chopper is not fully open until about
20 <inline-formula><mml:math id="M119" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>. Furthermore, labelling the water vapour channel was not attempted for this study, due
to its high natural variability in addition to instrumental variability.</p>
      <p id="d1e2632">We used 70 % of our data for the training phase and we kept 30 % of data for the test phase
(meaning that during the training phase 30 % of data stayed isolated and the algorithm was built
without considering features of the test data). In order to overcome the overfitting issue we used
the <inline-formula><mml:math id="M120" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-fold cross-validation technique, in which the data set is divided into <inline-formula><mml:math id="M121" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> equal subsets. In
this work we used 5-fold cross-validation. The accuracy score is the ratio of correct predictions to
the total number of predictions. We used accuracy as a metric of evaluating the performance of the
algorithms. We used the Python scikit-learn package to train our ML models. The prediction scores
resulting from the cross-validation method as well as from fitting the models on the test data set
is shown in Table <xref ref-type="table" rid="Ch1.T1"/>.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e2654">Accuracy scores for the training and the test set for SVM, RF, and GBT models. Results
are shown for HR, LR, and nitrogen channels.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col4">Test set </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Channel</oasis:entry>
         <oasis:entry colname="col2">SVM</oasis:entry>
         <oasis:entry colname="col3">RF</oasis:entry>
         <oasis:entry colname="col4">GBT</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">HR</oasis:entry>
         <oasis:entry colname="col2">83 %</oasis:entry>
         <oasis:entry colname="col3">97 %</oasis:entry>
         <oasis:entry colname="col4">98 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">LR</oasis:entry>
         <oasis:entry colname="col2">88 %</oasis:entry>
         <oasis:entry colname="col3">97 %</oasis:entry>
         <oasis:entry colname="col4">97 %</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Nitrogen</oasis:entry>
         <oasis:entry colname="col2">88 %</oasis:entry>
         <oasis:entry colname="col3">95 %</oasis:entry>
         <oasis:entry colname="col4">96 %</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col4">Training set </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Channel</oasis:entry>
         <oasis:entry colname="col2">SVM</oasis:entry>
         <oasis:entry colname="col3">RF</oasis:entry>
         <oasis:entry colname="col4">GBT</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">HR</oasis:entry>
         <oasis:entry colname="col2">90 %</oasis:entry>
         <oasis:entry colname="col3">98 %</oasis:entry>
         <oasis:entry colname="col4">99 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">LR</oasis:entry>
         <oasis:entry colname="col2">90 %</oasis:entry>
         <oasis:entry colname="col3">98 %</oasis:entry>
         <oasis:entry colname="col4">98 %</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Nitrogen</oasis:entry>
         <oasis:entry colname="col2">88 %</oasis:entry>
         <oasis:entry colname="col3">94 %</oasis:entry>
         <oasis:entry colname="col4">95 %</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e2811">We also used the confusion matrix for further evaluations where the good profiles are considered as
“positive” and the bad profiles are considered as “negative”. A confusion matrix can provide us
with the number of the following:
<list list-type="bullet"><list-item>
      <p id="d1e2816">True positives (TP) – the number of profiles that are correctly labelled as positive (clean profiles);</p></list-item><list-item>
      <p id="d1e2820">False positives (FP) – the number of profiles that are incorrectly labelled as positive;</p></list-item><list-item>
      <p id="d1e2824">True negatives (TN) – the number of profiles that are correctly labelled as negative (bad profiles);</p></list-item><list-item>
      <p id="d1e2828">False negatives (FN) – the number of profiles that are incorrectly labelled as negative.</p></list-item></list></p>
      <p id="d1e2831">A perfect algorithm will result in a confusion matrix in which FP and FN are zeros. Moreover, the
precision and recall can be employed to give us an insight into how our algorithm can distinguish
between good and bad profiles. The precision and recall are defined as follows:

                <disp-formula specific-use="align" content-type="numbered"><mml:math id="M122" display="block"><mml:mtable displaystyle="true"><mml:mtr><mml:mtd><mml:mstyle class="stylechange" displaystyle="true"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>precision</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mtext>true positive</mml:mtext><mml:mrow><mml:mtext>true positive</mml:mtext><mml:mo>+</mml:mo><mml:mtext>false positive</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mlabeledtr id="Ch1.E16"><mml:mtd><mml:mtext>16</mml:mtext></mml:mtd><mml:mtd><mml:mstyle displaystyle="true" class="stylechange"/></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>recall</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mtext>true positive</mml:mtext><mml:mrow><mml:mtext>true positive</mml:mtext><mml:mo>+</mml:mo><mml:mtext>false negative</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p>
      <p id="d1e2888">The precision and recall for the nitrogen channel for each category (clear, cloud, and bad) are
shown in Table <xref ref-type="table" rid="Ch1.T2"/> as well. The GBT and RF algorithms, both have high accuracy results
on HR and LR channels. The accuracy of the model on the training set on the LR channel for both RF
and GBT are 99 % and on the test set are 98 %. The precision and recall values for the clear profiles are close to unity and for the bad profiles they are 0.95 and 0.96 respectively. The HR
channel also has a high accuracy of 99 % in the training set for both RF and GBT, and the accuracy
score in the testing set is 98 %. The precision and recall values in Table <xref ref-type="table" rid="Ch1.T2"/> are
also similar to the LR channel.</p>
      <p id="d1e2895">For the nitrogen channel the GBT algorithm has the highest accuracy of 95 %, while the RF algorithm
has accuracy of 94 %. The confusion matrix of the test result for the GBT algorithm (the one with
the highest accuracy) is shown in Fig. <xref ref-type="fig" rid="Ch1.F3"/>a. The algorithm can perform
almost perfectly on distinguishing bad profiles (only one bad scan was wrongly labelled as
cloudy). The cloud and clear profiles for most profiles are labelled correctly; however, for a few
profiles the model mislabelled clouds as clear profiles.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e2903">Precision and recall values for the nitrogen, LR, and HR channels. The precision and
recall values are calculated using the GBT model.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col3">Nitrogen channel </oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Scan type</oasis:entry>
         <oasis:entry colname="col2">Precision</oasis:entry>
         <oasis:entry colname="col3">Recall</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Cloud</oasis:entry>
         <oasis:entry colname="col2">0.94</oasis:entry>
         <oasis:entry colname="col3">0.91</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Clear</oasis:entry>
         <oasis:entry colname="col2">0.96</oasis:entry>
         <oasis:entry colname="col3">0.98</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Bad</oasis:entry>
         <oasis:entry colname="col2">1.00</oasis:entry>
         <oasis:entry colname="col3">1.00</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col3">LR channel </oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Scan type</oasis:entry>
         <oasis:entry colname="col2">Precision</oasis:entry>
         <oasis:entry colname="col3">Recall</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Clear</oasis:entry>
         <oasis:entry colname="col2">0.99</oasis:entry>
         <oasis:entry colname="col3">0.99</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Bad</oasis:entry>
         <oasis:entry colname="col2">0.96</oasis:entry>
         <oasis:entry colname="col3">0.95</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col3">HR channel </oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Scan type</oasis:entry>
         <oasis:entry colname="col2">Precision</oasis:entry>
         <oasis:entry colname="col3">Recall</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Clear</oasis:entry>
         <oasis:entry colname="col2">0.98</oasis:entry>
         <oasis:entry colname="col3">1.00</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Bad</oasis:entry>
         <oasis:entry colname="col2">0.98</oasis:entry>
         <oasis:entry colname="col3">0.94</oasis:entry>
         <oasis:entry colname="col4"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e3095">The confusion matrices for nitrogen channel <bold>(a)</bold>, LR channel <bold>(b)</bold>, and HR
channel <bold>(c)</bold>. In a perfect model, the off-diagonal elements of the confusion matrix are
zeros.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f03.png"/>

        </fig>

</sec>
<?pagebreak page398?><sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Unsupervised ML results</title>
      <p id="d1e3121">The <inline-formula><mml:math id="M123" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm clusters the measurements by means of pairwise similarity. The clustering can
differ from night to night due to atmospheric and systematic variability. On nights where most
profiles are similar, fewer clusters are seen, and on other nights when the atmospheric or the
instrument conditions are more variable, more clusters are generated. The <inline-formula><mml:math id="M124" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE makes clusters but
does not estimate how many clusters are built, and the user must then estimate the number of
clusters. To automate this procedure, after applying the <inline-formula><mml:math id="M125" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE to lidar measurements we use the
DBSCAN algorithm to estimate the number of clusters. The second step of applying DBSCAN is used to
estimate the number of generated clusters by <inline-formula><mml:math id="M126" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE.</p>
      <p id="d1e3152">To demonstrate how clustering works, we show measurements from the PCL LR channel and the nitrogen
channels. Here, we use the <inline-formula><mml:math id="M127" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE on 15 May 2012 that contains both bad and good profiles. We also
show the clustering result for the nitrogen channel on 26 May 2012. We chose this night because at
the beginning of the measurements the sky was clear but the sky became cloudy.</p>
      <p id="d1e3162">On the night of 15 May 2012, the <inline-formula><mml:math id="M128" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm generates three distinct clusters for the LR
channel (Fig. <xref ref-type="fig" rid="Ch1.F4"/>). These clusters correspond to different types of lidar return
profiles. Figure <xref ref-type="fig" rid="Ch1.F5"/>a shows all the signals for each of the clusters. The maximum number
of photon counts and the value and the height of the background counts are the identifiers between
different clusters. Thus, cluster 3 with low background counts and high maximum counts represents a
group of profiles which are labelled as good profiles in our supervised algorithms. Cluster 1
represents the profiles with lower than normal laser powers, and cluster 2 shows profiles with
extremely low laser powers. To better understand the difference between these clusters,
Fig. <xref ref-type="fig" rid="Ch1.F5"/>b shows the average signal. Furthermore, the outliers of cluster 3 (shown in
black) identify the profiles with extremely high background counts. This result is consistent with
our supervised method, in which we had good profiles (cluster 3), and bad profiles which are
profiles with lower laser power (clusters 1 and 2).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e3181">Clustering of lidar return signal type using the <inline-formula><mml:math id="M129" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm for 339 profiles from the
low-gain Rayleigh measurement channel on the night of 15 May  2012. The profiles are automatically
clustered into three different groups selected by the algorithm. Cluster 3 has some
outliers.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f04.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><label>Figure 5</label><caption><p id="d1e3199"><bold>(a)</bold> All 339 profiles collected by the PCL system LR channel on the night of 15 May
2012. The sharp cutoff for all profiles below 20 <inline-formula><mml:math id="M130" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> is due to the system's mechanical
chopper. The green signals have extremely low power. The red line represents all signals with low
return signal and the blue line indicates the signals that are considered good profiles. The black
lines are signals with extremely high backgrounds. <bold>(b)</bold> Each line represents an average of
the signals within a cluster. The red line is the average signal for profiles with lower laser
power (cluster 1). The green line is the average signal for profiles with really low laser power
(cluster 2). The blue line is the average signal for profiles with strong laser power (cluster
3). The black line indicates the outliers that have extremely high background counts and are
outliers belonging to cluster 3 (blue curve). The background counts in the green line start at
about 50 <inline-formula><mml:math id="M131" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>, whereas for the red line the background starts at almost 70 <inline-formula><mml:math id="M132" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> and
for the blue line profiles the background starts at 90 <inline-formula><mml:math id="M133" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f05.png"/>

        </fig>

      <p id="d1e3245">Using the <inline-formula><mml:math id="M134" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE, we also have clustered profiles for the nitrogen channel with the measurements
taken on 26 May 2012. This night was selected because the sky conditions changed from clear to
cloudy. The measurements from this night allows us to test our algorithm and determine how well it
can distinguish cloudy profiles from the non-cloudy profiles. The result of clustering is shown in
Fig. <xref ref-type="fig" rid="Ch1.F6"/>a in which two well-distinguished clusters are generated,
where one cluster represents the cloudy and the other represents the non-cloudy profiles. The
averaged signal for each cluster is plotted in Fig. <xref ref-type="fig" rid="Ch1.F6"/>b. Moreover,
the particle extinction profile at altitudes between 3 and 10 <inline-formula><mml:math id="M135" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> is plotted in the same
figure <xref ref-type="bibr" rid="bib1.bibx5" id="paren.31"/>. The first 130 profiles are clean and the last 70 profiles are severely affected
by thick clouds; thus the extinction profile is consistent with our <inline-formula><mml:math id="M136" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE classification result.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><label>Figure 6</label><caption><p id="d1e3280"><bold>(a)</bold> Profiles for the nitrogen channel on the night of  15 May 2012 were clustered
into two different groups using the <inline-formula><mml:math id="M137" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm. <bold>(b)</bold> The red line (cluster 2)
is the average of all signals within this cluster and indicates the profiles in which clouds are
detectable. The blue line (cluster 2) is the average of all signals within this cluster and
indicates the clear profiles (non-cloudy condition). <bold>(c)</bold> The particle extinction profile
for the night shows the last 70 profiles are affected by thick clouds at about 4.5 <inline-formula><mml:math id="M138" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>
altitude.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f06.png"/>

        </fig>

      <p id="d1e3312">The <inline-formula><mml:math id="M139" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE method can be used as a visualization tool; however, to evaluate each cluster the user
either needs to examine profiles within each cluster or use one of the aforementioned classification
methods. For example Fig. <xref ref-type="fig" rid="Ch1.F4"/> shows this night of measurement had some major
differences among the collected profiles (if all the profiles were similar only one cluster would be
generated). But, to evaluate the cluster the profiles within each cluster must be examined by a
human, or a supervised ML should be used to label each cluster.</p>
</sec>
<?pagebreak page399?><sec id="Ch1.S3.SS3">
  <label>3.3</label><?xmltex \opttitle{PCL fire detection using the $t$-SNE algorithm}?><title>PCL fire detection using the <inline-formula><mml:math id="M140" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm</title>
      <p id="d1e3340">The <inline-formula><mml:math id="M141" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE can be used for anomaly detection. As fire's smoke in the stratosphere is a relatively
rare event, we can test the algorithms to identify these events. Here, we used the <inline-formula><mml:math id="M142" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE to explore
traces of aerosol in stratosphere within one month of measurements. We expect that the <inline-formula><mml:math id="M143" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE would
generate a single cluster for a month with no trace of stratospheric aerosols that means no “anomalies” have been detected. The algorithm should generate more than one cluster in the case of
detecting stratospheric aerosols. We use the DBSCAN algorithm to automatically estimate the number of
generated profiles. In DBSCAN, most of the bad profiles will be tagged as noise (meaning that they
do not belong to any cluster). Here we are showing two examples, in one of which the stratospheric
smoke exists and our algorithm generates more than one cluster. In the other example stratospheric
smoke is not present in the profiles, and the algorithm only generates one cluster. The nightly
measurements of June 2002 are used as an example of a month in which the <inline-formula><mml:math id="M144" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE can detect anomalies
(thus more than one cluster is generated), as the lidar measurement were affected by the wildfire in
Saskatchewan, and nights of measurements in July 2007 are used as an example of nights with no high
loads of aerosol in stratosphere (only one cluster is generated).</p>
      <p id="d1e3371">The wildfires in Saskatchewan during late June and early July 2002 produced a massive amount smoke that was
transported southward. As the smoke from the fire can reach to higher altitudes (reaching to lower
stratosphere), we are interested in seeing whether we can automatically detect stratospheric aerosol layers
during wildfire events. The PCL was operational on the nights of 8, 9, 10, 19, 21, 29, and 30 June
2002.  During these nights, 1961 lidar profiles were collected in the nitrogen channel. We used the
<inline-formula><mml:math id="M145" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm to<?pagebreak page400?> examine if the algorithm can detect and cluster the profiles with the trace of
wildfire in higher altitudes, using profiles in the altitude range of 8 to 25 <inline-formula><mml:math id="M146" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>. To
automatically estimate the number of produced clusters we used the DBSCAN algorithm. We set the
minPts condition to 30, and the <inline-formula><mml:math id="M147" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> value to 3. The DBSCAN algorithm estimated four clusters and
few profiles remained as the noise, which do not belong to any other clusters (shown in cyan, in
Fig. <xref ref-type="fig" rid="Ch1.F7"/>). To investigate whether profiles with layers are clustered together the
particle extinction profile for each generated cluster is plotted (Fig. <xref ref-type="fig" rid="Ch1.F8"/>). Most
of the profiles in cluster 1 are clean and no sign of particles can be seen in these profiles
(Fig. <xref ref-type="fig" rid="Ch1.F8"/>a), cluster 2 contains all profiles with mostly small traces of aerosol
between 10 and 14 <inline-formula><mml:math id="M148" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="Ch1.F8"/>b). The presence of high
loads of aerosol can clearly detected in the particle extinction profiles for both cluster 3 and 4;
the difference between the two clusters is in the height at which the presence of aerosol layer is
more distinguished (Fig. <xref ref-type="fig" rid="Ch1.F8"/>d and c). Profiles in the last two clusters belong to
the last two nights of measurements on June 2002 which are coincidental with the smoke being
transported to London from the wildfire in Saskatchewan.</p>
      <p id="d1e3415">We also examined a total of 2637 profiles in the altitude range of 8 to 25 <inline-formula><mml:math id="M149" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> obtained
from 10 nights of measurements in July 2007. As we expected, no anomalies were detected
(Fig. <xref ref-type="fig" rid="Ch1.F9"/>). The particle extinction profile of July 2007 also indicates that at the
altitude range of 8 to 25 <inline-formula><mml:math id="M150" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> no aerosol load exists (Fig. <xref ref-type="fig" rid="Ch1.F9"/>).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7"><?xmltex \currentcnt{7}?><label>Figure 7</label><caption><p id="d1e3441">Profiles for the nitrogen channel for the nights of June 2002 were clustered into four different groups using the <inline-formula><mml:math id="M151" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE algorithm. The small cluster in cyan indicates the group of profiles that do not belong to any of the other clusters in the DBSCAN algorithm.</p></caption>
          <?xmltex \igopts{width=170.716535pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f07.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><label>Figure 8</label><caption><p id="d1e3459">The <inline-formula><mml:math id="M152" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE generates four clusters for profiles of nitrogen channels in July
2007. <bold>(a)</bold> Most of the profiles are clean and no sign of particles can be
seen. <bold>(b)</bold> Profiles with mostly small traces of aerosol between 10 and 14 <inline-formula><mml:math id="M153" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>. <bold>(c, d)</bold> The presence of high loads of aerosol can clearly be detected.</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f08.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9" specific-use="star"><?xmltex \currentcnt{9}?><label>Figure 9</label><caption><p id="d1e3494"><bold>(a)</bold> The particle extinction profile of July 2007 indicates no
significant trace of stratospheric aerosols. <bold>(b)</bold> The <inline-formula><mml:math id="M154" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE generates a single cluster for all of 2637 profiles of nitrogen
channel for July 2007.</p></caption>
          <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/391/2021/amt-14-391-2021-f09.png"/>

        </fig>

      <p id="d1e3515">Thus, using the <inline-formula><mml:math id="M155" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE method we can detect anomalies in the UTLS. In the UTLS region, for the clear
atmosphere we expect to see a single cluster, and when aerosol loads exist at least two clusters
will be generated. We are implementing the <inline-formula><mml:math id="M156" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE on one month of measurements, and when the
algorithm generates more than one cluster we examine profiles within that cluster. However, at the
moment, because we only use the Raman channel it is not possible for us to distinguish between smoke
traces and cirrus clouds (unless the trace is detected in altitudes above 14 <inline-formula><mml:math id="M157" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>, similar to Fig. <xref ref-type="fig" rid="Ch1.F8"/> where we are more confident in claiming that the detected aerosol layers are
traces of smoke, as shown in <xref ref-type="bibr" rid="bib1.bibx8" id="altparen.32"/>).</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Summary and conclusion</title>
      <p id="d1e3554">We introduced a machine learning method to classify raw lidar (level-0) measurements. We used
different ML methods on elastic and inelastic measurements from the PCL lidar systems. The ML
methods we used and our results are summarized as follows.
<list list-type="order"><list-item>
      <p id="d1e3559">We tested different supervised ML algorithms, among which the RF and the GBT performed better,
with a success rate above 90 % for the PCL system.</p></list-item><list-item>
      <p id="d1e3563">The <inline-formula><mml:math id="M158" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE unsupervised algorithm can successfully cluster profiles on nights with both
consistent and varying lidar profiles due to both atmospheric conditions and system
alignment and performance. For example, if during the measurements the laser power dropped or clouds
became present, the <inline-formula><mml:math id="M159" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE showed different clusters representing these conditions.</p></list-item><list-item>
      <p id="d1e3581">Unlike the traditional method of defining a fixed threshold for the background counts, in
supervised ML approach the machine can distinguish high background counts by looking at the labels
of the training set. In the unsupervised ML approach, by looking at the similarities between the
two profiles and defining a distance scale, good profiles will be grouped together. High
background counts can be grouped in a smaller group. Most of the time the number of bad profiles
are small; thus they will be labelled as noise.</p></list-item></list></p>
      <p id="d1e3584">We successfully implemented supervised and unsupervised ML algorithms to classify lidar measurement
profiles. The ML is a robust method with high accuracy that enables us to precisely classify
thousands of lidar profiles within a short period of time. Thus, with accuracy of higher than 95 %
this method has a significant advantage over previous methods of classifying. For example, in the
supervised ML, we train the machine by showing (labelling) different profiles in different
conditions. When the machine has seen enough examples of each class (which is a small fraction of the
entire database), it can classify the un-labelled profiles with no need to pre-define any condition
for the system. Furthermore, in the unsupervised learning method, no labelling is needed, and the
whole classification is free from subjective biases of the individual marking the profiles (which is important
for large atmospheric data sets ranging over decades). Using ML avoids the problem of
different observers classifying profiles differently. We also showed that the unsupervised schema
has the potential to be used as an anomaly detector, which<?pagebreak page401?> can alert us when there is a trace of
aerosol in the UTLS region. We are planning to expand our unsupervised learning method to both
Rayleigh and nitrogen channels to be able to correctly identify and distinguish cirrus clouds from
smoke traces in the UTLS. Our results indicate that ML is a powerful technique that can be used in
lidar classifications. We encourage our colleagues in the lidar community to use both supervised and
unsupervised ML algorithms for their lidar profiles. For the supervised learning the GBT performs
exceptionally well, and the unsupervised learning has the potential of sorting anomalies.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e3592">The data used in this paper are publically available at
<uri>https://www.ndaccdemo.org/stations/london-ontario-canada</uri> (last access: 8 January 2021, <xref ref-type="bibr" rid="bib1.bibx16" id="altparen.33"/>) by clicking the DataLink button or via FTP at <uri>http://ftp.cpc.ncep.noaa.gov/ndacc/station/londonca/hdf/lidar/</uri> (last access: 8 January 2021). The data used in this study are also available from Robert Sica at The University of Western Ontario (sica@uwo.ca).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e3607">GF adopted, implemented, and interpreted machine learning in the context of lidar measurements and wrote the paper. RS supervised the PhD thesis, participated in the writing of paper, and provided lidar measurements used in the paper. MD helped with implementation and interpretation of the results.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e3613">The authors declare that they have no conflict of interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e3619">We would like to thank Sepideh Farsinejad for many interesting discussions about clustering
methods and statistics. We would like to thank Shayamila Mahagammulla Gamage for her inputs on
labelling methods, and Robin Wing for our chats about other statistical methods used for lidar
classification. We also recognize Dakota Cecil's insights and help in the labelling process of the
raw lidar profiles.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e3624">This paper was edited by Joanna Joiner and reviewed by Benoît Crouzy and one
anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Bishop(2006)</label><?label bishop2006pattern?><mixed-citation>
Bishop, C. M.: Pattern recognition and machine learning, Springer-Verlag, New York, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Breiman(2002)</label><?label breiman2002manual?><mixed-citation>
Breiman, L.: Random Forests, Mach. Learn., 45, 5–32, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Burges(1998)</label><?label burges1998tutorial?><mixed-citation> Burges, C. J.: A tutorial on support vector machines
for pattern recognition, Data Mining Knowledge Discovery, 2, 121–167, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Christian et al.(2019)</label><?label christianradiative?><mixed-citation>
Christian, K., Wang, J., Ge, C., Peterson, D., Hyer, E., Yorks, J., and McGill, M.: Radiative Forcing and Stratospheric Warming of Pyrocumulonimbus Smoke Aerosols: First Modeling Results With Multisensor (EPIC, CALIPSO, and CATS) Views from Space, Geophys. Res. Lett., 46,  10061–10071, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Doucet(2009)</label><?label paul?><mixed-citation>
Doucet, P. J.: First aerosol measurements with the Purple Crow Lidar: lofted particulate matter straddling the stratospheric boundary, Master's thesis, The University of Western Ontario, London, ON, Canada, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Feurer and Hutter(2019)</label><?label feurer2019hyperparameter?><mixed-citation> Feurer, M. and Hutter, F.:
Hyperparameter optimization, in: Automated Machine Learning, Springer, Cham, 3–33, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Foody and Mathur(2004)</label><?label SVS?><mixed-citation>
Foody, G. M. and Mathur, A.: A relative evaluation of multiclass image classification by support vector machines, IEEE T. Geosci. Remote Sens., 42,
1335–1343, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Fromm et al.(2010)Fromm, Lindsey, Servranckx, Yue, Trickl, Sica, Doucet, and
Godin-Beekmann</label><?label fromm2010untold?><mixed-citation> Fromm, M., Lindsey, D. T., Servranckx, R., Yue, G., Trickl,
T., Sica, R., Doucet, P., and Godin-Beekmann, S.: The untold story of pyrocumulonimbus,
B. Am. Meteorol. Soc., 91, 1193–1210, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Hastie et al.(2009)</label><?label unsupervised?><mixed-citation>
Hastie, T., Tibshirani, R., and Friedman, J.: Unsupervised learning, in: The elements of statistical learning, Springer Series in Statistics, New York, Chap. 14, 485–585, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Hinton and Roweis(2002)</label><?label hinton?><mixed-citation>Hinton, G. E. and Roweis, S. T.: Stochastic neighbor embedding, Advances in neural information processing systems, 15, 857–864, 2002.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx11"><label>Knerr et al.(1990)</label><?label knerr1990single?><mixed-citation> Knerr, S., Lé, P., and Dreyfus, G.: Single-layer
learning revisited: a stepwise procedure for building and training a neural network, in:
Neurocomputing, Springer, Berlin, Heidelberg, 41–50, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Lerman and Yitzhaki(1984)</label><?label lerman1984note?><mixed-citation> Lerman, R. I. and Yitzhaki, S.: A note on the
calculation and interpretation of the Gini index, Econ. Lett., 15, 363–368, 1984.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Liaw et al.(2002)Liaw, Wiener et al.</label><?label liaw2002classification?><mixed-citation> Liaw, A., Wiener, M.,
et al.: Classification and regression by randomForest, R News, 2, 18–22, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Maaten and Hinton(2008)</label><?label maaten2008visualizing?><mixed-citation>Maaten, L. and Hinton, G.: Visualizing
data using <inline-formula><mml:math id="M160" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-SNE, J. Machine Learn. Res., 9, 2579–2605, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Mantero et al.(2005)Mantero, Moser, and Serpico</label><?label SVm?><mixed-citation> Mantero, P., Moser, G., and
Serpico, S. B.: Partially supervised classification of remote sensing images through SVM-based
probability density estimation, IEEE T. Geosci. Remote Sens., 43, 559–570, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>NDACC(2021)</label><?label NDACC2021?><mixed-citation>NDACC: NDACC Measurements at the London, Ontario, Canada Station, NDACC, available at: <uri>https://www.ndaccdemo.org/stations/london-ontario-canada</uri> or via ftp at:
<uri>http://ftp.cpc.ncep.noaa.gov/ndacc/station/londonca/hdf/lidar/</uri>, last access: 8 January 2021.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Nicolae et al.(2018)Nicolae, Vasilescu, Talianu, Binietoglou,
Nicolae, Andrei, and Antonescu</label><?label NeuralNetlidar?><mixed-citation>Nicolae, D., Vasilescu, J., Talianu, C., Binietoglou, I., Nicolae, V., Andrei, S., and Antonescu, B.: A neural network aerosol-typing algorithm based on lidar data, Atmos. Chem. Phys., 18, 14511–14537, <ext-link xlink:href="https://doi.org/10.5194/acp-18-14511-2018" ext-link-type="DOI">10.5194/acp-18-14511-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Quinlan(1986)</label><?label quinlan1986induction?><mixed-citation> Quinlan, J. R.: Induction of decision trees, Machine
Learn., 1, 81–106, 1986.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Robert and Casella(2004)</label><?label Casella?><mixed-citation> Robert, C. P. and Casella, G.: Monte Carlo Statistical
Methods, Springer Texts in Statistics, Springer science &amp; business media, New York, NY, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Schapire(1990)</label><?label schapire1990strength?><mixed-citation> Schapire, R. E.: The strength of weak learnability,
Machine Learn., 5, 197–227, 1990.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Shannon(1948)</label><?label shannon1948mathematical?><mixed-citation> Shannon, C.: A mathematical theory of
communication, Bell Syst. Techn. J., 27, 379–423, 1948.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Sica et al.(1995)Sica, Sargoytchev, Argall, Borra, Girard, Sparrow, and
Flatt</label><?label sica1995lidar?><mixed-citation> Sica, R., Sargoytchev, S., Argall, P. S., Borra, E. F., Girard, L.,
Sparrow, C. T., and Flatt, S.: Lidar measurements taken with a large-aperture liquid
mirror. 1. Rayleigh-scatter system, Appl. Opt., 34, 6925–6936, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Vapnik(2013)</label><?label vapnik2013nature?><mixed-citation> Vapnik, V.: The nature of statistical learning theory,
Springer Science &amp; Business Media, Springer-Verlag New York, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Wing et al.(2018)Wing, Hauchecorne, Keckhut, Godin-Beekmann, Khaykin,
McCullough, Mariscal, and d'Almeida</label><?label Robin?><mixed-citation>Wing, R., Hauchecorne, A., Keckhut, P., Godin-Beekmann, S., Khaykin, S., McCullough, E. M., Mariscal, J.-F., and d'Almeida, É.: Lidar temperature series in the middle atmosphere as a reference data set – Part 1: Improved retrievals and a 20-year cross-validation of two co-located French lidars, Atmos. Meas. Tech., 11, 5531–5547, <ext-link xlink:href="https://doi.org/10.5194/amt-11-5531-2018" ext-link-type="DOI">10.5194/amt-11-5531-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Zeng et al.(2019)Zeng, Vaughan, Liu, Trepte, Kar, Omar, Winker,
Lucker, Hu, Getzewich, and Avery</label><?label FuzzyKmeans?><mixed-citation>Zeng, S., Vaughan, M., Liu, Z., Trepte, C., Kar, J., Omar, A., Winker, D., Lucker, P., Hu, Y., Getzewich, B., and Avery, M.: Application of high-dimensional fuzzy <inline-formula><mml:math id="M161" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means cluster analysis to CALIOP/CALIPSO version 4.1 cloud–aerosol discrimination, Atmos. Meas. Tech., 12, 2261–2285, <ext-link xlink:href="https://doi.org/10.5194/amt-12-2261-2019" ext-link-type="DOI">10.5194/amt-12-2261-2019</ext-link>, 2019.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Classification of lidar measurements using supervised and unsupervised machine learning methods</article-title-html>
<abstract-html><p>While it is relatively straightforward to automate the processing of lidar signals, it is more
difficult to choose periods of <q>good</q> measurements to process. Groups use various ad hoc
procedures involving either very simple (e.g. signal-to-noise ratio) or more complex procedures
(e.g. Wing et al., 2018) to perform a task that is easy to train humans to perform but is time-consuming. Here, we use machine learning techniques to train the machine to sort the measurements
before processing. The presented method is generic and can be applied to most lidars. We test the
techniques using measurements from the Purple Crow Lidar (PCL) system located in London,
Canada. The PCL has over 200&thinsp;000 raw profiles in Rayleigh and Raman channels available for
classification. We classify raw (level-0) lidar measurements as <q>clear</q> sky profiles with strong
lidar returns, <q>bad</q> profiles, and profiles which are significantly influenced by clouds or
aerosol loads.  We examined different supervised machine learning algorithms including the random
forest, the support vector machine, and the gradient boosting trees, all of which can successfully
classify profiles. The algorithms were trained using about 1500 profiles for each PCL channel,
selected randomly from different nights of measurements in different years. The success rate of identification for all the channels is above 95&thinsp;%.  We also used the <i>t</i>-distributed stochastic embedding (<i>t</i>-SNE) method, which is an unsupervised algorithm, to cluster our lidar profiles. Because the <i>t</i>-SNE is a data-driven method in which no labelling of the training set is needed, it is an attractive algorithm to find anomalies in lidar profiles. The method has been tested on several nights of measurements from the PCL measurements. The <i>t</i>-SNE can successfully
cluster the PCL data profiles into meaningful categories. To demonstrate the use of the technique,
we have used the algorithm to identify stratospheric aerosol layers due to wildfires.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Bishop(2006)</label><mixed-citation>
Bishop, C. M.: Pattern recognition and machine learning, Springer-Verlag, New York, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Breiman(2002)</label><mixed-citation>
Breiman, L.: Random Forests, Mach. Learn., 45, 5–32, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Burges(1998)</label><mixed-citation> Burges, C. J.: A tutorial on support vector machines
for pattern recognition, Data Mining Knowledge Discovery, 2, 121–167, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Christian et al.(2019)</label><mixed-citation>
Christian, K., Wang, J., Ge, C., Peterson, D., Hyer, E., Yorks, J., and McGill, M.: Radiative Forcing and Stratospheric Warming of Pyrocumulonimbus Smoke Aerosols: First Modeling Results With Multisensor (EPIC, CALIPSO, and CATS) Views from Space, Geophys. Res. Lett., 46,  10061–10071, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Doucet(2009)</label><mixed-citation>
Doucet, P. J.: First aerosol measurements with the Purple Crow Lidar: lofted particulate matter straddling the stratospheric boundary, Master's thesis, The University of Western Ontario, London, ON, Canada, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Feurer and Hutter(2019)</label><mixed-citation> Feurer, M. and Hutter, F.:
Hyperparameter optimization, in: Automated Machine Learning, Springer, Cham, 3–33, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Foody and Mathur(2004)</label><mixed-citation>
Foody, G. M. and Mathur, A.: A relative evaluation of multiclass image classification by support vector machines, IEEE T. Geosci. Remote Sens., 42,
1335–1343, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Fromm et al.(2010)Fromm, Lindsey, Servranckx, Yue, Trickl, Sica, Doucet, and
Godin-Beekmann</label><mixed-citation> Fromm, M., Lindsey, D. T., Servranckx, R., Yue, G., Trickl,
T., Sica, R., Doucet, P., and Godin-Beekmann, S.: The untold story of pyrocumulonimbus,
B. Am. Meteorol. Soc., 91, 1193–1210, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Hastie et al.(2009)</label><mixed-citation>
Hastie, T., Tibshirani, R., and Friedman, J.: Unsupervised learning, in: The elements of statistical learning, Springer Series in Statistics, New York, Chap. 14, 485–585, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Hinton and Roweis(2002)</label><mixed-citation>
Hinton, G. E. and Roweis, S. T.: Stochastic neighbor embedding, Advances in neural information processing systems, 15, 857–864, 2002.

</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Knerr et al.(1990)</label><mixed-citation> Knerr, S., Lé, P., and Dreyfus, G.: Single-layer
learning revisited: a stepwise procedure for building and training a neural network, in:
Neurocomputing, Springer, Berlin, Heidelberg, 41–50, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Lerman and Yitzhaki(1984)</label><mixed-citation> Lerman, R. I. and Yitzhaki, S.: A note on the
calculation and interpretation of the Gini index, Econ. Lett., 15, 363–368, 1984.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Liaw et al.(2002)Liaw, Wiener et al.</label><mixed-citation> Liaw, A., Wiener, M.,
et al.: Classification and regression by randomForest, R News, 2, 18–22, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Maaten and Hinton(2008)</label><mixed-citation> Maaten, L. and Hinton, G.: Visualizing
data using <i>t</i>-SNE, J. Machine Learn. Res., 9, 2579–2605, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Mantero et al.(2005)Mantero, Moser, and Serpico</label><mixed-citation> Mantero, P., Moser, G., and
Serpico, S. B.: Partially supervised classification of remote sensing images through SVM-based
probability density estimation, IEEE T. Geosci. Remote Sens., 43, 559–570, 2005.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>NDACC(2021)</label><mixed-citation>
NDACC: NDACC Measurements at the London, Ontario, Canada Station, NDACC, available at: <a href="https://www.ndaccdemo.org/stations/london-ontario-canada" target="_blank"/> or via ftp at:
<a href="http://ftp.cpc.ncep.noaa.gov/ndacc/station/londonca/hdf/lidar/" target="_blank"/>, last access: 8 January 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Nicolae et al.(2018)Nicolae, Vasilescu, Talianu, Binietoglou,
Nicolae, Andrei, and Antonescu</label><mixed-citation>
Nicolae, D., Vasilescu, J., Talianu, C., Binietoglou, I., Nicolae, V., Andrei, S., and Antonescu, B.: A neural network aerosol-typing algorithm based on lidar data, Atmos. Chem. Phys., 18, 14511–14537, <a href="https://doi.org/10.5194/acp-18-14511-2018" target="_blank">https://doi.org/10.5194/acp-18-14511-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Quinlan(1986)</label><mixed-citation> Quinlan, J. R.: Induction of decision trees, Machine
Learn., 1, 81–106, 1986.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Robert and Casella(2004)</label><mixed-citation> Robert, C. P. and Casella, G.: Monte Carlo Statistical
Methods, Springer Texts in Statistics, Springer science &amp; business media, New York, NY, 2004.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Schapire(1990)</label><mixed-citation> Schapire, R. E.: The strength of weak learnability,
Machine Learn., 5, 197–227, 1990.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Shannon(1948)</label><mixed-citation> Shannon, C.: A mathematical theory of
communication, Bell Syst. Techn. J., 27, 379–423, 1948.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Sica et al.(1995)Sica, Sargoytchev, Argall, Borra, Girard, Sparrow, and
Flatt</label><mixed-citation> Sica, R., Sargoytchev, S., Argall, P. S., Borra, E. F., Girard, L.,
Sparrow, C. T., and Flatt, S.: Lidar measurements taken with a large-aperture liquid
mirror. 1. Rayleigh-scatter system, Appl. Opt., 34, 6925–6936, 1995.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Vapnik(2013)</label><mixed-citation> Vapnik, V.: The nature of statistical learning theory,
Springer Science &amp; Business Media, Springer-Verlag New York, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Wing et al.(2018)Wing, Hauchecorne, Keckhut, Godin-Beekmann, Khaykin,
McCullough, Mariscal, and d'Almeida</label><mixed-citation>
Wing, R., Hauchecorne, A., Keckhut, P., Godin-Beekmann, S., Khaykin, S., McCullough, E. M., Mariscal, J.-F., and d'Almeida, É.: Lidar temperature series in the middle atmosphere as a reference data set – Part 1: Improved retrievals and a 20-year cross-validation of two co-located French lidars, Atmos. Meas. Tech., 11, 5531–5547, <a href="https://doi.org/10.5194/amt-11-5531-2018" target="_blank">https://doi.org/10.5194/amt-11-5531-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Zeng et al.(2019)Zeng, Vaughan, Liu, Trepte, Kar, Omar, Winker,
Lucker, Hu, Getzewich, and Avery</label><mixed-citation>
Zeng, S., Vaughan, M., Liu, Z., Trepte, C., Kar, J., Omar, A., Winker, D., Lucker, P., Hu, Y., Getzewich, B., and Avery, M.: Application of high-dimensional fuzzy <i>k</i>-means cluster analysis to CALIOP/CALIPSO version 4.1 cloud–aerosol discrimination, Atmos. Meas. Tech., 12, 2261–2285, <a href="https://doi.org/10.5194/amt-12-2261-2019" target="_blank">https://doi.org/10.5194/amt-12-2261-2019</a>, 2019.
</mixed-citation></ref-html>--></article>
