<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">AMT</journal-id><journal-title-group>
    <journal-title>Atmospheric Measurement Techniques</journal-title>
    <abbrev-journal-title abbrev-type="publisher">AMT</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Atmos. Meas. Tech.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1867-8548</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/amt-14-4335-2021</article-id><title-group><article-title>Deriving boundary layer height from aerosol lidar using <?xmltex \hack{\break}?> machine learning: KABL and ADABL algorithms</article-title><alt-title>ABL from aerosol lidar with machine learning</alt-title>
      </title-group><?xmltex \runningtitle{ABL from aerosol lidar with machine learning}?><?xmltex \runningauthor{T.~Rieutord et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Rieutord</surname><given-names>Thomas</given-names></name>
          <email>thomas.rieutord@meteo.fr</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Aubert</surname><given-names>Sylvain</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Machado</surname><given-names>Tiago</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-7705-5953</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>CNRM, Université de Toulouse, Météo-France, CNRS, Toulouse, France</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Météo-France, Direction des Systèmes d'Observation, Toulouse, France</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Thomas Rieutord (thomas.rieutord@meteo.fr)</corresp></author-notes><pub-date><day>11</day><month>June</month><year>2021</year></pub-date>
      
      <volume>14</volume>
      <issue>6</issue>
      <fpage>4335</fpage><lpage>4353</lpage>
      <history>
        <date date-type="received"><day>6</day><month>March</month><year>2020</year></date>
           <date date-type="rev-request"><day>7</day><month>April</month><year>2020</year></date>
           <date date-type="rev-recd"><day>16</day><month>April</month><year>2021</year></date>
           <date date-type="accepted"><day>21</day><month>April</month><year>2021</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2021 Thomas Rieutord et al.</copyright-statement>
        <copyright-year>2021</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021.html">This article is available from https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021.html</self-uri><self-uri xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021.pdf">The full text article is available as a PDF file from https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e107">The atmospheric boundary layer height (BLH) is a key parameter for many meteorological applications, including air quality forecasts.
Several algorithms have been proposed to automatically estimate BLH from lidar backscatter profiles. However recent advances in computing have enabled new approaches using machine learning that are seemingly well suited to this problem. Machine learning can handle complex classification problems and can be trained by a human expert. This paper describes and compares two machine-learning methods, the <inline-formula><mml:math id="M1" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means unsupervised algorithm and the AdaBoost supervised algorithm, to derive BLH from lidar backscatter profiles. The <inline-formula><mml:math id="M2" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means for Atmospheric Boundary Layer (KABL) and AdaBoost for Atmospheric Boundary Layer (ADABL) algorithm codes used in this study are free and open source. Both methods were compared to reference BLHs derived from colocated radiosonde data over a 2-year period (2017–2018) at two Météo-France operational network sites (Trappes and Brest).
A large discrepancy between the root-mean-square error (RMSE) and correlation with radiosondes was observed between the two sites. At the Trappes site, KABL and ADABL outperformed the manufacturer's algorithm, while the performance was clearly reversed at the Brest site. We conclude that ADABL is a promising algorithm (RMSE of 550 m at Trappes, 800 m for manufacturer) but has training issues that need to be resolved; KABL has a lower performance (RMSE of 800 m at Trappes) than ADABL but is much more versatile.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

      <?xmltex \hack{\newpage}?>
<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e135">The atmospheric boundary layer is the lowest part of the troposphere and is the region that is directly influenced by surface forcings. It is the layer within which most human activities take place, and all pollutants emitted at ground level are dispersed within this layer. The key parameter used to model this dilution is the depth of this layer, i.e., the boundary layer height (BLH). Because BLH can vary from a few tens of meters to approximately 2 km within a single day, the volume available for the dilution of pollutants can vary considerably and is a crucial parameter for reliable warnings of poor air quality <xref ref-type="bibr" rid="bib1.bibx54 bib1.bibx17" id="paren.1"/>. However, BLH is one of the largest sources of uncertainty in air quality models <xref ref-type="bibr" rid="bib1.bibx37" id="paren.2"/>, and there is a need to better evaluate this parameter <xref ref-type="bibr" rid="bib1.bibx1" id="paren.3"/>. Accurate representation of the physical processes within the boundary layer is also important for numerical weather prediction models <xref ref-type="bibr" rid="bib1.bibx50" id="paren.4"/>. In the study of physical processes in the boundary layer, with large eddy simulations <xref ref-type="bibr" rid="bib1.bibx34" id="paren.5"/> or with measurements <xref ref-type="bibr" rid="bib1.bibx5" id="paren.6"/>, BLH is often used as a normalization of the vertical profiles. Therefore, it is important to compare BLH calculated in models with that derived from measurements.</p>
      <?pagebreak page4336?><p id="d1e157">However, measuring BLH is not straightforward. As stated in <xref ref-type="bibr" rid="bib1.bibx47" id="text.7"/>, there are no systems that meet all of the requirements for making reliable BLH estimates. The best estimate of BLH can be achieved via the synergistic use of multiple instruments, but adding instruments limits the number of sites where estimates can be made.
In this paper, we focus on a single instrument, an aerosol lidar (see Sect. <xref ref-type="sec" rid="Ch1.S2.SS1.SSS1"/> for more information), that is already widely used <xref ref-type="bibr" rid="bib1.bibx23" id="paren.8"/>. Aerosol lidars are active remote sensing instruments that emit a laser pulse into the atmosphere and measure the amount of light backscattered from aerosols as a function of the vertical range from the instrument. Because aerosols are more concentrated in the boundary layer than in the overlying free troposphere, there is often a sharp decrease in the backscatter profile between these two layers. However, this decrease can be blurred or perturbed by other strong signals (e.g., clouds, aerosols residing in elevated or residual layers, and small-scale structures) and instrumental noise. For these reasons, numerous studies exist concerning the derivation of BLH from aerosol lidar. <xref ref-type="bibr" rid="bib1.bibx35" id="text.9"/> use a simple thresholding of the signal. Other methods are based on calculations of the derivative function of the backscatter profile. For example, <xref ref-type="bibr" rid="bib1.bibx25" id="text.10"/> take the minimum of the gradient, <xref ref-type="bibr" rid="bib1.bibx36" id="text.11"/> use the height where the second derivative is zero (the inflection point) as well as the variance of the signal, and <xref ref-type="bibr" rid="bib1.bibx52" id="text.12"/> use the derivative of the logarithm of the backscattered signal. One of the most used methods is the wavelet covariance transform, which searches for the maximum in the convolution between the backscatter profile and a Haar wavelet <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx11 bib1.bibx6" id="paren.13"/>. More recent studies have been based on backscatter signal analysis such as the Structure Of The Atmosphere (STRAT) algorithm <xref ref-type="bibr" rid="bib1.bibx38" id="paren.14"/> and the Characterising the Atmospheric Boundary layer based on ALC Measurements (CABAM, where ALC is automatic lidar and ceilometer) algorithm <xref ref-type="bibr" rid="bib1.bibx31" id="paren.15"/>. Graph theory has also been used to impose continuity constraints (both vertically and in time) in BLH estimates, e.g., the pathfinder algorithm <xref ref-type="bibr" rid="bib1.bibx15" id="paren.16"/>. Inspired by image processing, some methods use Canny edge detection in addition to backscatter signal analysis <xref ref-type="bibr" rid="bib1.bibx38 bib1.bibx23" id="paren.17"/>. An extension of pathfinder including the detection of the continuous aerosol layer was made in PathfinderTURB <xref ref-type="bibr" rid="bib1.bibx41" id="paren.18"/>. These studies demonstrate that estimating BLH from aerosol lidar is still an open area of research.</p>
      <p id="d1e200">In addition, artificial intelligence (AI), as a set of techniques aiming to reproduce human intelligence with machines, has reemerged in the last decade because of the simultaneous increase in the number of available data and amount of computational power. Both have reached levels that enable previously impractical applications. AI is capable of tackling complex classification problems, especially in image classification <xref ref-type="bibr" rid="bib1.bibx32" id="paren.19"/>.
Such breakthroughs were made possible by deep convolutional neural networks <xref ref-type="bibr" rid="bib1.bibx33" id="paren.20"/>; however, AI encompasses many other techniques that also benefit from larger datasets and increased computational power <xref ref-type="bibr" rid="bib1.bibx3" id="paren.21"/>. In this paper, we explore how the estimation of BLH from backscatter profiles can be formulated as a classification problem and how appropriate algorithms can be applied to solve this problem.
Machine-learning techniques are categorized into two broad families: supervised learning (mimicking a reliable reference) and unsupervised learning <xref ref-type="bibr" rid="bib1.bibx24" id="paren.22"><named-content content-type="pre">learning without a reference;</named-content></xref>.
<xref ref-type="bibr" rid="bib1.bibx56" id="text.23"/> have already described a method that falls within the scope of AI. They used unsupervised learning to classify whether measurement points were within the boundary layer. This method has yielded convincing results in previous studies <xref ref-type="bibr" rid="bib1.bibx57 bib1.bibx7 bib1.bibx43" id="paren.24"/> and is pursued here using the <inline-formula><mml:math id="M3" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means for Atmospheric Boundary Layer (KABL) algorithm. KABL has been extensively tested and is shared via an open-source code. In addition, we test an alternative adaptive boosting (AdaBoost) machine-learning algorithm, the AdaBoost for Atmospheric Boundary Layer (ADABL) algorithm. Both algorithms classify whether the measurement points are inside or outside of the boundary layer; however, ADABL learns the characteristics of both groups from a training set. The training set consists of atmospheric boundary layer identifications made by human experts, which is acknowledged as being more reliable than available automatic methods <xref ref-type="bibr" rid="bib1.bibx47" id="paren.25"/>. Algorithms classifying from a reference dataset (e.g., ADABL) are called supervised algorithms, while algorithms classifying without a reference dataset (e.g., KABL) are called unsupervised algorithms.
Supervised algorithms make it possible to automatically reproduce human expertise in boundary layer identification. To our knowledge, this is the first time that a supervised algorithm has been applied to this problem. This study is of practical interest because it includes the publication of the source code, which uses only free software.</p>
      <p id="d1e234">In Sect. <xref ref-type="sec" rid="Ch1.S2"/>, we describe the data used in this study, i.e., the lidar data in the algorithm inputs, reference radiosonde data, and ancillary data used to sort the meteorological conditions. In Sect. <xref ref-type="sec" rid="Ch1.S3"/>, we describe the two machine-learning algorithms (KABL and ADABL) and the procedures used to evaluate them. In Sect. <xref ref-type="sec" rid="Ch1.S4"/>, we present the results of our study, which consists of a sensitivity analysis of the KABL algorithm, a comparison of the methods with the radiosonde data over a 2-year period, and a case study. In Sect. <xref ref-type="sec" rid="Ch1.S5"/>, we discuss the results, limitations, and prospects of our study. The final section is dedicated to the conclusions that can be drawn from our study.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Material</title>
      <p id="d1e253">Our study used data from the Météo-France operational network. We used colocated radiosonde and aerosol lidar data over two sites: Brest (a coastal city in the northwestern region of France) and Trappes (an inland suburban area within Paris). The dataset spanned 2 years: 2017 and 2018. A case study was conducted on 2 August 2018 for the Trappes site.</p><?xmltex \hack{\newpage}?>
<?pagebreak page4337?><sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Lidar data</title>
<sec id="Ch1.S2.SS1.SSS1">
  <label>2.1.1</label><title>Lidar network</title>
      <p id="d1e271">Since 2016, Météo-France has deployed a network of six automatic backscatter lidars to help the Toulouse Volcanic Ash Advisory Centre characterize layers of volcanic ash and aerosol in the atmosphere. One of the six sensors can be quickly redeployed at a more suitable geographic location depending on the transport event being tracked. The network, fully operational since April 2017, is continuously functioning and has detected aerosol events at altitudes of up to 17 km. It is part of the wider automatic lidar and ceilometer network of the E-PROFILE program described in <xref ref-type="bibr" rid="bib1.bibx22" id="text.26"/>.</p>
      <p id="d1e277">Two sampling sites in this network were selected: Brest (48.444<inline-formula><mml:math id="M4" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 4.412<inline-formula><mml:math id="M5" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W; 94 m a.s.l. – meters above sea level) and Trappes (48.773<inline-formula><mml:math id="M6" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 2.0124<inline-formula><mml:math id="M7" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> E; 166 m a.s.l.). Both sites are equipped with a Mini Micro Pulse LiDAR (MiniMPL), built by Sigma Space Corporation with an exterior casing provided by Envicontrol. A MiniMPL unit from the Météo-France network is shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/>. MiniMPL is a compact version of the micro-pulse lidar (MPL) systems approved for the global NASA Micro-Pulse Lidar Network (MPLNET). A comprehensive description of MiniMPL can be found in <xref ref-type="bibr" rid="bib1.bibx58" id="text.27"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e324">Mini Micro Pulse LiDAR (MiniMPL) unit from the Météo-France network.</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f01.jpg"/>

          </fig>

</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <label>2.1.2</label><title>Data processing</title>
      <p id="d1e341">MiniMPL acquires profiles of atmospheric backscattering at high frequency (2500 Hz) using a low-energy pulse (3.5 <inline-formula><mml:math id="M8" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi></mml:mrow></mml:math></inline-formula>J) emitted by an Nd:YAG laser at 532 nm. The profiles are acquired in photon-counting mode and, in our present configuration, averaged over 5 min and 30 m vertical resolution bins. The instrument uses a monostatic coaxial design where the laser beam and the receiver optics share the same axis. Because of geometrical limitations, only a fraction of the signal can be recovered in the near field. Therefore, in our system, the first usable data are available at 120 m above ground level.</p>
      <p id="d1e352">The instrument has polarization capability, collecting backscattered photons in two channels with the measured raw signals in the “copolarized” and “cross-polarized” channels suffixed “co” and “cr”, respectively (the instrument uses both circular and linear depolarization; see <xref ref-type="bibr" rid="bib1.bibx18" id="altparen.28"/>, for more details). These raw signals are processed to obtain the quantity of interest, i.e., the range-corrected signal (RCS), which is also called the normalized relative backscatter.
This processing consists of several procedures including background, overlap, afterpulse, and dead-time corrections. A comprehensive description of the processing is given in <xref ref-type="bibr" rid="bib1.bibx9" id="text.29"/>. The copolarized and cross-polarized range-corrected signals, RCS<inline-formula><mml:math id="M9" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">co</mml:mi></mml:msub></mml:math></inline-formula> and RCS<inline-formula><mml:math id="M10" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">cr</mml:mi></mml:msub></mml:math></inline-formula>, respectively, as delivered by the manufacturer's software, are used as predictors for the machine-learning algorithms described in Sect. <xref ref-type="sec" rid="Ch1.S3"/>.</p>
      <p id="d1e381">The raw data type and format depends on the instrumental device used. To make the algorithms usable on other devices, we converted the files to a harmonized format using the raw2l1 software<fn id="Ch1.Footn1"><p id="d1e384">raw2l1, which is maintained by the Site Instrumental de Recherche par Télédétection Atmosphérique and is publicly available at <uri>https://gitlab.in2p3.fr/ipsl/sirta/raw2l1</uri> (last access: 3 June 2021).</p></fn>; we then used these files as the algorithm input.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Radiosonde data</title>
      <p id="d1e400">The algorithms were evaluated with respect to estimates derived from radiosonde (RS) profiles. Météo-France operates several RS sites for the World Meteorological Organization Global Observing System. Two RS sites are colocated with the lidars at Brest and Trappes. These sites are equipped with Meteomodem robotsondes and typically launch a Meteomodem M10 sonde at 11:15 and 23:15 UTC every day.</p>
      <p id="d1e403">Many methods exist to estimate BLH from RS data, several of which have been used in the literature. Some of these methods are listed below.
<?xmltex \hack{\newpage}?>
<list list-type="bullet"><list-item>
      <p id="d1e410"><italic>Parcel method.</italic> BLH is the height at which the profile of the potential temperature <inline-formula><mml:math id="M11" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula> reaches its ground value.</p></list-item><list-item>
      <p id="d1e423"><italic>Humidity gradient method.</italic> BLH is the height at which the gradient of the relative humidity is strongly negative.</p></list-item><list-item>
      <p id="d1e429"><italic>Bulk Richardson number method.</italic> BLH is the height at which the bulk Richardson number exceeds 0.25 (this threshold varies among studies).</p></list-item><list-item>
      <p id="d1e435"><italic>Surface-based inversion.</italic> BLH is the height at which the gradient temperature profile reaches zero.</p></list-item><list-item>
      <p id="d1e441"><italic>Stable layer inversion.</italic> BLH is the height at which the gradient of the potential temperature profile reaches zero.</p></list-item></list>
<xref ref-type="bibr" rid="bib1.bibx26" id="text.30"/> used the parcel and humidity gradient methods. <xref ref-type="bibr" rid="bib1.bibx12" id="text.31"/> used all the techniques mentioned above and recommend the bulk Richardson number method for all cases.
<xref ref-type="bibr" rid="bib1.bibx21" id="text.32"/> used the bulk Richardson number for a 2-year climatology. <xref ref-type="bibr" rid="bib1.bibx48" id="text.33"/> compared the parcel, humidity gradient, and surface-based inversion methods, as well as other methods, over a period of 10 years at 505 sites worldwide. <xref ref-type="bibr" rid="bib1.bibx49" id="text.34"/> compared several methods and recommend the bulk Richardson number method.</p>
      <p id="d1e463">Following the recommendations in Fig. 10 of <xref ref-type="bibr" rid="bib1.bibx47" id="text.35"/>, we chose to compute BLH using the parcel method for the 11:15 UTC sounding and the bulk Richardson number for the 23:15 UTC sounding and refer to this estimate as BLH-RS from now on.</p>
</sec>
<?pagebreak page4338?><sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Ancillary data</title>
      <p id="d1e478">Ancillary data were used to describe the meteorological situation at the observation sites. These data were not used by the machine-learning algorithms. All the instruments were colocated with the lidar and radiosonde launches.
<list list-type="bullet"><list-item>
      <p id="d1e483">Rain gauges were used to detect rain events.</p></list-item><list-item>
      <p id="d1e487">Vaisala Ceilometer CL31 instruments were used to detect the cloud base height and distinguish cases with clouds on top of, or inside, the boundary layer. Even though MiniMPL is capable of detecting clouds, we relied on the CL31 cloud detection because the MiniMPL algorithm was found to report non-existent clouds.</p></list-item><list-item>
      <p id="d1e491">Scatterometers were used to estimate the horizontal visibility and detect the occurrence of fog.</p></list-item></list></p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Machine-learning methods</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Supervised learning method</title>
      <p id="d1e510">Supervised methods learn from a reference. Such methods are divided into two families: classification, which aims to find the frontiers between groups, and regression, which aims to approximate a function. In this study, we treat the BLH estimation as a classification problem where we wish to classify the lidar measurement at each range gate as belonging to either “boundary layer” or “free atmosphere”. Then, the highest point of the boundary layer class indicates the BLH estimate. Several supervised algorithms were compared to maximize accuracy (see Sect. <xref ref-type="sec" rid="Ch1.S3.SS1.SSS3"/>); here we describe AdaBoost, which was the algorithm selected for this study. Boosting algorithms are a very powerful family of algorithms that were developed for classification but can also be used for regression <xref ref-type="bibr" rid="bib1.bibx24" id="paren.36"/>. In particular, the AdaBoost algorithm is designed for binary classification <xref ref-type="bibr" rid="bib1.bibx19" id="paren.37"/> and is therefore well suited to our problem.</p>
<sec id="Ch1.S3.SS1.SSS1">
  <label>3.1.1</label><title>AdaBoost algorithm</title>
      <p id="d1e528">Let us consider the following problem. We have <inline-formula><mml:math id="M12" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> vectors <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mi>p</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> (here, the number of predictors, <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> – seconds since midnight, height above ground, copolarized channel, and cross-polarized channel), and for each vector, we have a binary indicator <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> for boundary layer, 1 for free atmosphere). From the sample <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mo>)</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>]</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>]</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>]</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is the ensemble of integers from 1 to <inline-formula><mml:math id="M19" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>, we want to predict the output indicator <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi mathvariant="normal">new</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> of any new vector <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="normal">new</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. To do so, we must find a rule based on the <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="normal">new</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> coordinate values (the features) to cast it into the appropriate class. Decision tree classifiers <xref ref-type="bibr" rid="bib1.bibx4" id="paren.38"/> perform this casting one feature at a time. For example, in Fig. <xref ref-type="fig" rid="Ch1.F2"/>, there are black and white points in a two-dimensional space. The black points are mostly located where <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is low, hence the rule “if <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, then the point is black”. However, in the other region, where <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>&gt;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, there are still some black points, all with low <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. Therefore, we add the rule “if <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, then the point is black; otherwise it is white”. Decision trees are classifiers made up of such “if” statements with various depths and thresholds. The deeper the tree, the more accurate the border but the more complex the decision and the longer it takes to train. Deep trees are strongly subject to overfitting and are less efficient than other methods. However, shallow decision trees are valuable because of their simplicity and their speed, even though their performance is quite limited <xref ref-type="bibr" rid="bib1.bibx24" id="paren.39"/>. They are often used as <italic>weak learners</italic>, that is, classifiers with poor performance (but better than random) that are very simple <xref ref-type="bibr" rid="bib1.bibx19" id="paren.40"/>. In this study, weak learners in AdaBoost are trees with a maximum depth of five (a maximum of five forks between the root and the leaves).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e804">Illustration of binary classification with decision trees on two-dimensional artificial data.</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f02.png"/>

          </fig>

      <p id="d1e813">AdaBoost is based on decision tree classifiers. It aggregates these classifiers to determine the most accurate border. The concept behind AdaBoost is illustrated in Fig. <xref ref-type="fig" rid="Ch1.F3"/>. First, a shallow decision tree is fitted to the entire dataset using the classification and regression tree (CART) algorithm <xref ref-type="bibr" rid="bib1.bibx24" id="paren.41"/>. All points have the same weight in this first step. Some points in the dataset are misclassified, and the error<?pagebreak page4339?> in the classifier is the weighted average of the misclassified points. Another shallow decision tree is then fitted on a resampled dataset where the previously misclassified points are over-represented. This new tree has new misclassified points that will be over-represented in the training of the next tree and so on, up to the specified number of trees (<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula> in our case). The detailed algorithm is described in <xref ref-type="bibr" rid="bib1.bibx24" id="text.42"/>, algorithm 10.1, and in <xref ref-type="bibr" rid="bib1.bibx46" id="text.43"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e842">Illustration of boosting on two-dimensional artificial data with two classes.</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f03.png"/>

          </fig>

</sec>
<sec id="Ch1.S3.SS1.SSS2">
  <label>3.1.2</label><title>Training of the algorithm</title>
      <p id="d1e859">Such an algorithm needs to be trained using a trustworthy reference.
On days where the boundary layer is easily visible to a human expert, the top of the  boundary layer can be drawn by hand; all points below this limit are in the class boundary layer, and all points above this limit are in the class free atmosphere.</p>
      <p id="d1e862"><?xmltex \hack{\newpage}?>In this study, two dates were classified by hand. These two dates were chosen because the boundary layers on these dates were easily visible; the two hand-classified dates were at different sites in different seasons. The first hand-classified date was a clear summer day in Trappes, shown in Fig. <xref ref-type="fig" rid="Ch1.F4"/> (left); a stable boundary layer is present near the ground during the night, topped by a residual layer and a few clouds between 02:00 and 04:00 UTC. A mixed layer started to develop at 09:00 UTC and remained at approximately 2000 m for the rest of the day. At approximately 22:00 UTC, a new stable layer appeared to develop near the ground; however, it is not very clear where this layer started or what its extent was. The second hand-classified date was a clear winter day in Brest, shown in Fig. <xref ref-type="fig" rid="Ch1.F4"/> (right): a stable boundary layer was present near the ground during the night, topped by a residual layer, which was shallower than the layer observed at the Trappes site. The mixed layer started to develop at 08:00 UTC and remained at approximately 1000 m with the height of the layer gradually decreasing throughout the day. At approximately 17:00 UTC, aerosols appeared to accumulate in a thin layer close to the ground; therefore, we chose to select the top of this thin layer as the BLH.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e872">Hand-drawn reference classification and radiosonde estimates overlaying the lidar range-corrected signal for two dates: 2 August 2018 at the Trappes site <bold>(a)</bold> and 24 February 2018 at the Brest site <bold>(b)</bold>.</p></caption>
            <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f04.png"/>

          </fig>

      <p id="d1e888">The coordinates of the points on the hand-drawn BLHs were obtained using the Visual Geometry Group Image Annotator software.<fn id="Ch1.Footn2"><p id="d1e891">Publicly available online at <uri>https://www.robots.ox.ac.uk/~vgg/software/via/via-1.0.6.html</uri> (last access: 3 June 2021).</p></fn> Then, the output curves were interpolated with a cubic spline to match the lidar temporal resolution. Given the resolution of the lidar, this method of labeling the data results in <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">86</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">400</mml:mn></mml:mrow></mml:math></inline-formula> individuals in total (this number is the product of 288 profiles per day, 150 vertical range bins and 2 labeled days).</p>
</sec>
<sec id="Ch1.S3.SS1.SSS3">
  <label>3.1.3</label><title>Retained configuration</title>
      <p id="d1e921">Four predictors were used: the two lidar channels, time (number of seconds since midnight), and altitude (meters above ground level). The ADABL configuration used was
<list list-type="bullet"><list-item>
      <p id="d1e926">weak learner – decision tree of depth five;</p></list-item><list-item>
      <p id="d1e930">number of weak learners – 200; and</p></list-item><list-item>
      <p id="d1e934">predictors – time, altitude, RCS<inline-formula><mml:math id="M30" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">co</mml:mi></mml:msub></mml:math></inline-formula>, and RCS<inline-formula><mml:math id="M31" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">cr</mml:mi></mml:msub></mml:math></inline-formula>.</p></list-item></list></p>
      <?pagebreak page4340?><p id="d1e955">This configuration was chosen because more complex classifiers do not necessarily improve the performance. The computation time of the algorithm was still reasonable: training took 23 s on the full dataset and predicting BLH for a full day took 3.7 s with a modern laptop. AdaBoost was chosen after a comparison of multiple classification algorithms, i.e., random forest, nearest neighbor, decision trees, and label spreading (study not shown here). The benchmark score was the accuracy as measured by the percentage of individuals that were correctly classified. The accuracy was estimated by group <inline-formula><mml:math id="M32" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-fold cross-validation, where labeled datasets are grouped into chunks of 3 consecutive hours; one group was used as a testing set and all the rest as a training set. This operation was repeated until each group was used as the testing set. The resulting accuracy was 96 %. However, this figure overestimates the generalization ability of AdaBoost. A more correct estimation would be obtained with an independent validation set (e.g., a new hand-classified day). An independent validation set was not used here because the cross-validation accuracy was only used to discriminate between the classification algorithms.</p>
      <p id="d1e965">It is possible to quantify the relative importance of the predictors <xref ref-type="bibr" rid="bib1.bibx4 bib1.bibx24" id="paren.44"/>. After training, the relative importance of the time, RCS<inline-formula><mml:math id="M33" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">co</mml:mi></mml:msub></mml:math></inline-formula>, RCS<inline-formula><mml:math id="M34" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">cr</mml:mi></mml:msub></mml:math></inline-formula>, and altitude predictors was 30.3 %, 28.4 %, 26.5 %, and 14.8 %, respectively.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Unsupervised learning methods</title>
      <p id="d1e998">Unsupervised methods aim to identify groups in the data. In our case, we want to identify the group boundary layer. The BLH estimate is then the upper boundary of this group. Two unsupervised learning algorithms were tested: <inline-formula><mml:math id="M35" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means and expectation–maximization (EM).</p>
<sec id="Ch1.S3.SS2.SSS1">
  <label>3.2.1</label><?xmltex \opttitle{$K$-means algorithm}?><title><inline-formula><mml:math id="M36" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithm</title>
      <p id="d1e1022">The <inline-formula><mml:math id="M37" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithm is a well proven and commonly used algorithm for data segmentation <xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx40" id="paren.45"/> and consists of three steps, where <inline-formula><mml:math id="M38" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> is the number of clusters specified by the user.
<list list-type="order"><list-item>
      <p id="d1e1044"><italic>Initialization. </italic><inline-formula><mml:math id="M39" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> centroids <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, …, <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi>K</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are initialized at random locations inside the feature space.</p></list-item><list-item>
      <p id="d1e1078"><italic>Attribution.</italic> The distances from all points to all centroids <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:msub><mml:mo>)</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>]</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>]</mml:mo><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mspace linebreak="nobreak" width="-0.125em"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>N</mml:mi><mml:mo>]</mml:mo><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>]</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are computed, and points are attributed to the closest centroid: <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mi>C</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">argmin</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo mathvariant="italic">{</mml:mo><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>.
<?xmltex \hack{\newpage}?></p></list-item><list-item>
      <p id="d1e1203"><italic>Update.</italic> The centroids are re-defined as the average point of the cluster: <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">m</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mn mathvariant="bold">1</mml:mn><mml:mrow><mml:mi>C</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:msub><mml:mn mathvariant="bold">1</mml:mn><mml:mrow><mml:mi>C</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula>.</p></list-item></list>
Steps 2 and 3 are repeated until the centroids stop moving. It has been shown that this algorithm converges to a local minimum of the intra-cluster variance <xref ref-type="bibr" rid="bib1.bibx51" id="paren.46"/>. Figure <xref ref-type="fig" rid="Ch1.F5"/> (left) illustrates this algorithm.</p>
</sec>
<sec id="Ch1.S3.SS2.SSS2">
  <label>3.2.2</label><title>EM algorithm</title>
      <p id="d1e1301">The EM algorithm assumes that each group <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>∈</mml:mo><mml:mo>[</mml:mo><mml:mspace linebreak="nobreak" width="-0.125em"/><mml:mo>[</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>]</mml:mo><mml:mspace linebreak="nobreak" width="-0.125em"/><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is generated by a Gaussian distribution (<inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>). The algorithm iteratively estimates the parameters <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="bold">Σ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and the <italic>responsibility</italic> for each Gaussian <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="italic">γ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>, where the <italic>responsibility</italic> is the probability of the point <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> being generated by the <inline-formula><mml:math id="M52" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>th Gaussian. Points are then attributed to the group with the highest responsibility: <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mi>C</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">argmax</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="italic">γ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mn mathvariant="normal">1</mml:mn><mml:mi>i</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msubsup><mml:mover accent="true"><mml:mi mathvariant="italic">γ</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>K</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
Figure <xref ref-type="fig" rid="Ch1.F5"/> (right) illustrates this algorithm.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e1480">Illustration of the <inline-formula><mml:math id="M54" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means and expectation–maximization algorithm on two-dimensional artificial data with two clusters.</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f05.png"/>

          </fig>

      <p id="d1e1496">The <inline-formula><mml:math id="M55" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means and EM algorithms are very similar. If we assume that all Gaussian distributions have the same fixed variance and that this variance tends to zero, the EM and <inline-formula><mml:math id="M56" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithms are the same. However, <inline-formula><mml:math id="M57" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means does not rely on a Gaussian assumption.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e1523">Simplified flowchart of the <inline-formula><mml:math id="M58" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means for Atmospheric Boundary Layer (KABL) and AdaBoost for Atmospheric Boundary Layer (ADABL) algorithms with a focus on the KABL parameters. The parameters are described in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/> and in Table <xref ref-type="table" rid="Ch1.T2"/>.</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f06.png"/>

          </fig>

</sec>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Flowchart and description of KABL parameters</title>
      <p id="d1e1552">A simplified flowchart of KABL and ADABL is shown in Fig. <xref ref-type="fig" rid="Ch1.F6"/>. This section focuses on the KABL parameters to introduce the sensitivity analysis made in Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>. The parameters of the KABL software are detailed here.
<list list-type="bullet"><list-item>
      <p id="d1e1561"><monospace>algo</monospace> is the applied machine-learning algorithm; possible values are
<list list-type="bullet"><list-item>
      <p id="d1e1568">“gmm” for the EM algorithm (Gaussian mixture model) and</p></list-item><list-item>
      <p id="d1e1572">“kmeans” for the <inline-formula><mml:math id="M59" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithm.</p></list-item></list></p></list-item><list-item>
      <?pagebreak page4341?><p id="d1e1583"><monospace>classif_score</monospace> is the internal score used to automatically choose the number of clusters (only used when <monospace>n_clusters</monospace> is “auto”). See Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/> and Table <xref ref-type="table" rid="Ch1.T1"/> for a description of the internal scores.</p></list-item><list-item>
      <p id="d1e1596"><monospace>init</monospace> is the initialization strategy for both algorithms. Three choices are available:
<?xmltex \hack{\newpage}?>
<list list-type="bullet"><list-item>
      <p id="d1e1605">“random” – randomly pick an individual as the starting point (both <inline-formula><mml:math id="M60" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means and EM),</p></list-item><list-item>
      <p id="d1e1616">“advanced” – use a more sophisticated initialization (kmeans<inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mo>+</mml:mo></mml:mrow></mml:math></inline-formula> for <inline-formula><mml:math id="M62" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means – <xref ref-type="bibr" rid="bib1.bibx2" id="altparen.47"/> – and the output a <inline-formula><mml:math id="M63" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means pass for EM), and</p></list-item><list-item>
      <p id="d1e1647">“given” – start at explicitly selected point coordinates.</p></list-item></list></p></list-item><list-item>
      <p id="d1e1651"><monospace>max_height</monospace> is the height (meters above ground level) at which the profiles are cut.</p></list-item><list-item>
      <p id="d1e1657"><monospace>n_clusters</monospace> is the number of clusters to be formed (between two and six). This is either explicitly given or determined automatically to optimize the score given in <monospace>classif_score</monospace>.</p></list-item><list-item>
      <p id="d1e1666"><monospace>n_inits</monospace> is the number of repetitions of the algorithm. When this number is larger, the algorithm is more likely to find the global optimum but requires more time.</p></list-item><list-item>
      <p id="d1e1672"><monospace>n_profiles</monospace> is the number of profiles concatenated prior to the application of the algorithm. For example, if <monospace>n_profiles</monospace> <inline-formula><mml:math id="M64" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 1, only the current profile is used. If <monospace>n_profiles</monospace> <inline-formula><mml:math id="M65" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 3, the current profile and the two previous profiles are concatenated and input into the algorithm.</p></list-item><list-item>
      <p id="d1e1698"><monospace>predictors</monospace> is the list of variables used in the classification. These variables can be different at night and during the day. For both time periods, the variables can be chosen from
<list list-type="bullet"><list-item>
      <p id="d1e1705">RCS<inline-formula><mml:math id="M66" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">co</mml:mi></mml:msub></mml:math></inline-formula>, the copolarized range-corrected backscatter signal, and</p></list-item><list-item>
      <p id="d1e1718">RCS<inline-formula><mml:math id="M67" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">cr</mml:mi></mml:msub></mml:math></inline-formula>, the cross-polarized range-corrected backscatter signal.</p></list-item></list></p></list-item></list></p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e1733">Table of metrics used to measure the performance of the <inline-formula><mml:math id="M68" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means for Atmospheric Boundary Layer (KABL) algorithm. The metrics are described in detail in Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Metric</oasis:entry>
         <oasis:entry colname="col2">Type</oasis:entry>
         <oasis:entry colname="col3">Description</oasis:entry>
         <oasis:entry colname="col4">Best, worst</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">score</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">corr</oasis:entry>
         <oasis:entry colname="col2">External</oasis:entry>
         <oasis:entry colname="col3">Pearson correlation coefficient</oasis:entry>
         <oasis:entry colname="col4">1, 0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">errl2</oasis:entry>
         <oasis:entry colname="col2">External</oasis:entry>
         <oasis:entry colname="col3">Root-mean-square error (RMSE)</oasis:entry>
         <oasis:entry colname="col4">0, <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">s_score</oasis:entry>
         <oasis:entry colname="col2">Internal</oasis:entry>
         <oasis:entry colname="col3">Silhouette score</oasis:entry>
         <oasis:entry colname="col4">1, <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">db_score</oasis:entry>
         <oasis:entry colname="col2">Internal</oasis:entry>
         <oasis:entry colname="col3">Davies–Bouldin index</oasis:entry>
         <oasis:entry colname="col4">0, <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ch_score</oasis:entry>
         <oasis:entry colname="col2">Internal</oasis:entry>
         <oasis:entry colname="col3">Calinski–Harabasz index</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula>, 0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">chrono</oasis:entry>
         <oasis:entry colname="col2">Other</oasis:entry>
         <oasis:entry colname="col3">Time to estimate BLH for a 24 h period</oasis:entry>
         <oasis:entry colname="col4">0, <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">n_invalid</oasis:entry>
         <oasis:entry colname="col2">Other</oasis:entry>
         <oasis:entry colname="col3">Number of invalid BLH estimates (NaN or Inf) for a 24 h period</oasis:entry>
         <oasis:entry colname="col4">0, 288</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e1943">The parameters of the KABL software are written in typewriter font in the following explanation of the KABL algorithm. A netCDF file generated by the raw2l1 software needs to be provided as input data to KABL. The data, namely, the altitude vector <inline-formula><mml:math id="M74" display="inline"><mml:mi mathvariant="bold-italic">z</mml:mi></mml:math></inline-formula> (size <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), the time vector <inline-formula><mml:math id="M76" display="inline"><mml:mi mathvariant="bold-italic">t</mml:mi></mml:math></inline-formula> (size <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and the range-corrected signals <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">RCS</mml:mi><mml:mi mathvariant="normal">co</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold">RCS</mml:mi><mml:mi mathvariant="normal">cr</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>×</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> matrices), are extracted from this file. Such data are prepared to fulfill the machine-learning algorithm requirements. For each time, the <monospace>n_profiles</monospace> last profiles are extracted. Then, the data they contain are normalized (by removing the mean and dividing by the standard deviation); this provides a matrix <inline-formula><mml:math id="M81" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> (<inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>×</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M83" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is <monospace>n_profiles</monospace> <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mo>⋅</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">|</mml:mi></mml:mrow></mml:math></inline-formula><monospace>predictors</monospace><inline-formula><mml:math id="M86" display="inline"><mml:mi mathvariant="normal">|</mml:mi></mml:math></inline-formula> is the number of elements in the list). The matrix <inline-formula><mml:math id="M87" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula> is the usual input for a machine-learning algorithm; it has one line for each individual observation and one column for each variable (or predictor) observed. For the BLH retrieval, the preparation also provides a vector <inline-formula><mml:math id="M88" display="inline"><mml:mi mathvariant="bold-italic">Z</mml:mi></mml:math></inline-formula> (size <inline-formula><mml:math id="M89" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>) containing the altitude of each individual observation. The algorithm (either <inline-formula><mml:math id="M90" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means or EM, as specified by <monospace>algo</monospace>) is applied to the matrix <inline-formula><mml:math id="M91" display="inline"><mml:mi mathvariant="bold">X</mml:mi></mml:math></inline-formula>, with the parameters <monospace>n_clusters</monospace>, <monospace>init</monospace>, and <monospace>n_inits</monospace>. This results in a vector of <italic>labels</italic> (size <inline-formula><mml:math id="M92" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>) that contains the cluster attribution of each individual. Finally, since by definition the boundary layer is the layer directly influenced by the ground, we look for the first change in the cluster attribution, starting from the ground level. This gives us the value of BLH for this profile. These operations are repeated until reaching the end of the netCDF file.</p>
</sec>
<?pagebreak page4342?><sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Performance metrics</title>
      <p id="d1e2156">Two types of metrics were used.
<list list-type="bullet"><list-item>
      <p id="d1e2161"><italic>External scores.</italic> These metrics compare the result to a trustworthy reference. They have the advantage of providing a meaningful evaluation of the performance but depend strongly on the quality of the reference (i.e., its accuracy and availability).</p></list-item><list-item>
      <p id="d1e2167"><italic>Internal scores.</italic> These metrics rate how well the classification performs based only on the distances between points. They have the advantage of being always computable but are not linked to any physical property and therefore are not always meaningful.</p></list-item></list>
None of these metrics are perfect; however, the information they provide allows a broader understanding of the algorithm performance.</p>
<sec id="Ch1.S3.SS4.SSS1">
  <label>3.4.1</label><title>External scores</title>
      <p id="d1e2180">External scores use a reference to assess the quality of the result. In our case, the references are BLH-RS (RS-derived BLH) and, when available, the human expert hand-classified BLH. Two external scores are used in this study. If we denote <inline-formula><mml:math id="M93" display="inline"><mml:mover accent="true"><mml:mi>Z</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> as the estimated BLH (by any of the previously introduced algorithms) and <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> as the reference, the external scores are as follows:
<list list-type="bullet"><list-item>
      <p id="d1e2206">RMSE, where lower values are better,<disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M95" display="block"><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi><mml:mfenced open="[" close="]"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mover accent="true"><mml:mi>Z</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:msqrt><mml:mo>;</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item><list-item>
      <p id="d1e2247">the Pearson correlation, where higher values are better,<disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M96" display="block"><mml:mrow><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi mathvariant="normal">cov</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mover accent="true"><mml:mi>Z</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mi>Z</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mi mathvariant="italic">σ</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p></list-item></list></p>
      <p id="d1e2301">Here, <inline-formula><mml:math id="M97" display="inline"><mml:mover accent="true"><mml:mi>Z</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are random variables, <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi><mml:mo>[</mml:mo><mml:mo>⋅</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> denotes the mathematical expectation, and <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the standard deviation. When these scores are estimated, the random variables are replaced by a sample vector and the expectation and standard deviation are replaced by their usual estimators.</p>
</sec>
<?pagebreak page4343?><sec id="Ch1.S3.SS4.SSS2">
  <label>3.4.2</label><title>Internal scores</title>
      <p id="d1e2361">The quality of a classification can be quantified using scores that are based only on the labels and the distances between points. Such scores estimate how trustworthy an estimation is without any external input. Many such scores exist with different formulations and different strengths and weaknesses <xref ref-type="bibr" rid="bib1.bibx16" id="paren.48"/>. In this study, three internal scores were used:
<list list-type="bullet"><list-item>
      <p id="d1e2369">the silhouette score <xref ref-type="bibr" rid="bib1.bibx45" id="paren.49"/>,<disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M101" display="block"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">sil</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>b</mml:mi><mml:mo>-</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math id="M102" display="inline"><mml:mi>a</mml:mi></mml:math></inline-formula> is the average distance to its own group and <inline-formula><mml:math id="M103" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> is the average distance to the neighboring group; <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">sil</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> is the best classification; <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">sil</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> is neutral, and <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">sil</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> is the worst classification;</p></list-item><list-item>
      <p id="d1e2475">the Calinski–Harabasz index <xref ref-type="bibr" rid="bib1.bibx8" id="paren.50"/>,<disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M107" display="block"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">ch</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>-</mml:mo><mml:mi>K</mml:mi><mml:mo>)</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>K</mml:mi></mml:munderover><mml:msub><mml:mi>W</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math id="M108" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of points, <inline-formula><mml:math id="M109" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> is the number of clusters, <inline-formula><mml:math id="M110" display="inline"><mml:mi>B</mml:mi></mml:math></inline-formula> is the between-cluster dispersion, and <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the within-cluster dispersion of the cluster <inline-formula><mml:math id="M112" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>; higher <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">ch</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> indicates a better classification;</p></list-item><list-item>
      <p id="d1e2591">the Davies–Bouldin index <xref ref-type="bibr" rid="bib1.bibx13" id="paren.51"/>,<disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M114" display="block"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">db</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">max</mml:mi><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>≠</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">δ</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">δ</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula><mml:math id="M115" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> are the two cluster numbers, <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi mathvariant="italic">δ</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the average distance between points and their cluster center for the cluster <inline-formula><mml:math id="M118" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the distance between the cluster centers <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mrow><mml:msup><mml:mi>k</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>; lower <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">db</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> indicates a better classification.</p></list-item></list>
These three scores were chosen to diversify the metrics and are all implemented in scikit-learn (version <inline-formula><mml:math id="M123" display="inline"><mml:mo>≥</mml:mo></mml:math></inline-formula>0.20).</p>
</sec>
<sec id="Ch1.S3.SS4.SSS3">
  <label>3.4.3</label><title>Other metrics</title>
      <p id="d1e2793">In addition to the internal and external scores, the computation time and the number of invalid values (NaN or Inf) were recorded. BLH estimates of NaN or Inf can occur when all the points of the profile are assigned to the same cluster; this reflects a faulty configuration of the algorithm. Even though these metrics do not measure how well a program is performing, they are useful to the user.</p>
      <p id="d1e2796">All the metrics used to measure the performance of KABL are summarized in Table <xref ref-type="table" rid="Ch1.T1"/>.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Sensitivity analysis of the KABL algorithm</title>
      <p id="d1e2819">A sensitivity analysis was performed on the KABL code to identify the “best” configuration. Various KABL configurations were extensively tested on a single day, 2 August 2018, at the Trappes site, for which we have a hand-classified reference (Fig. <xref ref-type="fig" rid="Ch1.F4"/>, left). The most relevant configurations were retained and tested on the 2-year lidar dataset.</p>
      <p id="d1e2824">There are eight parameters in the KABL code (see Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/> for their descriptions). To assess the sensitivity of KABL to these parameters, the performance metrics (given in Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>) were calculated using the hand-classified BLH as <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and with the output of KABL as <inline-formula><mml:math id="M125" display="inline"><mml:mover accent="true"><mml:mi>Z</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> for different combinations of input parameters. The output metrics given in Table <xref ref-type="table" rid="Ch1.T1"/> were tested using the input values given in Table <xref ref-type="table" rid="Ch1.T2"/>. We refer to a set of values for the KABL parameters as a <italic>configuration</italic>. Screening all the possible values listed in Table <xref ref-type="table" rid="Ch1.T2"/> would require 3240 different configurations.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e2865">Possible values for the parameters of the KABL code. The parameters are described in details in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>. The dependencies between parameters result in 3240 different configurations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">

         <oasis:entry colname="col1">Parameter</oasis:entry>

         <oasis:entry colname="col2">Possible values</oasis:entry>

         <oasis:entry colname="col3">Meaning</oasis:entry>

       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="1"><monospace>algo</monospace></oasis:entry>

         <oasis:entry colname="col2">kmeans</oasis:entry>

         <oasis:entry colname="col3">The <inline-formula><mml:math id="M126" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithm is used</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">gmm</oasis:entry>

         <oasis:entry colname="col3">The EM algorithm is used (Gaussian mixture model)</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="2"><monospace>classif_score</monospace></oasis:entry>

         <oasis:entry colname="col2">silh</oasis:entry>

         <oasis:entry colname="col3">The silhouette score is used</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">db</oasis:entry>

         <oasis:entry colname="col3">The Davies–Bouldin index is used</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">ch</oasis:entry>

         <oasis:entry colname="col3">The Calinski–Harabasz index is used</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="2"><monospace>init</monospace></oasis:entry>

         <oasis:entry colname="col2">random</oasis:entry>

         <oasis:entry colname="col3">Starting points are chosen randomly</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">advanced</oasis:entry>

         <oasis:entry colname="col3">Starting points are chosen with a smarter strategy</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">given</oasis:entry>

         <oasis:entry colname="col3">Starting points are explicitly given</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="1"><monospace>max_height</monospace></oasis:entry>

         <oasis:entry colname="col2">3500</oasis:entry>

         <oasis:entry colname="col3">The altitude above which profile data are discarded</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">4500</oasis:entry>

         <oasis:entry colname="col3">(meters above ground level)</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col1" morerows="4"><monospace>n_clusters</monospace></oasis:entry>

         <oasis:entry colname="col2">2</oasis:entry>

         <oasis:entry colname="col3"/>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">3</oasis:entry>

         <oasis:entry colname="col3">The number of clusters to be formed is explicitly</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">4</oasis:entry>

         <oasis:entry colname="col3">passed and is always the same</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">5</oasis:entry>

         <oasis:entry colname="col3"/>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">auto</oasis:entry>

         <oasis:entry colname="col3">The number of clusters is automatically chosen to</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col1"/>

         <oasis:entry colname="col2"/>

         <oasis:entry colname="col3">optimize <monospace>classif_score</monospace></oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="1"><monospace>n_inits</monospace></oasis:entry>

         <oasis:entry colname="col2">10</oasis:entry>

         <oasis:entry colname="col3">The number of times the algorithm is repeated with</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">80</oasis:entry>

         <oasis:entry colname="col3">different initializations (when <monospace>init</monospace> is not given)</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="3"><monospace>n_profiles</monospace></oasis:entry>

         <oasis:entry colname="col2">1</oasis:entry>

         <oasis:entry colname="col3">Only the current profile is used</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">2</oasis:entry>

         <oasis:entry colname="col3">The current profile and the previous profile are used</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">3</oasis:entry>

         <oasis:entry colname="col3">The current profile and the two previous profiles are used</oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">4</oasis:entry>

         <oasis:entry colname="col3">The current profile and the three previous profiles are used</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col1" morerows="5"><monospace>predictors</monospace></oasis:entry>

         <oasis:entry colname="col2">co</oasis:entry>

         <oasis:entry colname="col3">The copolarized range-corrected signal is used at all times</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">co/co <inline-formula><mml:math id="M127" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> cr</oasis:entry>

         <oasis:entry colname="col3">The copolarized range-corrected signal is used during</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2"/>

         <oasis:entry colname="col3">the daytime, and both polarization channels are used</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2"/>

         <oasis:entry colname="col3">independently during the nighttime</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2">co <inline-formula><mml:math id="M128" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> cr</oasis:entry>

         <oasis:entry colname="col3">Both polarization channels are used independently at</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col2"/>

         <oasis:entry colname="col3">all times</oasis:entry>

       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <?pagebreak page4344?><p id="d1e3210">To obtain an overview of these 3240 configurations, we started by estimating the influence of the parameters (listed in Table <xref ref-type="table" rid="Ch1.T2"/>) on the different metrics (listed in Table <xref ref-type="table" rid="Ch1.T1"/>). The influence of the parameters was quantified using first-order Sobol indices <xref ref-type="bibr" rid="bib1.bibx53 bib1.bibx29 bib1.bibx42" id="paren.52"/>, that is, the ratio of the variance of the metric when the parameter was fixed over the total variance of the metric. If we denote <inline-formula><mml:math id="M129" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula> as the metric and <inline-formula><mml:math id="M130" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula> as the vector of the parameters, where all the parameters are treated as random variables, the first-order Sobol index of the <inline-formula><mml:math id="M131" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th parameter is defined as
<inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>V</mml:mi><mml:mspace linebreak="nobreak" width="-0.125em"/><mml:mo>(</mml:mo><mml:mi mathvariant="double-struck">E</mml:mi><mml:mo>[</mml:mo><mml:mi>Y</mml:mi><mml:mi mathvariant="normal">|</mml:mi><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>]</mml:mo><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi>V</mml:mi><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>(</mml:mo><mml:mi>Y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:mi>V</mml:mi><mml:mspace width="-0.125em" linebreak="nobreak"/><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the variance and <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi><mml:mo>[</mml:mo><mml:mo>⋅</mml:mo><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> denotes the expectation. A higher Sobol index indicates a larger influence.</p>
      <p id="d1e3318">Figure <xref ref-type="fig" rid="Ch1.F7"/> shows the Sobol indices obtained with the KABL computer code. Examining the matrix line by line, one can see that the different metrics are sensitive to different parameters. For example, the silhouette score is very sensitive to <monospace>n_clusters</monospace> while the Calinski–Harabasz index is sensitive to <monospace>n_profiles</monospace> and <monospace>predictors</monospace>. Examining the matrix column by column, one can see that some parameters are more influential than others (e.g., <monospace>classif_score</monospace> is much less influential than <monospace>n_clusters</monospace>). This matrix highlights the main effects of changing a parameter and, therefore, how to set each parameter appropriately. For each parameter, we examined the metrics that it influences and determined the preferred configuration.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e3341">Relative influence of parameters on the different metrics. The <inline-formula><mml:math id="M135" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis indicates the parameters of the code, and the <inline-formula><mml:math id="M136" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis indicates the metrics. The shading represents the influence of the parameter on the metric with darker shading indicating a larger influence. The abbreviations for the parameters are described in Table <xref ref-type="table" rid="Ch1.T2"/>, and the abbreviations for the metrics are described in Table <xref ref-type="table" rid="Ch1.T1"/>.</p></caption>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f07.png"/>

        </fig>

      <?xmltex \floatpos{p}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e3370">Distribution of the relevant outputs for the critical inputs. The effect of <monospace>algo</monospace> on <bold>(a)</bold> the computing time and <bold>(b)</bold> the Davies–Bouldin index. The effect of <monospace>init</monospace> on <bold>(c)</bold> the correlation and <bold>(d)</bold> the computing time. The effect of <monospace>n_clusters</monospace> on <bold>(e)</bold> the root-mean-square error (RMSE) and <bold>(f)</bold> the silhouette score. The effect of the predictors on <bold>(g)</bold> the silhouette score and <bold>(h)</bold> the Calinski–Harabasz index. For each panel, the best parameter value is highlighted by a yellow star.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f08.png"/>

        </fig>

      <p id="d1e3413">Critical parameters are indicated in Fig. <xref ref-type="fig" rid="Ch1.F7"/> by the darkest blue columns, namely, <monospace>n_clusters</monospace>, <monospace>algo</monospace>, <monospace>predictors</monospace>, and <monospace>init</monospace>.<fn id="Ch1.Footn3"><p id="d1e3431">Even though <monospace>n_profiles</monospace> has a large Sobol index for the Calinski–Harabasz index, this influence was not explored because it is known: it is due to the linear increase in this index with the number of points.</p></fn> For each parameter, Fig. <xref ref-type="fig" rid="Ch1.F8"/> shows the distribution of the relevant output given the parameter value (violin plots are explained in <xref ref-type="bibr" rid="bib1.bibx28" id="altparen.53"/>). For example, Fig. <xref ref-type="fig" rid="Ch1.F8"/>a has the value of <monospace>algo</monospace> on its <inline-formula><mml:math id="M137" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis and the computing time on its <inline-formula><mml:math id="M138" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis. The 3240 different configurations were divided into two groups; those with <monospace>algo</monospace> <inline-formula><mml:math id="M139" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> kmeans and those with <monospace>algo</monospace> <inline-formula><mml:math id="M140" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> gmm. Figure <xref ref-type="fig" rid="Ch1.F8"/>a shows a smoothed histogram of the computing time for the divided populations. The other panels in Fig. <xref ref-type="fig" rid="Ch1.F8"/> were constructed in the same manner. Each line corresponds to a critical parameter, and we represent the two most influenced outputs according to Fig. <xref ref-type="fig" rid="Ch1.F7"/>.</p>
      <?pagebreak page4345?><p id="d1e3491">The parameter values were chosen to give the optimal values for the metrics they influence. The optimal values are indicated by a yellow star in each plot. To set <monospace>algo</monospace>, we examined the computing time (Fig. <xref ref-type="fig" rid="Ch1.F8"/>a) and the Davies–Bouldin index (Fig. <xref ref-type="fig" rid="Ch1.F8"/>b). These figures indicate that kmeans is the best choice for both metrics (resulting in a lower computing time and a lower Davies–Bouldin index). To set <monospace>init</monospace>, we examined the correlation (Fig. <xref ref-type="fig" rid="Ch1.F8"/>c) and the computing time (Fig. <xref ref-type="fig" rid="Ch1.F8"/>d). In this case, given appears to be the best choice. To set <monospace>n_clusters</monospace>, we examined RMSE (Fig. <xref ref-type="fig" rid="Ch1.F8"/>e) and the silhouette score (Fig. <xref ref-type="fig" rid="Ch1.F8"/>f). They indicate that the best numbers of clusters are three and auto, respectively. We chose to give priority to RMSE because the silhouette score has very high values for two clusters, which is suspicious given the presence of a cloud and a residual layer on this day. To set <monospace>predictors</monospace>, we examined the silhouette score (Fig. <xref ref-type="fig" rid="Ch1.F8"/>g) and the Calinski–Harabasz index (Fig. <xref ref-type="fig" rid="Ch1.F8"/>h); here, co appears to be the best choice. Following this methodology, we can identify a few configurations worth trying. These configurations were tested on the 2-year dataset. The configuration used to generate the results in Sect. <xref ref-type="sec" rid="Ch1.S4.SS2.SSS1"/> is given in Table <xref ref-type="table" rid="Ch1.T3"/>. This configuration was chosen to maximize the correlation between KABL and RS at the Trappes site.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e3531">Retained values for the parameters of the KABL code after the sensitivity analysis.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Parameter</oasis:entry>
         <oasis:entry colname="col2">Retained values</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>algo</monospace></oasis:entry>
         <oasis:entry colname="col2">kmeans</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>classif_score</monospace></oasis:entry>
         <oasis:entry colname="col2">db</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>init</monospace></oasis:entry>
         <oasis:entry colname="col2">given</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>max_height</monospace></oasis:entry>
         <oasis:entry colname="col2">4500</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>n_clusters</monospace></oasis:entry>
         <oasis:entry colname="col2">3</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>n_inits</monospace></oasis:entry>
         <oasis:entry colname="col2">10</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>n_profiles</monospace></oasis:entry>
         <oasis:entry colname="col2">1</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>predictors</monospace></oasis:entry>
         <oasis:entry colname="col2">co</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>A 2-year comparison</title>
      <p id="d1e3647">BLH estimates from the three methods (KABL, ADABL, and the manufacturer's algorithm) were compared to BLH-RS over a 2-year period.</p><?xmltex \hack{\newpage}?>
<sec id="Ch1.S4.SS2.SSS1">
  <label>4.2.1</label><title>Overall comparison</title>
      <p id="d1e3658">As explained in Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>, two external scores, RMSE and the correlation, were used to assess the quality of the estimates.
In Eqs. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) and (<xref ref-type="disp-formula" rid="Ch1.E2"/>), the reference BLH <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msub><mml:mi>Z</mml:mi><mml:mi mathvariant="normal">ref</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> was set to the BLH-RS, as described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>. To compute the scores, the BLH estimates from the lidar and radiosonde must be colocated. For each BLH-RS, the corresponding lidar BLH estimate is the average of all available estimates within the 10 min following the release of the radiosonde (this translates to one or two lidar estimates). Using the ancillary measurements presented in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>, the following meteorological conditions were discarded:
<list list-type="bullet"><list-item>
      <p id="d1e3685">rain (rain gauge measures the rainfall as <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> mm),</p></list-item><list-item>
      <p id="d1e3699">fog (scatterometer measures the visibility as <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1000</mml:mn></mml:mrow></mml:math></inline-formula> m),</p></list-item><list-item>
      <p id="d1e3713">low-level cloud (ceilometer measures the cloud base height as <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">3000</mml:mn></mml:mrow></mml:math></inline-formula> m),</p></list-item><list-item>
      <p id="d1e3727">BLH-RS below 120 m (blind zone for lidar), and</p></list-item><list-item>
      <p id="d1e3731">nighttime (RS launched at 23:15 UTC).</p></list-item></list>
This selection rejects a large part of the dataset but ensures that only well-defined cases are retained for the comparison. In total, 178 RS measurements from Trappes and 101 RS measurements from Brest were used for the overall comparison.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F9"><?xmltex \currentcnt{9}?><?xmltex \def\figurename{Figure}?><label>Figure 9</label><caption><p id="d1e3737">Results of a 2-year comparison with the radiosonde (RS) estimates at both sites for two metrics: RMSE and correlation. INDUS refers to the manufacturer's algorithm; KABL and ADABL refer to the eponymous estimates. Cases at night or with rain, fog, an RS-estimated boundary layer height (BLH) of under 120 m, or clouds under 3000 m were removed. The 95 % confidence intervals were estimated using percentile bootstrapping <xref ref-type="bibr" rid="bib1.bibx14" id="paren.54"/>.</p></caption>
            <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f09.png"/>

          </fig>

      <?pagebreak page4347?><p id="d1e3749">Figure <xref ref-type="fig" rid="Ch1.F9"/> displays how the three methods, KABL (blue bars), ADABL (brown bars) and manufacturer (orange bars), compare to BLH-RS.
The first column represents RMSE <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> (lower is better), and the second column represents the correlation <inline-formula><mml:math id="M146" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> (higher is better). The upper row shows the results for the Brest site; the lower row shows the results for the Trappes site. While both KABL and ADABL outperform the manufacturer's algorithm at the Trappes site, neither algorithm does at the Brest site. While the correlation for both KABL and ADABL is higher than that for the manufacturer's algorithm at the Trappes site, it collapses to close to zero for KABL at the Brest site (0.07 for ADABL). The RMSE values can be compared to the values given in <xref ref-type="bibr" rid="bib1.bibx23" id="text.55"/>. For KABL, we find 770 m at the Brest site and 798 m at the Trappes site, while for ADABL, we find 675 m at the Brest site and 552 m at the Trappes site. Our values are notably higher than those in <xref ref-type="bibr" rid="bib1.bibx23" id="text.56"/>. This is likely due to the larger extent of our dataset (178 RS at Trappes and 101 at Brest, spanning a 2-year period) and the low maturity of the algorithms. ADABL has better correlation and RMSE values than KABL at both sites. The manufacturer's algorithm performs well without any specific tuning on our part. It uses a wavelet covariance transform, as described in <xref ref-type="bibr" rid="bib1.bibx6" id="text.57"/>. This result is not surprising because the wavelet method has been shown to be robust in numerous studies, especially in <xref ref-type="bibr" rid="bib1.bibx7" id="text.58"/>, who included a cluster analysis method and concluded that the wavelet method is preferred.</p>
</sec>
<sec id="Ch1.S4.SS2.SSS2">
  <label>4.2.2</label><title>Seasonal and diurnal cycles</title>
      <p id="d1e3793">To quantify the ability of the algorithms to provide a consistent BLH estimate, Fig. <xref ref-type="fig" rid="Ch1.F10"/> shows the seasonal cycle (monthly average) and the diurnal cycle (6 min average) at both sites. For each estimator, the thick line represents the average BLH estimate and the shaded area represents the inter-quartile gap. Rain, fog, and low-cloud conditions were discarded. For the monthly average, the nighttime values were also removed. The seasonal cycle is reversed when only nighttime values are studied, with BLH-RS being lower in summer than in winter, on average. For other estimators, we do not see such a difference between the day and night seasonal cycles (not shown).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F10" specific-use="star"><?xmltex \currentcnt{10}?><?xmltex \def\figurename{Figure}?><label>Figure 10</label><caption><p id="d1e3800"><bold>(a, c)</bold> Seasonal and <bold>(b, d)</bold> diurnal cycles of all BLH estimates at both sites. INDUS indicates the manufacturer's algorithm. Thick lines represent the average, and the shaded area represents the quartiles.</p></caption>
            <?xmltex \igopts{width=455.244094pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f10.png"/>

          </fig>

      <p id="d1e3814">At the Brest site (Fig. <xref ref-type="fig" rid="Ch1.F10"/>a), estimates made by the manufacturer's algorithm are lower than those made by KABL and ADABL and estimates made by ADABL are usually higher than those made by KABL (except in July). BLH-RS values were low in summer (June–October), high in February and March (higher than the KABL estimates), and between the manufacturer's and KABL estimates during the rest of the year. Overall, the manufacturer's algorithm displays the seasonal cycle that is closest to BLH-RS, while KABL and ADABL both overestimate BLH. The inter-quartile range (shaded areas) is large for all estimates.</p>
      <p id="d1e3820">At the Trappes site (Fig. <xref ref-type="fig" rid="Ch1.F10"/>c), KABL and ADABL also overestimate BLH in comparison to the BLH-RS, while the manufacturer's estimate is close. The seasonal cycle is more visible at Trappes than at Brest, and all BLH estimates are higher in summer than in winter. The most pronounced cycle is given by KABL, while the least pronounced cycle is given by BLH-RS. The inter-quartile ranges are also very large, especially in summer, reflecting the variation in BLH between  day and night.</p>
      <p id="d1e3825">Figure <xref ref-type="fig" rid="Ch1.F10"/>b and d show the diurnal cycle, where all values within the same 6 min period in the day were averaged. Because the radiosondes are only launched twice a day, at 11:15 and 23:15 UTC, an equivalent BLH-RS diurnal cycle cannot be drawn. However, we used the average and quartile values at these times as checkpoints for the other estimates. The manufacturer's and KABL estimates both have very smooth diurnal cycles, with lower BLH at night and maximum BLH at around 15:00 UTC at the Trappes site and around 13:00 UTC at the Brest site. The KABL average is always higher than that calculated by the manufacturer's algorithm. The ADABL estimation has a very different diurnal cycle, similar to the conceptual image we have of the boundary layer. Indeed, ADABL was trained using hand-classified BLHs that reflect this conceptual image. Therefore, it is not surprising that ADABL reproduces this image well; however, it may fail to adapt to special cases. It appears that the “time” predictor (the number of seconds since midnight) has a large influence that is not balanced by the other predictors. This is likely because ADABL was trained on only two dates, resulting in an unbalanced importance for sunrise and sunset on these particular days and at these locations. To balance this importance, the AdaBoost algorithm needs to be trained on more days and at more sites with a representative selection of cases.</p>
</sec>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Case study</title>
      <p id="d1e3839">The chosen case study was for 19 April 2017, at the Trappes site. The boundary layer was clearly visible and had nearly all the features of the conceptual image. The case study was for a day that was not included in the ADABL training set so as not to bias the comparison in favor of ADABL.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F11" specific-use="star"><?xmltex \currentcnt{11}?><?xmltex \def\figurename{Figure}?><label>Figure 11</label><caption><p id="d1e3844">Case study: BLH estimates using different methods on 19 April 2017, at the Trappes site. Superimposed on the lidar range-corrected signal are BLH estimates from KABL (dotted blue line), the manufacturer’s algorithm (denoted INDUS, dotted orange line), and ADABL (dotted green line). The two crosses at 11:15 and 23:15 UTC indicate the radiosonde estimates for that day, and the two icons at 05:50 and 19:49 UTC represent the time of sunrise and sunset, respectively.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f11.png"/>

        </fig>

      <p id="d1e3853">Figure <xref ref-type="fig" rid="Ch1.F11"/> represents the range-corrected copolarized backscatter signal (RCS<inline-formula><mml:math id="M147" display="inline"><mml:msub><mml:mi/><mml:mi mathvariant="normal">co</mml:mi></mml:msub></mml:math></inline-formula>) in shaded colors. The <inline-formula><mml:math id="M148" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis indicates the hour of the day (UTC), and the <inline-formula><mml:math id="M149" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis indicates the height (meters above ground level). The different BLH estimates are represented by dotted lines: blue indicates KABL, orange indicates the manufacturer's algorithm, and green indicates ADABL. At the beginning of the day, there is a thick residual layer containing some plumes. Both KABL and the manufacturer's algorithm include these plumes in the boundary layer. Conversely, ADABL gives a very low<?pagebreak page4349?> estimate where there is no visible frontier. In the morning (from 08:00 to 12:00 UTC), all the algorithms capture the morning transition reasonably well. However, KABL includes more irrelevant estimates (selecting remnants of the surface layer) than the other methods and ADABL gives an estimate that is too high for no apparent reason around 12:00 UTC. During the day, ADABL sticks to the top of the boundary layer, the manufacturer's algorithm sticks to the surface layer (which is very visible), and KABL oscillates between the two. The evening transition is blurry; the signal from the surface layer slowly increases, as the mixed layer decays into a residual layer. KABL locates this transition very early (around 17:00 UTC), when it stops oscillating and sticks to the surface layer. ADABL makes the transition more smoothly, from 19:00 to 22:00 UTC. The manufacturer's algorithm is the last to make the transition, at around 23:00 UTC, and the transition then occurs very sharply. We can conclude from this case study that none of the algorithms perfectly capture the boundary layer. Some of the limitations are physical; e.g., the evening transition is ill defined, resulting in disagreement between the algorithms. BLH-RS at 23:15 UTC is close to the lower boundary of the lidar range. This highlights the fact that BLHs below 120 m are not rare and will not be detected by lidar if the BLH is in the lidar blind zone. Some of the other limitations are algorithmic; KABL has an unfortunate tendency to oscillate between several candidates for the top of the boundary layer (surface layer or clouds), and ADABL too closely reproduces the features of the days it has been trained on (e.g., night estimates and morning transitions).</p>
</sec>
</sec>
<sec id="Ch1.S5">
  <label>5</label><title>Discussion and prospects</title>
<sec id="Ch1.S5.SS1">
  <label>5.1</label><title>Algorithm maturity</title>
      <p id="d1e3897">Both algorithms examined here are not yet mature when applied in this context. <inline-formula><mml:math id="M150" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithms have already been used to detect BLHs in previous studies <xref ref-type="bibr" rid="bib1.bibx56 bib1.bibx7 bib1.bibx57 bib1.bibx43" id="paren.59"/>; therefore, they comprise a more mature method. This is visible in this paper via the level of investigation, which was much higher for KABL than for ADABL.
Concerning boosting, this is the first time, to our knowledge, that such an algorithm has been tested on this type of problem; therefore, ADABL is a completely new algorithm. Yet, it outperforms KABL and competes favorably with the manufacturer's algorithm despite raising training issues.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F12" specific-use="star"><?xmltex \currentcnt{12}?><?xmltex \def\figurename{Figure}?><label>Figure 12</label><caption><p id="d1e3912">Cluster labels for KABL output on a time–altitude grid for 19 April 2017, at Trappes, with KABL BLH (black line) estimated using <bold>(a)</bold> the centroids initialized at random and <bold>(b)</bold> the centroids initialized at given values.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/4335/2021/amt-14-4335-2021-f12.png"/>

        </fig>

</sec>
<sec id="Ch1.S5.SS2">
  <label>5.2</label><title>Time and altitude continuity</title>
      <p id="d1e3935">The oscillations observed in Fig. <xref ref-type="fig" rid="Ch1.F11"/> are unrealistic and need to be avoided. They occur with KABL because clusters do not always have vertical persistence, as shown in Fig. <xref ref-type="fig" rid="Ch1.F12"/>. One can see the cluster labels on a time–altitude grid for the same day as in Fig. <xref ref-type="fig" rid="Ch1.F11"/>. When the initialization is random (Fig. <xref ref-type="fig" rid="Ch1.F12"/>a, <monospace>init</monospace> <inline-formula><mml:math id="M151" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> random, default settings in <inline-formula><mml:math id="M152" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means), the labels are also random. Only the transition of the labels on a profile is important. When the initialization is given (Fig. <xref ref-type="fig" rid="Ch1.F12"/>b, <monospace>init</monospace> <inline-formula><mml:math id="M153" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> given, retained settings in KABL), the labels can be identified. The blue cluster starts from a very high attenuated backscatter coefficient (it detects clouds and the shallow morning boundary layer); the red cluster starts from a high attenuated backscatter coefficient (it detects the mixed layer or residual layer); and the green cluster starts from a low attenuated backscatter coefficient (it detects the free atmosphere). Oscillations occur when some points are identified as free atmosphere in the middle of the boundary layer. In the case study presented here (Figs. <xref ref-type="fig" rid="Ch1.F11"/> and <xref ref-type="fig" rid="Ch1.F12"/>b), this happens in the afternoon, when the blue cluster (starting from a very high attenuated backscatter coefficient) gathers a few points near the first measurements and the overhead artifacts because a very high attenuated backscatter coefficient is irrelevant under these conditions.</p>
      <p id="d1e3981">Several prospects exist to enforce vertical persistence of the clusters in KABL; we list here four examples. First, the problem identified in Fig. <xref ref-type="fig" rid="Ch1.F12"/>b could be solved by choosing the number of clusters automatically. However, this option was tested (case where <monospace>n_clusters</monospace> is auto in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>) and it did not solve this issue. The sensitivity analysis showed that <monospace>n_clusters</monospace> <inline-formula><mml:math id="M154" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> auto led to higher discrepancies with respect to BLH-RS. In fact, this setting usually gives low estimates of BLH because the automatically chosen number of clusters is usually high. A more advanced strategy to automatically choose the number of clusters <xref ref-type="bibr" rid="bib1.bibx55" id="paren.60"><named-content content-type="pre">e.g.,</named-content></xref> might get around this issue. The vertical persistence of clusters can be enforced by adding altitude to the KABL predictors. This can be done thanks to post-processing, for example, with a moving average or by imposing a maximum BLH growth rate <xref ref-type="bibr" rid="bib1.bibx41" id="paren.61"/>. The distance used in <inline-formula><mml:math id="M155" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means could also be modified to incorporate these constraints, for example, by adding penalty terms.</p>
      <p id="d1e4017">In ADABL, time and altitude continuity are ensured because they are within the predictors. However, ADABL yields BLH estimates that are too similar to the BLHs in the training set. Removing time and/or altitude from the predictors should be considered to force the algorithm to rely more on the measurements. Further, the sensitivity analysis presented here for KABL needs to be performed for ADABL.</p>
</sec>
<sec id="Ch1.S5.SS3">
  <label>5.3</label><title>Real-time estimation</title>
      <p id="d1e4028">Even though it was not necessary for this study, all of the algorithms studied here can be used in real time. As soon as a lidar profile is available, BLH estimations can be performed instantaneously.<fn id="Ch1.Footn4"><p id="d1e4031">Both KABL and ADABL need less than 1 s to run a single profile.</p></fn> However, KABL suffers from undesirable oscillations from one profile to the next. A method to filter these oscillations is needed but would disable the “real-time” feature of the algorithm. In addition, the hour of the day needs to be explicitly passed in a periodic function. This has not been done here because we worked only on 24 h time periods.</p>
</sec>
<?pagebreak page4350?><sec id="Ch1.S5.SS4">
  <label>5.4</label><title>Quality of the evaluation</title>
      <p id="d1e4043">Even though we made an effort to sort the meteorological conditions using ancillary data, the 2-year comparison still mixes heterogeneous conditions. In addition, the results are clearly different at the sites studied here, emphasizing the importance of local conditions. A more precise casting of the meteorological conditions with atmospheric stability indices or large-scale insights would lead to a better understanding of the strengths and weaknesses of the algorithms. The importance of the sites needs to be investigated by extending the study to a larger number of sites with different environments. A more careful examination of cloudy days also needs to be performed. Cases where the cloud bases were below 3 km were filtered out from our study. However, cases where clouds reside inside the inversion should be detected by KABL as an extra cluster; further studies are required to confirm this behavior. In addition, ADABL was not specifically trained to deal with cloudy situations. Further studies to determine how ADABL behaves without training and how it could be appropriately trained are required.</p>
</sec>
<sec id="Ch1.S5.SS5">
  <label>5.5</label><title>Quality of the reference</title>
      <p id="d1e4055">Radiosonde profiles are usually regarded as the best reference for BLH. However, the derivation of BLH from such measurements is contentious because several methods exist and some strongly disagree. Moreover, RS measurements cannot be used to assess the full diurnal cycle of BLH. This is a clear limitation of this study because the RS measurements cannot determine if the difference between the diurnal cycle of ADABL and those of the other methods represents an improvement. Therefore, a very interesting project would be to use a dedicated field experiment with high-frequency tethersonde or other continuously running instruments as a reference. For example, microwave radiometers are good candidates because they provide information that is not based on aerosols and the derivation of BLH from these instruments is routine <xref ref-type="bibr" rid="bib1.bibx10" id="paren.62"/>.</p>
</sec>
<sec id="Ch1.S5.SS6">
  <label>5.6</label><title>ADABL – training</title>
      <p id="d1e4069">ADABL already shows good performance when trained on only two dates. Most of its bad estimates result from the short length of its training period. Therefore, a short-term project would be to label more dates with various meteorological conditions. However, the dependence of ADABL on training makes it sensitive to instrumentation settings and calibrations. Even though the effect of a calibration or the evolution of an instrumental device has not been studied, it is likely that training needs to be repeated after each calibration or change in the instrumental device. Therefore, two strategies are possible for training ADABL: remove the influence of calibration prior to training (this would require knowing the instrumental constants for all of the devices) or train it to deal with differences (this would require including as many different devices as possible in the training set, which would then become very large). In any case, the main limitation will be the need to label the entire training dataset (a priori by human experts).</p>
</sec>
<sec id="Ch1.S5.SS7">
  <label>5.7</label><title>KABL – training-less</title>
      <p id="d1e4080">KABL appeared to perform the least well in this study; however, there are interesting prospects to improve its performance. KABL does not require any training; therefore, it is less dependent on instrumentation settings and calibrations. Because it is not strongly dependent on the instrumental devices, it can be used on backscatter profiles made by other instruments (e.g., ceilometers). Moreover, other profiles<?pagebreak page4351?> besides the backscatter intensity can be added as additional predictors for unsupervised learning after normalization. Therefore, the concept of KABL can be advanced further to create synergy between multiple remote sensing instruments. Microwave radiometers are good candidates because they have comparable time resolutions to lidar and provide independent information concerning the thermal stratification of the boundary layer. Cloud radars also have comparable time resolutions to lidar and provide additional independent information.</p>
</sec>
<sec id="Ch1.S5.SS8">
  <label>5.8</label><title>Quality flags</title>
      <p id="d1e4091">Currently, no quality flags for the estimation are provided. One approach would be to use the internal scores (i.e., silhouette, Davies–Bouldin, and Calinski–Harabasz defined in Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>) as quality flags; however, further study is required to determine whether these metrics can serve as reliable quality flags.</p>
</sec>
</sec>
<sec id="Ch1.S6" sec-type="conclusions">
  <label>6</label><title>Conclusions</title>
      <p id="d1e4105">This paper described two algorithms based on machine learning to estimate the boundary layer height from aerosol lidar measurements. The first, KABL, is based on the <inline-formula><mml:math id="M156" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithm. The second, ADABL, is based on the AdaBoost algorithm. Both algorithms take the same input file, 1 d of data generated by the raw2l1 routine, and produce similar output, a BLH time series for the input day. KABL is a non-supervised algorithm that looks for a natural separation in the backscatter signals between the boundary layer and the free atmosphere. ADABL is a supervised algorithm that fits a large number of decision trees in a labeled dataset and aggregates them in an intelligent manner to provide a good prediction. KABL, ADABL, and the lidar manufacturer's algorithm were tested on a 2-year dataset taken from the Météo-France operational lidar network. The Trappes and Brest sites were chosen because of their different climates and the availability of regular RS measurements, which were used as a reference.</p>
      <p id="d1e4115">A large discrepancy between RMSE and the correlation with the radiosondes was observed between the two sites. At the Trappes site, KABL and ADABL outperformed the manufacturer's algorithm, while the opposite occurred at the Brest site. At both sites, ADABL performed better than KABL (higher correlation and lower error) and the manufacturer's algorithm performed well.
By analyzing the seasonal and diurnal cycles, we determined that the KABL and manufacturer's estimates have similar behavior; however, the KABL estimates are always higher by approximately 200 m. ADABL generates the most pronounced diurnal cycle, with a pattern that is very similar to the expected diurnal cycle; however, its results depend greatly on the days it has been trained on. In particular, the sunset and sunrise times of these days over-influenced the ADABL estimate. In the case study, we saw that both algorithms perform well overall; however, we identified several algorithmic limitations; e.g., KABL tended to oscillate between several candidates for the top of the boundary layer (surface layers or clouds) and ADABL was overly constrained by the days it was trained on (e.g., the night estimate and morning transition). In summary, ADABL is promising but has training issues that need to be resolved; KABL has lower performance but is much more versatile, and the manufacturer's algorithm using a wavelet covariance transform performs well with little tuning but is not open source. A wide range of future developments is available for ADABL and KABL, the most immediate being that the training set of ADABL can be enhanced, time and altitude continuity can be enforced in the KABL estimation, and both can be compared to high-temporal-resolution RS measurements.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e4122">The KABL source code is available to and usable by all users, including commercial users. The code is freely available under an open-source license at the following link:
<uri>https://github.com/ThomasRieutord/kabl</uri> (last access: 3 June 2021) <xref ref-type="bibr" rid="bib1.bibx44" id="paren.63"/>. It is made in Python 3.7 with regular statistics and machine-learning packages, namely, scikit-learn 0.20 <xref ref-type="bibr" rid="bib1.bibx39" id="paren.64"/> and SALib 1.3.7 <xref ref-type="bibr" rid="bib1.bibx27" id="paren.65"/>, which are open source and available under free licenses. The repository contains all the necessary features to run the code on raw2l1 outputs. Several days of data are also provided as examples.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e4140">TM implemented an initial version of the KABL code and performed the first comparisons to the RS data. SA extracted and processed all the data (lidar, radiosonde, and ancillary), made one of the hand-labeled BLHs, and participated actively in the writing of the manuscript. TR implemented the current versions of KABL and ADABL, made one of the hand-labeled BLHs, produced the figures, and actively participated in the writing of the manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e4146">The authors declare that they have no conflicts of interest.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d1e4152">This article is part of the special issue “Tropospheric profiling (ISTP11) (AMT/ACP inter-journal SI)”. It is a result of the 11th edition of the International Symposium on Tropospheric Profiling (ISTP), Toulouse, France, 20–24 May 2019.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e4158">The authors would like to thank Alexandre Paci, Alain Dabas, and Olivier Traullé for their helpful reading and comments. We would like to thank Marc-Antoine Drouin, from the Site Instrumental de Recherche par Télédétection Atmosphérique, for providing us with the link to raw2l1 and all the Météo-France agents who install and maintain the lidar network. We thank Martha Evonuk from Evonuk Scientific Editing<?pagebreak page4352?> (<uri>http://evonukscientificediting.com</uri>, last access: 14 December 2020) for editing a draft of this paper.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e4167">This paper was edited by E. J. O'Connor and reviewed by Anton Sokolov and three anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Arciszewska and McClatchey(2001)}}?><label>Arciszewska and McClatchey(2001)</label><?label arciszewska2001importance?><mixed-citation>
Arciszewska, C. and McClatchey, J.: The importance of meteorological data for
modelling air pollution using ADMS-Urban, Meteorol. Appl., 8, 345–350, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Arthur and Vassilvitskii(2007)}}?><label>Arthur and Vassilvitskii(2007)</label><?label arthur2007kmeanspp?><mixed-citation>Arthur, D. and Vassilvitskii, S.: <inline-formula><mml:math id="M157" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means<inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mo>+</mml:mo></mml:mrow></mml:math></inline-formula>: The advantages of careful seeding, in: Proceedings of the eighteenth annual ACM-SIAM symposium on Discrete algorithms, Society for Industrial and Applied Mathematics, 1027–1035, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Besse et~al.(2018)Besse, Guillouet, and Laurent}}?><label>Besse et al.(2018)Besse, Guillouet, and Laurent</label><?label besse2018wikistat?><mixed-citation>
Besse, P., Guillouet, B., and Laurent, B.: Wikistat 2.0: Educational Resources for Artificial Intelligence, arXiv: preprint, arXiv:1810.02688, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Breiman et~al.(1984)Breiman, Friedman, Olshen, and
Stone}}?><label>Breiman et al.(1984)Breiman, Friedman, Olshen, and
Stone</label><?label breiman1984classification?><mixed-citation>
Breiman, L., Friedman, J., Olshen, R., and Stone, C.: Classification and
Regression Trees, Wadsworth, 1984.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Brilouet et~al.(2017)Brilouet, Durand, and
Canut}}?><label>Brilouet et al.(2017)Brilouet, Durand, and
Canut</label><?label brilouet2017marine?><mixed-citation>
Brilouet, P.-E., Durand, P., and Canut, G.: The marine atmospheric boundary
layer under strong wind conditions: Organized turbulence structure and flux
estimates by airborne measurements, J. Geophys. Res.-Atmos., 122, 2115–2130, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Brooks(2003)}}?><label>Brooks(2003)</label><?label brooks2003finding?><mixed-citation>
Brooks, I. M.: Finding boundary layer top: Application of a wavelet covariance transform to lidar backscatter profiles, J. Atmos. Ocean. Tech., 20, 1092–1105, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Caicedo et~al.(2017)Caicedo, Rappengl{\"{u}}ck, Lefer, Morris, Toledo, and Delgado}}?><label>Caicedo et al.(2017)Caicedo, Rappenglück, Lefer, Morris, Toledo, and Delgado</label><?label caicedo2017comparison?><mixed-citation>Caicedo, V., Rappenglück, B., Lefer, B., Morris, G., Toledo, D., and Delgado, R.: Comparison of aerosol lidar retrieval methods for boundary layer height detection using ceilometer aerosol backscatter data, Atmos. Meas. Tech., 10, 1609–1622, <ext-link xlink:href="https://doi.org/10.5194/amt-10-1609-2017" ext-link-type="DOI">10.5194/amt-10-1609-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{Cali\'{n}ski and Harabasz(1974)}}?><label>Caliński and Harabasz(1974)</label><?label calinski1974dendrite?><mixed-citation>
Caliński, T. and Harabasz, J.: A dendrite method for cluster analysis,
Commun. Stat.-Theor. Meth., 3, 1–27, 1974.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{Campbell et~al.(2002)Campbell, Hlavka, Welton, Flynn, Turner,
Spinhirne, Scott, and Hwang}}?><label>Campbell et al.(2002)Campbell, Hlavka, Welton, Flynn, Turner,
Spinhirne, Scott, and Hwang</label><?label campbell2002mplprocessing?><mixed-citation>
Campbell, J. R., Hlavka, D. L., Welton, E. J., Flynn, C. J., Turner, D. D.,
Spinhirne, J. D., Scott III, V. S., and Hwang, I. H.: Full-Time, Eye-Safe
Cloud and Aerosol Lidar Observation at Atmospheric Radiation Measurement
Program Sites: Instruments and Data Processing, J. Atmos. Ocean. Tech., 19, 431–442, 2002.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Cimini et~al.(2013)Cimini, De~Angelis, Dupont, Pal, and
Haeffelin}}?><label>Cimini et al.(2013)Cimini, De Angelis, Dupont, Pal, and
Haeffelin</label><?label cimini2013mixing?><mixed-citation>Cimini, D., De Angelis, F., Dupont, J.-C., Pal, S., and Haeffelin, M.: Mixing layer height retrievals by multichannel microwave radiometer observations, Atmos. Meas. Tech., 6, 2941–2951, <ext-link xlink:href="https://doi.org/10.5194/amt-6-2941-2013" ext-link-type="DOI">10.5194/amt-6-2941-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Cohn and Angevine(2000)}}?><label>Cohn and Angevine(2000)</label><?label cohn2000boundary?><mixed-citation>
Cohn, S. A. and Angevine, W. M.: Boundary layer height and entrainment zone
thickness measured by lidars and wind-profiling radars, J. Appl. Meteorol., 39, 1233–1247, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{Collaud~Coen et~al.(2014)Collaud~Coen, Praz, Haefele, Ruffieux,
Kaufmann, and Calpini}}?><label>Collaud Coen et al.(2014)Collaud Coen, Praz, Haefele, Ruffieux,
Kaufmann, and Calpini</label><?label collaud2014determination?><mixed-citation>Collaud Coen, M., Praz, C., Haefele, A., Ruffieux, D., Kaufmann, P., and Calpini, B.: Determination and climatology of the planetary boundary layer height above the Swiss plateau by in situ and remote sensing measurements as well as by the COSMO-2 model, Atmos. Chem. Phys., 14, 13205–13221, <ext-link xlink:href="https://doi.org/10.5194/acp-14-13205-2014" ext-link-type="DOI">10.5194/acp-14-13205-2014</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{Davies and Bouldin(1979)}}?><label>Davies and Bouldin(1979)</label><?label davies1979cluster?><mixed-citation>Davies, D. L. and Bouldin, D. W.: A cluster separation measure, in: IEEE
transactions on pattern analysis and machine intelligence, 12–14 April 1978,
Princeton, NJ, 224–227, <ext-link xlink:href="https://doi.org/10.1109/TPAMI.1979.4766909" ext-link-type="DOI">10.1109/TPAMI.1979.4766909</ext-link>, 1979.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Davison and Hinkley(1997)}}?><label>Davison and Hinkley(1997)</label><?label davison1997bootstrap?><mixed-citation>
Davison, A. C. and Hinkley, D. V.: Bootstrap methods and their application,
Cambridge University Press, Cambridge, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{de~Bruine et~al.(2017)De~Bruine, Apituley, Donovan, Klein~Baltink,
and de~Haij}}?><label>de Bruine et al.(2017)De Bruine, Apituley, Donovan, Klein Baltink,
and de Haij</label><?label debruine2017pathfinder?><mixed-citation>de Bruine, M., Apituley, A., Donovan, D. P., Klein Baltink, H., and de Haij, M. J.: Pathfinder: applying graph theory to consistent tracking of daytime mixed layer height with backscatter lidar, Atmos. Meas. Tech., 10, 1893–1909, <ext-link xlink:href="https://doi.org/10.5194/amt-10-1893-2017" ext-link-type="DOI">10.5194/amt-10-1893-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Desgraupes(2013)}}?><label>Desgraupes(2013)</label><?label desgraupes2013clustering?><mixed-citation>Desgraupes, B.: Clustering indices, University of Paris Ouest-Lab Modal'X, available at: <uri>https://cran.biodisk.org/web/packages/clusterCrit/vignettes/clusterCrit.pdf</uri> (last access: 7 June 2021), 34 pp., 2013.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Dupont et~al.(2016)Dupont, Haeffelin, Badosa, Elias, Favez, Petit,
Meleux, Sciare, Crenn, and Bonne}}?><label>Dupont et al.(2016)Dupont, Haeffelin, Badosa, Elias, Favez, Petit,
Meleux, Sciare, Crenn, and Bonne</label><?label dupont2016role?><mixed-citation>
Dupont, J.-C., Haeffelin, M., Badosa, J., Elias, T., Favez, O., Petit, J.,
Meleux, F., Sciare, J., Crenn, V., and Bonne, J.: Role of the boundary layer
dynamics effects on an extreme air pollution event in Paris, Atmos. Environ., 141, 571–579, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Flynn et~al.(2007)Flynn, Mendoza, Zheng, and
Mathur}}?><label>Flynn et al.(2007)Flynn, Mendoza, Zheng, and
Mathur</label><?label flynn2007depolarmpl?><mixed-citation>
Flynn, C. J., Mendoza, A., Zheng, Y., and Mathur, S.: Novel polarization-sensitive micropulse lidar measurement technique, Opt. Express,
15, 2785–2790, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{Freund and Schapire(1997)}}?><label>Freund and Schapire(1997)</label><?label freund1997decision?><mixed-citation>
Freund, Y. and Schapire, R. E.: A decision-theoretic generalization of on-line learning and an application to boosting, J. Comput. Syst. Sci., 55, 119–139, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Gamage and Hagelberg(1993)}}?><label>Gamage and Hagelberg(1993)</label><?label gamage1993detection?><mixed-citation>
Gamage, N. and Hagelberg, C.: Detection and analysis of microfronts and
associated coherent events using localized transforms, J. Atmos. Sci., 50, 750–756, 1993.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Guo et~al.(2016)Guo, Miao, Zhang, Liu, Li, Zhang, He, Lou, Yan, Bian et~al.}}?><label>Guo et al.(2016)Guo, Miao, Zhang, Liu, Li, Zhang, He, Lou, Yan, Bian et al.</label><?label guo2016climatology?><mixed-citation>Guo, J., Miao, Y., Zhang, Y., Liu, H., Li, Z., Zhang, W., He, J., Lou, M., Yan, Y., Bian, L., and Zhai, P.: The climatology of planetary boundary layer height in China derived from radiosonde and reanalysis data, Atmos. Chem. Phys., 16, 13309–13319, <ext-link xlink:href="https://doi.org/10.5194/acp-16-13309-2016" ext-link-type="DOI">10.5194/acp-16-13309-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Haefele et~al.(2016)Haefele, Hervo, Turp, Lampin J-L, Lehmann
et~al.}}?><label>Haefele et al.(2016)Haefele, Hervo, Turp, Lampin J-L, Lehmann
et al.</label><?label eprof2016teco?><mixed-citation>
Haefele, A., Hervo, M., Turp, M., Lampin, J.-L., Haeffelin, M., and Lehmann, V.: The E-PROFILE network for the operational measurement of wind and aerosol profiles over Europe, in: Proceedings of WMO Technical Conference on Meteorological and Environmental Instruments and Methods of Observation, CIMO TECO 2016, Madrid, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Haeffelin et~al.(2012)Haeffelin, Angelini, Morille, Martucci, Frey,
Gobbi, Lolli, O'dowd, Sauvage, Xueref-R{\'{e}}my
et~al.}}?><label>Haeffelin et al.(2012)Haeffelin, Angelini, Morille, Martucci, Frey,
Gobbi, Lolli, O'dowd, Sauvage, Xueref-Rémy
et al.</label><?label haeffelin2012evaluation?><mixed-citation>
Haeffelin, M., Angelini, F., Morille, Y., Martucci, G., Frey, S., Gobbi, G.,
Lolli, S., O'dowd, C., Sauvage, L., Xueref-Rémy, I., Wastine, B., and Feist, D. G.: Evaluation of mixing-height retrievals from automatic profiling lidars and ceilometers in view of future integrated networks in Europe, Bound.-Lay. Meteorol., 143, 49–75, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Hastie et~al.(2009)Hastie, Tibshirani, and
Friedman}}?><label>Hastie et al.(2009)Hastie, Tibshirani, and
Friedman</label><?label hastie2009elements?><mixed-citation>
Hastie, T., Tibshirani, R., and Friedman, J.: The Elements of Statistical
Learning: Data Mining, Inference, and Prediction, Springer Science &amp; Business Media, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Hayden et~al.(1997)Hayden, Anlauf, Hoff, Strapp, Bottenheim, Wiebe,
Froude, Martin, Steyn, and McKendry}}?><label>Hayden et al.(1997)Hayden, Anlauf, Hoff, Strapp, Bottenheim, Wiebe,
Froude, Martin, Steyn, and McKendry</label><?label hayden1997?><mixed-citation>
Hayden, K., Anlauf, K., Hoff, R., Strapp, J., Bottenheim, J., Wiebe, H.,
Froude, F., Martin, J., Steyn, D., and McKendry, I.: The vertical chemical
and meteorological structure of the boundary layer in the Lower Fraser Valley
during Pacific'93, Atmos. Environ., 31, 2089–2105, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Hennemuth and Lammert(2006)}}?><label>Hennemuth and Lammert(2006)</label><?label hennemuth2006determination?><mixed-citation>
Hennemuth, B. and Lammert, A.: Determination of the atmospheric boundary layer height from radiosonde and lidar backscatter, Bound.-Lay. Meteorol., 120, 181–200, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Herman and Usher(2017)}}?><label>Herman and Usher(2017)</label><?label salib2017?><mixed-citation>Herman, J. and Usher, W.: SALib: An open-source Python library for
Sensitivity Analysis, J. Open Source Softw., 2, 97, <ext-link xlink:href="https://doi.org/10.21105/joss.00097" ext-link-type="DOI">10.21105/joss.00097</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Hintze and Nelson(1998)}}?><label>Hintze and Nelson(1998)</label><?label hintze1998?><mixed-citation>
Hintze, J. L. and Nelson, R. D.: Violin Plots: A Box Plot-Density Trace
Synergism, Am. Statist., 52, 181–184, 1998.</mixed-citation></ref>
      <?pagebreak page4353?><ref id="bib1.bibx29"><?xmltex \def\ref@label{{Iooss and Lema\^{i}tre(2015)}}?><label>Iooss and Lemaître(2015)</label><?label iooss2015review?><mixed-citation>
Iooss, B. and Lemaître, P.: A review on global sensitivity analysis
methods, in: Uncertainty management in simulation-optimization of complex
systems, Springer, 101–122, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Jain et~al.(1999)Jain, Murty, and Flynn}}?><label>Jain et al.(1999)Jain, Murty, and Flynn</label><?label jain1999data?><mixed-citation>
Jain, A. K., Murty, M. N., and Flynn, P. J.: Data clustering: a review, ACM
Comput. Surv., 31, 264–323, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Kotthaus and Grimmond(2018)}}?><label>Kotthaus and Grimmond(2018)</label><?label kotthaus2018cabam?><mixed-citation>
Kotthaus, S. and Grimmond, C. S. B.: Atmospheric boundary-layer characteristics from ceilometer measurements. Part 1: A new method to track mixed layer height and classify clouds, Q. J. Roy. Meteorol. Soc., 144, 1525–1538, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Krizhevsky et~al.(2012)Krizhevsky, Sutskever, and
Hinton}}?><label>Krizhevsky et al.(2012)Krizhevsky, Sutskever, and
Hinton</label><?label krizhevsky2012imagenet?><mixed-citation>Krizhevsky, A., Sutskever, I., and Hinton, G. E.: Imagenet classification with deep convolutional neural networks, in: Advances in neural information
processing systems, Curran Associates, Inc., 1097–1105, available at: <uri>https://proceedings.neurips.cc/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf</uri>
(last access: 7 June 2021), 2012.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{{LeCun et~al.(2015)LeCun, Bengio, and Hinton}}?><label>LeCun et al.(2015)LeCun, Bengio, and Hinton</label><?label lecun2015deep?><mixed-citation>
LeCun, Y., Bengio, Y., and Hinton, G.: Deep learning, Nature, 521, 436–444,
2015.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{Lenschow et~al.(2012)Lenschow, Lothon, Mayor, Sullivan, and
Canut}}?><label>Lenschow et al.(2012)Lenschow, Lothon, Mayor, Sullivan, and
Canut</label><?label lenschow2012comparison?><mixed-citation>
Lenschow, D. H., Lothon, M., Mayor, S. D., Sullivan, P. P., and Canut, G.: A
comparison of higher-order vertical velocity moments in the convective boundary layer from lidar with in situ measurements and large-eddy
simulation, Bound.-Lay. Meteorol., 143, 107–123, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Melfi et~al.(1985)Melfi, Spinhirne, Chou, and Palm}}?><label>Melfi et al.(1985)Melfi, Spinhirne, Chou, and Palm</label><?label melfi1985lidar?><mixed-citation>
Melfi, S., Spinhirne, J., Chou, S., and Palm, S.: Lidar observations of
vertically organized convection in the planetary boundary layer over the
ocean, J. Clim. Appl. Meteorol., 24, 806–821, 1985.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{Menut et~al.(1999)Menut, Flamant, Pelon, and
Flamant}}?><label>Menut et al.(1999)Menut, Flamant, Pelon, and
Flamant</label><?label menut1999urban?><mixed-citation>
Menut, L., Flamant, C., Pelon, J., and Flamant, P. H.: Urban boundary-layer
height determination from lidar measurements over the Paris area, Appl. Optics, 38, 945–954, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{{Mohan et~al.(2011)Mohan, Bhati, Sreenivas, and
Marrapu}}?><label>Mohan et al.(2011)Mohan, Bhati, Sreenivas, and
Marrapu</label><?label mohan2011performance?><mixed-citation>
Mohan, M., Bhati, S., Sreenivas, A., and Marrapu, P.: Performance evaluation of AERMOD and ADMS-urban for total suspended particulate matter concentrations in megacity Delhi, Aerosol Air Qual. Res., 11, 883–894, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{{Morille et~al.(2007)Morille, Haeffelin, Drobinski, and
Pelon}}?><label>Morille et al.(2007)Morille, Haeffelin, Drobinski, and
Pelon</label><?label morille2007strat?><mixed-citation>
Morille, Y., Haeffelin, M., Drobinski, P., and Pelon, J.: STRAT: An automated
algorithm to retrieve the vertical structure of the atmosphere from
single-channel lidar data, J. Atmos. Ocean. Tech., 24, 761–775, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{{Pedregosa et~al.(2011)Pedregosa, Varoquaux, Gramfort, Michel,
Thirion, Grisel, Blondel, Prettenhofer, Weiss, Dubourg, Vanderplas, Passos,
Cournapeau, Brucher, Perrot, and Duchesnay}}?><label>Pedregosa et al.(2011)Pedregosa, Varoquaux, Gramfort, Michel,
Thirion, Grisel, Blondel, Prettenhofer, Weiss, Dubourg, Vanderplas, Passos,
Cournapeau, Brucher, Perrot, and Duchesnay</label><?label scikit-learn?><mixed-citation>
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,
O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J.,
Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E.:
Scikit-learn: Machine Learning in Python, J. Mach. Learn. Res., 12, 2825–2830, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{{Pollard(1981)}}?><label>Pollard(1981)</label><?label pollard1981?><mixed-citation>Pollard, D.: Strong consistency of <inline-formula><mml:math id="M159" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means clustering, Ann. Statist., 9, 135–140, 1981.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{{Poltera et~al.(2017)Poltera, Martucci, Collaud~Coen, Hervo,
Emmenegger, Henne, Brunner, and Haefele}}?><label>Poltera et al.(2017)Poltera, Martucci, Collaud Coen, Hervo,
Emmenegger, Henne, Brunner, and Haefele</label><?label poltera2017pathfinderturb?><mixed-citation>Poltera, Y., Martucci, G., Collaud Coen, M., Hervo, M., Emmenegger, L., Henne, S., Brunner, D., and Haefele, A.: PathfinderTURB: an automatic boundary layer algorithm. Development, validation and application to study the impact on in situ measurements at the Jungfraujoch, Atmos. Chem. Phys., 17, 10051–10070, <ext-link xlink:href="https://doi.org/10.5194/acp-17-10051-2017" ext-link-type="DOI">10.5194/acp-17-10051-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{{Rieutord(2017)}}?><label>Rieutord(2017)</label><?label rieutord2017phd?><mixed-citation>Rieutord, T.: Sensitivity analysis of a filtering algorithm for wind lidar
measurements, PhD thesis, Institut National Polytechnique de Toulouse, Toulouse, available at: <uri>https://oatao.univ-toulouse.fr/19457/</uri> (last access: 7 June 2021), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{{Rieutord et~al.(2014)Rieutord, Brewer, and
Hardesty}}?><label>Rieutord et al.(2014)Rieutord, Brewer, and
Hardesty</label><?label rieutord2014automatic?><mixed-citation>
Rieutord, T., Brewer, W. A., and Hardesty, R. M.: Automatic detection of
boundary layer height using Doppler lidar measurements, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx44"><?xmltex \def\ref@label{{Rieutord et~al.(2021)}}?><label>Rieutord et al.(2021)</label><?label Rieutord2021?><mixed-citation>Rieutord, T., Aubert, S., and Machado, T.: KABL, Github, available at: <uri>https://github.com/ThomasRieutord/kabl</uri>, last access: 3 June 2021.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{Rousseeuw(1987)}}?><label>Rousseeuw(1987)</label><?label rousseeuw1987silhouettes?><mixed-citation>
Rousseeuw, P. J.: Silhouettes: a graphical aid to the interpretation and
validation of cluster analysis, J. Comput. Appl. Math., 20, 53–65, 1987.</mixed-citation></ref>
      <ref id="bib1.bibx46"><?xmltex \def\ref@label{{Schapire(2013)}}?><label>Schapire(2013)</label><?label schapire2013explaining?><mixed-citation>
Schapire, R. E.: Explaining adaboost, in: Empirical inference, Springer, 37–52, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{{Seibert et~al.(2000)Seibert, Beyrich, Gryning, Joffre, Rasmussen, and Tercier}}?><label>Seibert et al.(2000)Seibert, Beyrich, Gryning, Joffre, Rasmussen, and Tercier</label><?label seibert2000review?><mixed-citation>
Seibert, P., Beyrich, F., Gryning, S.-E., Joffre, S., Rasmussen, A., and
Tercier, P.: Review and intercomparison of operational methods for the
determination of the mixing height, Atmos. Environ., 34, 1001–1027, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{Seidel et~al.(2010)Seidel, Ao, and Li}}?><label>Seidel et al.(2010)Seidel, Ao, and Li</label><?label seidel2010estimating?><mixed-citation>Seidel, D. J., Ao, C. O., and Li, K.: Estimating climatological planetary
boundary layer heights from radiosonde observations: Comparison of methods
and uncertainty analysis, J. Geophys. Res.-Atmos., 115, D16113, <ext-link xlink:href="https://doi.org/10.1029/2009JD013680" ext-link-type="DOI">10.1029/2009JD013680</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{Seidel et~al.(2012)Seidel, Zhang, Beljaars, Golaz, Jacobson, and
Medeiros}}?><label>Seidel et al.(2012)Seidel, Zhang, Beljaars, Golaz, Jacobson, and
Medeiros</label><?label seidel2012climatology?><mixed-citation>Seidel, D. J., Zhang, Y., Beljaars, A., Golaz, J.-C., Jacobson, A. R., and
Medeiros, B.: Climatology of the planetary boundary layer over the continental United States and Europe, J. Geophys. Res.-Atmos., 117, D17106, <ext-link xlink:href="https://doi.org/10.1029/2012JD018143" ext-link-type="DOI">10.1029/2012JD018143</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx50"><?xmltex \def\ref@label{{Seity et~al.(2011)Seity, Brousseau, Malardel, Hello, B{\'{e}}nard,
Bouttier, Lac, and Masson}}?><label>Seity et al.(2011)Seity, Brousseau, Malardel, Hello, Bénard,
Bouttier, Lac, and Masson</label><?label seity2011arome?><mixed-citation>
Seity, Y., Brousseau, P., Malardel, S., Hello, G., Bénard, P., Bouttier,
F., Lac, C., and Masson, V.: The AROME-France convective-scale operational
model, Mon. Weather Rev., 139, 976–991, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx51"><?xmltex \def\ref@label{{Selim and Ismail(1984)}}?><label>Selim and Ismail(1984)</label><?label selim1984?><mixed-citation>Selim, S. Z. and Ismail, M. A.: <inline-formula><mml:math id="M160" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula>-means-type algorithms: a generalized
convergence theorem and characterization of local optimality, IEEE T. Pattern Anal. Mach. Intel., PAMI-6, 81–87, <ext-link xlink:href="https://doi.org/10.1109/TPAMI.1984.4767478" ext-link-type="DOI">10.1109/TPAMI.1984.4767478</ext-link>, 1984.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx52"><?xmltex \def\ref@label{{Senff et~al.(1996)Senff, B{\"{o}}senberg, Peters, and
Schaberl}}?><label>Senff et al.(1996)Senff, Bösenberg, Peters, and
Schaberl</label><?label senff1996remote?><mixed-citation>
Senff, C., Bösenberg, J., Peters, G., and Schaberl, T.: Remote sensing of
turbulent ozone fluxes and the ozone budget in the convective boundary layer
with DIAL and Radar-RASS: A case study, Contrib. Atmos. Phys., 69, 161–176, 1996.</mixed-citation></ref>
      <ref id="bib1.bibx53"><?xmltex \def\ref@label{{Sobol(2001)}}?><label>Sobol(2001)</label><?label sobol2001global?><mixed-citation>
Sobol, I. M.: Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates, Math. Comput. Simul., 55, 271–280, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx54"><?xmltex \def\ref@label{{Stull(1988)}}?><label>Stull(1988)</label><?label stull1988?><mixed-citation>
Stull, R. B.: An introduction to boundary layer meteorology, in: vol. 13, Springer, 1988.</mixed-citation></ref>
      <ref id="bib1.bibx55"><?xmltex \def\ref@label{{Tibshirani et~al.(2001)Tibshirani, Walther, and
Hastie}}?><label>Tibshirani et al.(2001)Tibshirani, Walther, and
Hastie</label><?label tibshirani2001estimating?><mixed-citation>
Tibshirani, R., Walther, G., and Hastie, T.: Estimating the number of clusters in a data set via the gap statistic, J. Roy. Stat. Soc. B, 63, 411–423, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx56"><?xmltex \def\ref@label{{Toledo et~al.(2014)Toledo, C{\'{o}}rdoba-Jabonero, and
Gil-Ojeda}}?><label>Toledo et al.(2014)Toledo, Córdoba-Jabonero, and
Gil-Ojeda</label><?label toledo2014cluster?><mixed-citation>
Toledo, D., Córdoba-Jabonero, C., and Gil-Ojeda, M.: Cluster analysis: A
new approach applied to lidar measurements for atmospheric boundary layer
height estimation, J. Atmos. Ocean. Tech., 31, 422–436, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx57"><?xmltex \def\ref@label{{Toledo et~al.(2017)Toledo, C{\'{o}}rdoba-Jabonero, Adame, De~La~Morena, and Gil-Ojeda}}?><label>Toledo et al.(2017)Toledo, Córdoba-Jabonero, Adame, De La Morena, and Gil-Ojeda</label><?label toledo2017estimation?><mixed-citation>
Toledo, D., Córdoba-Jabonero, C., Adame, J. A., De La Morena, B., and
Gil-Ojeda, M.: Estimation of the atmospheric boundary layer height during
different atmospheric conditions: a comparison on reliability of several
methods applied to lidar measurements, Int. J. Remote Sens., 38, 3203–3218, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx58"><?xmltex \def\ref@label{{Ware et~al.(2016)Ware, Kort, DeCola, and Duren}}?><label>Ware et al.(2016)Ware, Kort, DeCola, and Duren</label><?label ware2016minimpl?><mixed-citation>
Ware, J., Kort, E. A., DeCola, P., and Duren, R.: Aerosol lidar observations of atmospheric mixing in Los Angeles: Climatology and implications for
greenhouse gas observations, J. Geophys. Res.-Atmos., 121, 9862–9878, 2016.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Deriving boundary layer height from aerosol lidar using  machine learning: KABL and ADABL algorithms</article-title-html>
<abstract-html><p>The atmospheric boundary layer height (BLH) is a key parameter for many meteorological applications, including air quality forecasts.
Several algorithms have been proposed to automatically estimate BLH from lidar backscatter profiles. However recent advances in computing have enabled new approaches using machine learning that are seemingly well suited to this problem. Machine learning can handle complex classification problems and can be trained by a human expert. This paper describes and compares two machine-learning methods, the <i>K</i>-means unsupervised algorithm and the AdaBoost supervised algorithm, to derive BLH from lidar backscatter profiles. The <i>K</i>-means for Atmospheric Boundary Layer (KABL) and AdaBoost for Atmospheric Boundary Layer (ADABL) algorithm codes used in this study are free and open source. Both methods were compared to reference BLHs derived from colocated radiosonde data over a 2-year period (2017–2018) at two Météo-France operational network sites (Trappes and Brest).
A large discrepancy between the root-mean-square error (RMSE) and correlation with radiosondes was observed between the two sites. At the Trappes site, KABL and ADABL outperformed the manufacturer's algorithm, while the performance was clearly reversed at the Brest site. We conclude that ADABL is a promising algorithm (RMSE of 550&thinsp;m at Trappes, 800&thinsp;m for manufacturer) but has training issues that need to be resolved; KABL has a lower performance (RMSE of 800&thinsp;m at Trappes) than ADABL but is much more versatile.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Arciszewska and McClatchey(2001)</label><mixed-citation>
Arciszewska, C. and McClatchey, J.: The importance of meteorological data for
modelling air pollution using ADMS-Urban, Meteorol. Appl., 8, 345–350, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Arthur and Vassilvitskii(2007)</label><mixed-citation>
Arthur, D. and Vassilvitskii, S.: <i>k</i>-means+ + : The advantages of careful seeding, in: Proceedings of the eighteenth annual ACM-SIAM symposium on Discrete algorithms, Society for Industrial and Applied Mathematics, 1027–1035, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Besse et al.(2018)Besse, Guillouet, and Laurent</label><mixed-citation>
Besse, P., Guillouet, B., and Laurent, B.: Wikistat 2.0: Educational Resources for Artificial Intelligence, arXiv: preprint, arXiv:1810.02688, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Breiman et al.(1984)Breiman, Friedman, Olshen, and
Stone</label><mixed-citation>
Breiman, L., Friedman, J., Olshen, R., and Stone, C.: Classification and
Regression Trees, Wadsworth, 1984.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Brilouet et al.(2017)Brilouet, Durand, and
Canut</label><mixed-citation>
Brilouet, P.-E., Durand, P., and Canut, G.: The marine atmospheric boundary
layer under strong wind conditions: Organized turbulence structure and flux
estimates by airborne measurements, J. Geophys. Res.-Atmos., 122, 2115–2130, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Brooks(2003)</label><mixed-citation>
Brooks, I. M.: Finding boundary layer top: Application of a wavelet covariance transform to lidar backscatter profiles, J. Atmos. Ocean. Tech., 20, 1092–1105, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Caicedo et al.(2017)Caicedo, Rappenglück, Lefer, Morris, Toledo, and Delgado</label><mixed-citation>
Caicedo, V., Rappenglück, B., Lefer, B., Morris, G., Toledo, D., and Delgado, R.: Comparison of aerosol lidar retrieval methods for boundary layer height detection using ceilometer aerosol backscatter data, Atmos. Meas. Tech., 10, 1609–1622, <a href="https://doi.org/10.5194/amt-10-1609-2017" target="_blank">https://doi.org/10.5194/amt-10-1609-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Caliński and Harabasz(1974)</label><mixed-citation>
Caliński, T. and Harabasz, J.: A dendrite method for cluster analysis,
Commun. Stat.-Theor. Meth., 3, 1–27, 1974.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Campbell et al.(2002)Campbell, Hlavka, Welton, Flynn, Turner,
Spinhirne, Scott, and Hwang</label><mixed-citation>
Campbell, J. R., Hlavka, D. L., Welton, E. J., Flynn, C. J., Turner, D. D.,
Spinhirne, J. D., Scott III, V. S., and Hwang, I. H.: Full-Time, Eye-Safe
Cloud and Aerosol Lidar Observation at Atmospheric Radiation Measurement
Program Sites: Instruments and Data Processing, J. Atmos. Ocean. Tech., 19, 431–442, 2002.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Cimini et al.(2013)Cimini, De Angelis, Dupont, Pal, and
Haeffelin</label><mixed-citation>
Cimini, D., De Angelis, F., Dupont, J.-C., Pal, S., and Haeffelin, M.: Mixing layer height retrievals by multichannel microwave radiometer observations, Atmos. Meas. Tech., 6, 2941–2951, <a href="https://doi.org/10.5194/amt-6-2941-2013" target="_blank">https://doi.org/10.5194/amt-6-2941-2013</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Cohn and Angevine(2000)</label><mixed-citation>
Cohn, S. A. and Angevine, W. M.: Boundary layer height and entrainment zone
thickness measured by lidars and wind-profiling radars, J. Appl. Meteorol., 39, 1233–1247, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Collaud Coen et al.(2014)Collaud Coen, Praz, Haefele, Ruffieux,
Kaufmann, and Calpini</label><mixed-citation>
Collaud Coen, M., Praz, C., Haefele, A., Ruffieux, D., Kaufmann, P., and Calpini, B.: Determination and climatology of the planetary boundary layer height above the Swiss plateau by in situ and remote sensing measurements as well as by the COSMO-2 model, Atmos. Chem. Phys., 14, 13205–13221, <a href="https://doi.org/10.5194/acp-14-13205-2014" target="_blank">https://doi.org/10.5194/acp-14-13205-2014</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Davies and Bouldin(1979)</label><mixed-citation>
Davies, D. L. and Bouldin, D. W.: A cluster separation measure, in: IEEE
transactions on pattern analysis and machine intelligence, 12–14 April 1978,
Princeton, NJ, 224–227, <a href="https://doi.org/10.1109/TPAMI.1979.4766909" target="_blank">https://doi.org/10.1109/TPAMI.1979.4766909</a>, 1979.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Davison and Hinkley(1997)</label><mixed-citation>
Davison, A. C. and Hinkley, D. V.: Bootstrap methods and their application,
Cambridge University Press, Cambridge, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>de Bruine et al.(2017)De Bruine, Apituley, Donovan, Klein Baltink,
and de Haij</label><mixed-citation>
de Bruine, M., Apituley, A., Donovan, D. P., Klein Baltink, H., and de Haij, M. J.: Pathfinder: applying graph theory to consistent tracking of daytime mixed layer height with backscatter lidar, Atmos. Meas. Tech., 10, 1893–1909, <a href="https://doi.org/10.5194/amt-10-1893-2017" target="_blank">https://doi.org/10.5194/amt-10-1893-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Desgraupes(2013)</label><mixed-citation>
Desgraupes, B.: Clustering indices, University of Paris Ouest-Lab Modal'X, available at: <a href="https://cran.biodisk.org/web/packages/clusterCrit/vignettes/clusterCrit.pdf" target="_blank"/> (last access: 7 June 2021), 34&thinsp;pp., 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Dupont et al.(2016)Dupont, Haeffelin, Badosa, Elias, Favez, Petit,
Meleux, Sciare, Crenn, and Bonne</label><mixed-citation>
Dupont, J.-C., Haeffelin, M., Badosa, J., Elias, T., Favez, O., Petit, J.,
Meleux, F., Sciare, J., Crenn, V., and Bonne, J.: Role of the boundary layer
dynamics effects on an extreme air pollution event in Paris, Atmos. Environ., 141, 571–579, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Flynn et al.(2007)Flynn, Mendoza, Zheng, and
Mathur</label><mixed-citation>
Flynn, C. J., Mendoza, A., Zheng, Y., and Mathur, S.: Novel polarization-sensitive micropulse lidar measurement technique, Opt. Express,
15, 2785–2790, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Freund and Schapire(1997)</label><mixed-citation>
Freund, Y. and Schapire, R. E.: A decision-theoretic generalization of on-line learning and an application to boosting, J. Comput. Syst. Sci., 55, 119–139, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Gamage and Hagelberg(1993)</label><mixed-citation>
Gamage, N. and Hagelberg, C.: Detection and analysis of microfronts and
associated coherent events using localized transforms, J. Atmos. Sci., 50, 750–756, 1993.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Guo et al.(2016)Guo, Miao, Zhang, Liu, Li, Zhang, He, Lou, Yan, Bian et al.</label><mixed-citation>
Guo, J., Miao, Y., Zhang, Y., Liu, H., Li, Z., Zhang, W., He, J., Lou, M., Yan, Y., Bian, L., and Zhai, P.: The climatology of planetary boundary layer height in China derived from radiosonde and reanalysis data, Atmos. Chem. Phys., 16, 13309–13319, <a href="https://doi.org/10.5194/acp-16-13309-2016" target="_blank">https://doi.org/10.5194/acp-16-13309-2016</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Haefele et al.(2016)Haefele, Hervo, Turp, Lampin J-L, Lehmann
et al.</label><mixed-citation>
Haefele, A., Hervo, M., Turp, M., Lampin, J.-L., Haeffelin, M., and Lehmann, V.: The E-PROFILE network for the operational measurement of wind and aerosol profiles over Europe, in: Proceedings of WMO Technical Conference on Meteorological and Environmental Instruments and Methods of Observation, CIMO TECO 2016, Madrid, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Haeffelin et al.(2012)Haeffelin, Angelini, Morille, Martucci, Frey,
Gobbi, Lolli, O'dowd, Sauvage, Xueref-Rémy
et al.</label><mixed-citation>
Haeffelin, M., Angelini, F., Morille, Y., Martucci, G., Frey, S., Gobbi, G.,
Lolli, S., O'dowd, C., Sauvage, L., Xueref-Rémy, I., Wastine, B., and Feist, D. G.: Evaluation of mixing-height retrievals from automatic profiling lidars and ceilometers in view of future integrated networks in Europe, Bound.-Lay. Meteorol., 143, 49–75, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Hastie et al.(2009)Hastie, Tibshirani, and
Friedman</label><mixed-citation>
Hastie, T., Tibshirani, R., and Friedman, J.: The Elements of Statistical
Learning: Data Mining, Inference, and Prediction, Springer Science &amp; Business Media, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Hayden et al.(1997)Hayden, Anlauf, Hoff, Strapp, Bottenheim, Wiebe,
Froude, Martin, Steyn, and McKendry</label><mixed-citation>
Hayden, K., Anlauf, K., Hoff, R., Strapp, J., Bottenheim, J., Wiebe, H.,
Froude, F., Martin, J., Steyn, D., and McKendry, I.: The vertical chemical
and meteorological structure of the boundary layer in the Lower Fraser Valley
during Pacific'93, Atmos. Environ., 31, 2089–2105, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Hennemuth and Lammert(2006)</label><mixed-citation>
Hennemuth, B. and Lammert, A.: Determination of the atmospheric boundary layer height from radiosonde and lidar backscatter, Bound.-Lay. Meteorol., 120, 181–200, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Herman and Usher(2017)</label><mixed-citation>
Herman, J. and Usher, W.: SALib: An open-source Python library for
Sensitivity Analysis, J. Open Source Softw., 2, 97, <a href="https://doi.org/10.21105/joss.00097" target="_blank">https://doi.org/10.21105/joss.00097</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Hintze and Nelson(1998)</label><mixed-citation>
Hintze, J. L. and Nelson, R. D.: Violin Plots: A Box Plot-Density Trace
Synergism, Am. Statist., 52, 181–184, 1998.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Iooss and Lemaître(2015)</label><mixed-citation>
Iooss, B. and Lemaître, P.: A review on global sensitivity analysis
methods, in: Uncertainty management in simulation-optimization of complex
systems, Springer, 101–122, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Jain et al.(1999)Jain, Murty, and Flynn</label><mixed-citation>
Jain, A. K., Murty, M. N., and Flynn, P. J.: Data clustering: a review, ACM
Comput. Surv., 31, 264–323, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Kotthaus and Grimmond(2018)</label><mixed-citation>
Kotthaus, S. and Grimmond, C. S. B.: Atmospheric boundary-layer characteristics from ceilometer measurements. Part 1: A new method to track mixed layer height and classify clouds, Q. J. Roy. Meteorol. Soc., 144, 1525–1538, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Krizhevsky et al.(2012)Krizhevsky, Sutskever, and
Hinton</label><mixed-citation>
Krizhevsky, A., Sutskever, I., and Hinton, G. E.: Imagenet classification with deep convolutional neural networks, in: Advances in neural information
processing systems, Curran Associates, Inc., 1097–1105, available at: <a href="https://proceedings.neurips.cc/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf" target="_blank"/>
(last access: 7 June 2021), 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>LeCun et al.(2015)LeCun, Bengio, and Hinton</label><mixed-citation>
LeCun, Y., Bengio, Y., and Hinton, G.: Deep learning, Nature, 521, 436–444,
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Lenschow et al.(2012)Lenschow, Lothon, Mayor, Sullivan, and
Canut</label><mixed-citation>
Lenschow, D. H., Lothon, M., Mayor, S. D., Sullivan, P. P., and Canut, G.: A
comparison of higher-order vertical velocity moments in the convective boundary layer from lidar with in situ measurements and large-eddy
simulation, Bound.-Lay. Meteorol., 143, 107–123, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Melfi et al.(1985)Melfi, Spinhirne, Chou, and Palm</label><mixed-citation>
Melfi, S., Spinhirne, J., Chou, S., and Palm, S.: Lidar observations of
vertically organized convection in the planetary boundary layer over the
ocean, J. Clim. Appl. Meteorol., 24, 806–821, 1985.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Menut et al.(1999)Menut, Flamant, Pelon, and
Flamant</label><mixed-citation>
Menut, L., Flamant, C., Pelon, J., and Flamant, P. H.: Urban boundary-layer
height determination from lidar measurements over the Paris area, Appl. Optics, 38, 945–954, 1999.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Mohan et al.(2011)Mohan, Bhati, Sreenivas, and
Marrapu</label><mixed-citation>
Mohan, M., Bhati, S., Sreenivas, A., and Marrapu, P.: Performance evaluation of AERMOD and ADMS-urban for total suspended particulate matter concentrations in megacity Delhi, Aerosol Air Qual. Res., 11, 883–894, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Morille et al.(2007)Morille, Haeffelin, Drobinski, and
Pelon</label><mixed-citation>
Morille, Y., Haeffelin, M., Drobinski, P., and Pelon, J.: STRAT: An automated
algorithm to retrieve the vertical structure of the atmosphere from
single-channel lidar data, J. Atmos. Ocean. Tech., 24, 761–775, 2007.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Pedregosa et al.(2011)Pedregosa, Varoquaux, Gramfort, Michel,
Thirion, Grisel, Blondel, Prettenhofer, Weiss, Dubourg, Vanderplas, Passos,
Cournapeau, Brucher, Perrot, and Duchesnay</label><mixed-citation>
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,
O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J.,
Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E.:
Scikit-learn: Machine Learning in Python, J. Mach. Learn. Res., 12, 2825–2830, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Pollard(1981)</label><mixed-citation>
Pollard, D.: Strong consistency of <i>k</i>-means clustering, Ann. Statist., 9, 135–140, 1981.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Poltera et al.(2017)Poltera, Martucci, Collaud Coen, Hervo,
Emmenegger, Henne, Brunner, and Haefele</label><mixed-citation>
Poltera, Y., Martucci, G., Collaud Coen, M., Hervo, M., Emmenegger, L., Henne, S., Brunner, D., and Haefele, A.: PathfinderTURB: an automatic boundary layer algorithm. Development, validation and application to study the impact on in situ measurements at the Jungfraujoch, Atmos. Chem. Phys., 17, 10051–10070, <a href="https://doi.org/10.5194/acp-17-10051-2017" target="_blank">https://doi.org/10.5194/acp-17-10051-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Rieutord(2017)</label><mixed-citation>
Rieutord, T.: Sensitivity analysis of a filtering algorithm for wind lidar
measurements, PhD thesis, Institut National Polytechnique de Toulouse, Toulouse, available at: <a href="https://oatao.univ-toulouse.fr/19457/" target="_blank"/> (last access: 7 June 2021), 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Rieutord et al.(2014)Rieutord, Brewer, and
Hardesty</label><mixed-citation>
Rieutord, T., Brewer, W. A., and Hardesty, R. M.: Automatic detection of
boundary layer height using Doppler lidar measurements, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Rieutord et al.(2021)</label><mixed-citation>
Rieutord, T., Aubert, S., and Machado, T.: KABL, Github, available at: <a href="https://github.com/ThomasRieutord/kabl" target="_blank"/>, last access: 3 June 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Rousseeuw(1987)</label><mixed-citation>
Rousseeuw, P. J.: Silhouettes: a graphical aid to the interpretation and
validation of cluster analysis, J. Comput. Appl. Math., 20, 53–65, 1987.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Schapire(2013)</label><mixed-citation>
Schapire, R. E.: Explaining adaboost, in: Empirical inference, Springer, 37–52, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Seibert et al.(2000)Seibert, Beyrich, Gryning, Joffre, Rasmussen, and Tercier</label><mixed-citation>
Seibert, P., Beyrich, F., Gryning, S.-E., Joffre, S., Rasmussen, A., and
Tercier, P.: Review and intercomparison of operational methods for the
determination of the mixing height, Atmos. Environ., 34, 1001–1027, 2000.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Seidel et al.(2010)Seidel, Ao, and Li</label><mixed-citation>
Seidel, D. J., Ao, C. O., and Li, K.: Estimating climatological planetary
boundary layer heights from radiosonde observations: Comparison of methods
and uncertainty analysis, J. Geophys. Res.-Atmos., 115, D16113, <a href="https://doi.org/10.1029/2009JD013680" target="_blank">https://doi.org/10.1029/2009JD013680</a>, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Seidel et al.(2012)Seidel, Zhang, Beljaars, Golaz, Jacobson, and
Medeiros</label><mixed-citation>
Seidel, D. J., Zhang, Y., Beljaars, A., Golaz, J.-C., Jacobson, A. R., and
Medeiros, B.: Climatology of the planetary boundary layer over the continental United States and Europe, J. Geophys. Res.-Atmos., 117, D17106, <a href="https://doi.org/10.1029/2012JD018143" target="_blank">https://doi.org/10.1029/2012JD018143</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Seity et al.(2011)Seity, Brousseau, Malardel, Hello, Bénard,
Bouttier, Lac, and Masson</label><mixed-citation>
Seity, Y., Brousseau, P., Malardel, S., Hello, G., Bénard, P., Bouttier,
F., Lac, C., and Masson, V.: The AROME-France convective-scale operational
model, Mon. Weather Rev., 139, 976–991, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Selim and Ismail(1984)</label><mixed-citation>
Selim, S. Z. and Ismail, M. A.: <i>K</i>-means-type algorithms: a generalized
convergence theorem and characterization of local optimality, IEEE T. Pattern Anal. Mach. Intel., PAMI-6, 81–87, <a href="https://doi.org/10.1109/TPAMI.1984.4767478" target="_blank">https://doi.org/10.1109/TPAMI.1984.4767478</a>, 1984.

</mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Senff et al.(1996)Senff, Bösenberg, Peters, and
Schaberl</label><mixed-citation>
Senff, C., Bösenberg, J., Peters, G., and Schaberl, T.: Remote sensing of
turbulent ozone fluxes and the ozone budget in the convective boundary layer
with DIAL and Radar-RASS: A case study, Contrib. Atmos. Phys., 69, 161–176, 1996.
</mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Sobol(2001)</label><mixed-citation>
Sobol, I. M.: Global sensitivity indices for nonlinear mathematical models and their Monte Carlo estimates, Math. Comput. Simul., 55, 271–280, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Stull(1988)</label><mixed-citation>
Stull, R. B.: An introduction to boundary layer meteorology, in: vol. 13, Springer, 1988.
</mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Tibshirani et al.(2001)Tibshirani, Walther, and
Hastie</label><mixed-citation>
Tibshirani, R., Walther, G., and Hastie, T.: Estimating the number of clusters in a data set via the gap statistic, J. Roy. Stat. Soc. B, 63, 411–423, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Toledo et al.(2014)Toledo, Córdoba-Jabonero, and
Gil-Ojeda</label><mixed-citation>
Toledo, D., Córdoba-Jabonero, C., and Gil-Ojeda, M.: Cluster analysis: A
new approach applied to lidar measurements for atmospheric boundary layer
height estimation, J. Atmos. Ocean. Tech., 31, 422–436, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Toledo et al.(2017)Toledo, Córdoba-Jabonero, Adame, De La Morena, and Gil-Ojeda</label><mixed-citation>
Toledo, D., Córdoba-Jabonero, C., Adame, J. A., De La Morena, B., and
Gil-Ojeda, M.: Estimation of the atmospheric boundary layer height during
different atmospheric conditions: a comparison on reliability of several
methods applied to lidar measurements, Int. J. Remote Sens., 38, 3203–3218, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Ware et al.(2016)Ware, Kort, DeCola, and Duren</label><mixed-citation>
Ware, J., Kort, E. A., DeCola, P., and Duren, R.: Aerosol lidar observations of atmospheric mixing in Los Angeles: Climatology and implications for
greenhouse gas observations, J. Geophys. Res.-Atmos., 121, 9862–9878, 2016.
</mixed-citation></ref-html>--></article>
