<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">AMT</journal-id><journal-title-group>
    <journal-title>Atmospheric Measurement Techniques</journal-title>
    <abbrev-journal-title abbrev-type="publisher">AMT</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Atmos. Meas. Tech.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1867-8548</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/amt-14-5637-2021</article-id><title-group><article-title>Machine learning calibration of low-cost NO<inline-formula><mml:math id="M1" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M2" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors: non-linear algorithms and their impact on site transferability</article-title><alt-title>Machine learning calibration</alt-title>
      </title-group><?xmltex \runningtitle{Machine learning calibration}?><?xmltex \runningauthor{P. Nowack et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2 aff3 aff4">
          <name><surname>Nowack</surname><given-names>Peer</given-names></name>
          <email>p.nowack@uea.ac.uk</email>
        <ext-link>https://orcid.org/0000-0003-4588-7832</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff5">
          <name><surname>Konstantinovskiy</surname><given-names>Lev</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff5">
          <name><surname>Gardiner</surname><given-names>Hannah</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff5">
          <name><surname>Cant</surname><given-names>John</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Grantham Institute – Climate Change and the Environment, Imperial College London, London SW7 2AZ, UK</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Physics, Imperial College London, London SW7 2AZ, UK</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>Data Science Institute, Imperial College London, London SW7 2AZ, UK</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>Climatic Research Unit, School of Environmental Sciences, University of East Anglia, Norwich NR4 7TJ, UK</institution>
        </aff>
        <aff id="aff5"><label>5</label><institution>AirPublic Ltd, London, UK</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Peer Nowack (p.nowack@uea.ac.uk)</corresp></author-notes><pub-date><day>18</day><month>August</month><year>2021</year></pub-date>
      
      <volume>14</volume>
      <issue>8</issue>
      <fpage>5637</fpage><lpage>5655</lpage>
      <history>
        <date date-type="received"><day>30</day><month>November</month><year>2020</year></date>
           <date date-type="rev-request"><day>22</day><month>December</month><year>2020</year></date>
           <date date-type="rev-recd"><day>24</day><month>June</month><year>2021</year></date>
           <date date-type="accepted"><day>10</day><month>July</month><year>2021</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2021 </copyright-statement>
        <copyright-year>2021</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://amt.copernicus.org/articles/.html">This article is available from https://amt.copernicus.org/articles/.html</self-uri><self-uri xlink:href="https://amt.copernicus.org/articles/.pdf">The full text article is available as a PDF file from https://amt.copernicus.org/articles/.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e154">Low-cost air pollution sensors often fail to attain sufficient performance compared with state-of-the-art measurement stations, and they typically require expensive laboratory-based calibration procedures. A repeatedly proposed strategy to overcome these limitations is calibration through co-location with public measurement stations. Here we test the idea of using machine learning algorithms for such calibration tasks using hourly-averaged co-location data for nitrogen dioxide (NO<inline-formula><mml:math id="M3" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>) and particulate matter of particle sizes smaller than 10 <inline-formula><mml:math id="M4" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> (PM<inline-formula><mml:math id="M5" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>) at three different locations in the urban area of London, UK. We compare the performance of ridge regression, a linear statistical learning algorithm, to two non-linear algorithms in the form of random forest regression (RFR) and Gaussian process regression (GPR). We further benchmark the performance of all three machine learning methods relative to the more common multiple linear regression (MLR). We obtain very good out-of-sample <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores (coefficient of determination) <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula>, frequently exceeding 0.8, for the machine learning calibrated low-cost sensors. In contrast, the performance of MLR is more dependent on random variations in the sensor hardware and co-located signals, and it is also more sensitive to the length of the co-location period. We find that, subject to certain conditions, GPR is typically the best-performing method in our calibration setting, followed by ridge regression and RFR. We also highlight several key limitations of the machine learning methods, which will be crucial to consider in any co-location calibration. In particular, all methods are fundamentally limited in how well they can reproduce pollution levels that lie outside those encountered at training stage. We find, however, that the linear ridge regression outperforms the non-linear methods in extrapolation settings. GPR can allow for a small degree of extrapolation, whereas RFR can only predict values within the training range. This algorithm-dependent ability to extrapolate is one of the key limiting factors when the calibrated sensors are deployed away from the co-location site itself. Consequently, we find that ridge regression is often performing as good as or even better than GPR after sensor relocation. Our results highlight the potential of co-location approaches paired with machine learning calibration techniques to reduce costs of air pollution measurements, subject to careful consideration of the co-location training conditions, the choice of calibration variables and the features of the calibration algorithm.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <?pagebreak page5638?><p id="d1e215">Air pollutants such as nitrogen dioxide (NO<inline-formula><mml:math id="M8" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>) and particulate matter (PM) have harmful impacts on human health, the ecosystem and public infrastructure <xref ref-type="bibr" rid="bib1.bibx13" id="paren.1"/>. Moving towards reliable and high-density air pollution measurements is consequently of prime importance. The development of new low-cost sensors, hand in hand with novel sensor calibration methods, has been at the forefront of current research efforts in this discipline <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx30 bib1.bibx23 bib1.bibx49 bib1.bibx42 bib1.bibx47 bib1.bibx11 bib1.bibx43" id="paren.2"><named-content content-type="pre">e.g.</named-content></xref>. Here we present insights from a case study using low-cost air pollution sensors for measurements at three separate locations in the urban area of London, UK. Our focus is on testing the advantages and disadvantages of machine learning calibration techniques for low-cost NO<inline-formula><mml:math id="M9" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M10" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors. The principal idea is to calibrate the sensors through co-location with established high-performance air pollution measurement stations (Fig. <xref ref-type="fig" rid="Ch1.F1"/>). Such calibration techniques, if successful, could complement more expensive laboratory-based calibration approaches, thereby further reducing the costs of the overall measurement process <xref ref-type="bibr" rid="bib1.bibx45 bib1.bibx49 bib1.bibx31" id="paren.3"><named-content content-type="pre">e.g.</named-content></xref>. For the sensor calibration, we compare three machine learning regression techniques in the form of ridge regression, random forest regression (RFR) and Gaussian process regression (GPR), and we contrast the results to those obtained with standard multiple linear regression (MLR). RFR has been studied in the context of NO<inline-formula><mml:math id="M11" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> co-location calibrations before, with very promising results <xref ref-type="bibr" rid="bib1.bibx49" id="paren.4"/>. Equally for NO<inline-formula><mml:math id="M12" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> (but not for PM<inline-formula><mml:math id="M13" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>) different linear versions of GPR have been tested by <xref ref-type="bibr" rid="bib1.bibx8" id="text.5"/> and <xref ref-type="bibr" rid="bib1.bibx25" id="text.6"/>. To the best of our knowledge, we are the first to test ridge regression both for NO<inline-formula><mml:math id="M14" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M15" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and GPR for PM<inline-formula><mml:math id="M16" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>. Finally, we also investigate well-known issues concerning site transferability <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx14 bib1.bibx16 bib1.bibx25" id="paren.7"/>, i.e. if a calibration through co-location at one location gives rise to reliable measurements at a different location.</p>
      <p id="d1e328">A key motivation for our study is the potential of low-cost sensors (costs of the order of GBP 10 to GBP 100) to transform the level of availability of air pollution measurements. Installation costs of state-of-the-art measurement stations typically range between GBP 10 000 and GBP 100 000 per site, and those already high costs are further exacerbated through subsequent maintenance and calibration requirements <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx22 bib1.bibx6" id="paren.8"/>. Lower measurement costs would allow for the deployment of denser air pollution sensor networks and for portable devices possibly even at the exposure level of individuals <xref ref-type="bibr" rid="bib1.bibx29" id="paren.9"/>. A central complication is the sensitivity of sensors to environmental conditions such as temperature and relative humidity <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx45 bib1.bibx20 bib1.bibx22 bib1.bibx46 bib1.bibx6" id="paren.10"/> or to cross-sensitivities with other gases (e.g. nitrogen oxide), which can significantly impede their measurement performance <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx37 bib1.bibx38 bib1.bibx23 bib1.bibx24" id="paren.11"/>. Low-cost sensors thus require, in the same way as many other measurement devices, sophisticated calibration techniques. Machine learning regressions have seen increased use in this context due to their ability to calibrate for many simultaneous, non-linear dependencies. These dependencies, in turn, can for example be assessed in relatively expensive laboratory settings.  However, even laboratory calibrations do not always perform well in the field <xref ref-type="bibr" rid="bib1.bibx6 bib1.bibx49" id="paren.12"/>. Here, we instead evaluate the performance of low-cost sensor calibrations based on co-location measurements with established reference stations <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx45 bib1.bibx12 bib1.bibx22 bib1.bibx7 bib1.bibx16 bib1.bibx4 bib1.bibx5 bib1.bibx8 bib1.bibx9 bib1.bibx49 bib1.bibx5 bib1.bibx31 bib1.bibx25 bib1.bibx26" id="paren.13"><named-content content-type="pre">e.g.</named-content></xref>. If sufficiently successful, these methods could help to substantially reduce the overall costs and simplify the process of calibrating low-cost sensors.</p>
      <p id="d1e352">Another challenge in relation to co-location calibration procedures is “site transferability”. This term refers to the measurement performance implications of moving a calibrated device from one location (where the calibration was conducted) to another location. Some significant performance losses after site transfers have been reported <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx4 bib1.bibx17 bib1.bibx48" id="paren.14"><named-content content-type="pre">e.g.</named-content></xref>, with reasons typically not being straightforward to assign. A key driver might be that often devices are calibrated in an environment not representative of situations found in later measurement locations. As we discuss in greater detail below, for machine-learning-based calibrations this behaviour can, to a degree, be fairly intuitively explained by the fact that they do not tend to perform well when extrapolating beyond their training domain. As we will show, this issue can easily occur in situations where already calibrated sensors have to measure pollution levels well beyond the range of values encountered in their training environment.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e363">Sketch of the co-location calibration methodology. We co-locate several low-cost sensors for PM (for various particle sizes) and NO<inline-formula><mml:math id="M17" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> with higher-cost reference measurement stations for PM<inline-formula><mml:math id="M18" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and NO<inline-formula><mml:math id="M19" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>. The low-cost sensors also measure relative humidity and temperature as key environmental variables that can interfere with the sensor signals, and for NO<inline-formula><mml:math id="M20" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> calibrations, we further include nitrogen oxide (NO) sensors. We formulate the calibration task as a regression problem in which the low-cost sensor signals and the environmental variables are the predictors (<inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and the reference station signal the predictand (<inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), both measured at the same location <inline-formula><mml:math id="M23" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>. The time resolution is set to hourly averages to match publicly available reference data. We train separate calibration functions for each NO<inline-formula><mml:math id="M24" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M25" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensor, and we compare three different machine learning algorithms (ridge, random forest and Gaussian process regressions) with multiple linear regression in terms of their respective calibration performances. The performance is evaluated on out-of-sample test data, i.e. on data not used during training. Once trained and cross-validated, we use these calibration functions to predict PM<inline-formula><mml:math id="M26" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> and NO<inline-formula><mml:math id="M27" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations given new low-cost measurements <inline-formula><mml:math id="M28" display="inline"><mml:mi>X</mml:mi></mml:math></inline-formula>, either measured at the same location <inline-formula><mml:math id="M29" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> or at a new location <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msup><mml:mi>r</mml:mi><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula>. The latter location is to test the feasibility and impacts of changing measurement sites post-calibration. The time series (right) are for illustration purposes only.</p></caption>
        <?xmltex \igopts{width=455.244094pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f01.png"/>

      </fig>

      <p id="d1e500">We highlight that, in particular concerning the performance of low-cost PM<inline-formula><mml:math id="M31" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors, a huge gap in the scientific literature has been identified regarding issues related to co-location calibrations <xref ref-type="bibr" rid="bib1.bibx38" id="paren.15"/>. We therefore expect that our study will provide novel insights into the effects of different calibration techniques on sensor performances, and a data sample that other measurement studies from academia and industry can compare their results against. We will mainly use the <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score (coefficient of determination) and root mean squared error (RMSE) as metrics to evaluate our calibration results, which are widely used and should thus facilitate intercomparisons. To provide a reference for calibration results perceived as “good” for PM<inline-formula><mml:math id="M33" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, we point towards a sensor comparison by <xref ref-type="bibr" rid="bib1.bibx38" id="text.16"/>, who found that low-cost sensors generally displayed moderate to excellent linearity (<inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>) across various calibration settings. The sensors typically perform particularly well (<inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>) when tested in idealized laboratory conditions. However, their performance is generally lower in field deployments <xref ref-type="bibr" rid="bib1.bibx23" id="paren.17"><named-content content-type="pre">see also</named-content></xref>. For PM<inline-formula><mml:math id="M36" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, this performance deterioration was, inter alia, attributed to changing conditions of particle composition, particle sizes and environmental factors such as humidity and temperature, which are thus important factors to account for in our calibrations.</p>
      <?pagebreak page5639?><p id="d1e583">The structure of our paper is as follows. In Sect. <xref ref-type="sec" rid="Ch1.S2"/>, we introduce the low-cost sensor hardware used, the reference measurement sources, the three measurement site characteristics and measurement periods, the four calibration regression methods, and the calibration settings (e.g. measured signals used) for NO<inline-formula><mml:math id="M37" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M38" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>. In Sect. <xref ref-type="sec" rid="Ch1.S3"/>, we first introduce multi-sensor calibration results for NO<inline-formula><mml:math id="M39" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> at a single site, depending on the sensor signals included in the calibrations and the number of training samples used to train the regressions. This is followed by a discussion of single-site PM<inline-formula><mml:math id="M40" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> calibration results before we test the feasibility and challenges of site transfers. We discuss our results and draw conclusions in Sect. 4.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methods and data</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Sensor hardware</title>
      <p id="d1e642">Depending on the measurement location, we deployed one set or several sets of air pollution sensors, and we refer to each set (provided by London-based AirPublic Ltd) as a multi-sensor “node”. Each of these nodes consists of multiple electrochemical and metal oxide sensors for PM and NO<inline-formula><mml:math id="M41" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, as well as sensors for environmental quantities and other chemical species known for potential interference with their sensor signals (required for calibration). Each node thus allows for simultaneous measurement of multiple air pollutants, but we will focus on individual calibrations for NO<inline-formula><mml:math id="M42" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M43" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> here, because these species were of particular interest to our own measurement campaigns. We note that other species such as PM<inline-formula><mml:math id="M44" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> are also included in the measured set of variables. We will make our low-cost sensor data available (see “Data availability” section), which will allow other users to test similar calibration procedures for other variables of interest (e.g. PM<inline-formula><mml:math id="M45" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula>, ozone). A caveat is that appropriate co-location data from higher-cost reference measurements might not always be available.</p>
      <p id="d1e690">For NO<inline-formula><mml:math id="M46" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, we incorporated three different types of sensors in our set-up for which purchasing prices differed by an order of magnitude. One aspect of our study will therefore be to evaluate the performance gained by using the more expensive (but still relatively low-cost) sensor types. Of course, our results will only be validated in the context of our specific calibration method so that more general conclusions have to be drawn with care.</p>
      <p id="d1e702">Each multi-sensor node contained the following (i.e. all nodes consist of the same types of individual sensors):
<list list-type="bullet"><list-item>
      <p id="d1e707">Two MiCS-2714 NO<inline-formula><mml:math id="M47" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensors produced by SGX Sensortech. These are the cheapest measurement devices deployed in our set with market costs of approximately GBP 5 per sensor.</p></list-item><list-item>
      <p id="d1e720">Two Plantower PMS5003T series PM sensors (PMSs), which measure particles of various size categories including PM<inline-formula><mml:math id="M48" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> based on laser scattering using Mie theory. We note that particle composition does play a role in any PM calibration process as, for example, organic materials tend to absorb a higher proportion of incident light as compared to inorganic materials <xref ref-type="bibr" rid="bib1.bibx38" id="paren.18"/>. Below we therefore effectively make the assumption that we measure and calibrate within composition-wise similar environments. By taking into account various particle size measures in the calibration, we likely do indirectly account for some aspects of composition though, because to a degree, particle sizes might be correlated with particle composition. Each PMS device also contains a temperature and relative humidity sensor, and these variables were also included in<?pagebreak page5640?> our calibrations. The minimum distinguishable particle diameter for the PMS devices is 0.3 <inline-formula><mml:math id="M49" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula>. The market cost is GBP 20 for one sensor.</p></list-item><list-item>
      <p id="d1e746">An NO2-A43F four-electrode NO<inline-formula><mml:math id="M50" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor produced by AlphaSense (market cost GBP 45).</p></list-item><list-item>
      <p id="d1e759">An NO2-B43F four-electrode NO<inline-formula><mml:math id="M51" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor produced by AlphaSense (market cost GBP 45).</p></list-item><list-item>
      <p id="d1e772">An NO-A4 four-electrode nitric oxide sensor produced by AlphaSense to calibrate against the sometimes significant interference of NO<inline-formula><mml:math id="M52" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> signals with NO (market cost GBP 45).</p></list-item><list-item>
      <p id="d1e785">An OX-A431 four-electrode oxidizing gas sensor measuring a combined signal from ozone and NO<inline-formula><mml:math id="M53" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> produced by AlphaSense. We used this signal to calibrate against possible interference of electrochemical NO<inline-formula><mml:math id="M54" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> measurements by ozone (market cost GBP 45).</p></list-item><list-item>
      <p id="d1e807">A separate temperature sensor built into the AlphaSense set. It is needed to monitor the warm-up phase of the sensors.</p></list-item></list>
In normal operation mode, each node provided measurements around every 30 s. These signals were time-averaged to hourly values for calibration against hourly public reference measurements.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Measurement sites and reference monitors</title>
      <p id="d1e819">We conducted measurements at three sites in the Greater London area during distinct multi-week periods (Table <xref ref-type="table" rid="Ch1.T1"/>). Two of the sites are located in the London Borough of Croydon, which we label CR7 and CR9 according to their UK postcodes. The third site is located in the car park of the company AlphaSense in Essex, hereafter referred to as site “CarPark” (Fig. <xref ref-type="fig" rid="Ch1.F2"/>a). At CR7, the sensor nodes were located kerbside on a medium busy street with two lanes in either direction. At CR9, the nodes were located on a traffic island in the middle of a very busy road with three lanes in either direction (Fig. <xref ref-type="fig" rid="Ch1.F2"/>b). For the two Croydon sites, reference measurements were obtained from co-located, publicly available measurements of London's Air Quality Network (LAQN; <uri>https://www.londonair.org.uk</uri>, last access: 30 April 2020​​​​​​​). In the two locations in question, the LAQN used chemiluminescence detectors for NO<inline-formula><mml:math id="M55" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and Thermo Scientific tapered element oscillating microbalance (TEOM) continuous ambient particulate monitors, with the Volatile Correction Model <xref ref-type="bibr" rid="bib1.bibx15" id="paren.19"/> for PM<inline-formula><mml:math id="M56" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> measurements at CR9. For the CarPark site, PM<inline-formula><mml:math id="M57" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> measurements were conducted with a Palas Fidas optical particle counter AFL-W07A. These CarPark reference measurements were provided at 15 min intervals. For consistency, these measurements were averaged to hourly values to match the measurement frequency of publicly available data at the other two sites.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T1" specific-use="star"><?xmltex \currentcnt{1}?><label>Table 1</label><caption><p id="d1e865">Overview of the measurement sites and the corresponding maximum co-location periods, which vary for each specific sensor node and co-location site due to practical aspects such as sensor availability and random occurrences of sensor failures. Note that reference measurements for NO<inline-formula><mml:math id="M58" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M59" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> are only available for two of the three sites each. Sensors that were co-located for at least 820 active measurement hours are identified by their sensor IDs in the last column. Further note that the only sensor used to measure at multiple sites is sensor 19, which is therefore used to test the feasibility of site transfers.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Site</oasis:entry>
         <oasis:entry colname="col2">Max co-location period</oasis:entry>
         <oasis:entry colname="col3">Reference sensors</oasis:entry>
         <oasis:entry colname="col4">Low-cost sensor IDs</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">CR7</oasis:entry>
         <oasis:entry colname="col2">22 October–5 December 2018</oasis:entry>
         <oasis:entry colname="col3">NO<inline-formula><mml:math id="M60" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> (London Air Quality Network)</oasis:entry>
         <oasis:entry colname="col4">3–7, 11, 13–24, 27, 28</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CR9</oasis:entry>
         <oasis:entry colname="col2">24 September 2019–19 January 2020</oasis:entry>
         <oasis:entry colname="col3">NO<inline-formula><mml:math id="M61" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M62" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (London Air Quality Network)</oasis:entry>
         <oasis:entry colname="col4">19</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">CarPark</oasis:entry>
         <oasis:entry colname="col2">29 January–26 April 2019</oasis:entry>
         <oasis:entry colname="col3">PM<inline-formula><mml:math id="M63" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (Palas Fidas optical particle counter AFL-W07A)</oasis:entry>
         <oasis:entry colname="col4">19, 25, 26</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e1004">Examples of the co-location set-up of the AirPublic low-cost sensor nodes with reference measurement stations at sites <bold>(a)</bold> CarPark and <bold>(b)</bold> CR9.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f02.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Co-location set-up and calibration variables</title>
      <p id="d1e1027">In total, we co-located up to 30 nodes, labelled by identifiers (IDs) 1 to 30. For our NO<inline-formula><mml:math id="M64" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> measurements, we considered the following 15 sensor signals per node to be important for the calibration process: the NO sensor (plus its baseline signal to remove noise), the NO<inline-formula><mml:math id="M65" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> <inline-formula><mml:math id="M66" display="inline"><mml:mo>+</mml:mo></mml:math></inline-formula> O<inline-formula><mml:math id="M67" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula> sensor (plus baseline), the two intermediate cost NO<inline-formula><mml:math id="M68" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensors (NO2-A43F, NO2-B43F) plus their respective baselines, the two cheaper MiCS sensors, three different temperature sensors, and two relative humidity sensors. All 15 signals can be used for calibration against the reference measurements obtained with the co-located higher-cost measurement devices. We discuss the relative importance of the different signals, e.g. the relative importance of the different NO<inline-formula><mml:math id="M69" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensors or the influence of temperature and humidity in Sect. <xref ref-type="sec" rid="Ch1.S3"/>. For the PM<inline-formula><mml:math id="M70" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> calibrations, we used two devices of the same type of low-cost PM sensor, resulting in 2<inline-formula><mml:math id="M71" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula>10 different particle measures used in the PM<inline-formula><mml:math id="M72" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> calibrations. In addition, we included the respective sensor signals for temperature and relative humidity, providing us with in total 24 calibration signals for PM<inline-formula><mml:math id="M73" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Calibration algorithms</title>
      <p id="d1e1127">We evaluate four regression calibration strategies for low-cost NO<inline-formula><mml:math id="M74" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M75" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> devices, by means of co-location of the devices with the aforementioned air quality measurement reference stations. The four different regression methods – which are multiple linear regression (MLR), ridge regression, random forest regression (RFR) and Gaussian process regression (GPR) – are introduced in detail in the following subsections. As we will show in Sect. <xref ref-type="sec" rid="Ch1.S3"/>, the relative skill of the calibration methods depends on the chemical species to be measured, sample size available for calibration, and certain user preferences. We will additionally consider the issue of site transferability for sensor node 19, including its dependence on the calibration algorithm used. We note that we do not include the manufacturer calibration of the low-cost sensors in our comparison here mainly because we found that this method, which is a simple linear regression based on certain laboratory measurement relationships, provided us with negative <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores when compared with reference sensors in the field. This result is in line with other studies that reported differences between sensor performances in the field and under laboratory calibrations <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx23 bib1.bibx38" id="paren.20"><named-content content-type="pre">see e.g.</named-content></xref>.</p>
<sec id="Ch1.S2.SS4.SSS1">
  <label>2.4.1</label><title>Ridge and multiple linear regression</title>
      <p id="d1e1173">Ridge regression is a linear least squares regression augmented by L<inline-formula><mml:math id="M77" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> regularization to address the bias-variance trade-off <xref ref-type="bibr" rid="bib1.bibx18 bib1.bibx19 bib1.bibx33 bib1.bibx34" id="paren.21"/>. Using statistical cross-validation, the regression fit is optimized by minimizing the cost function
              <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M78" display="block"><mml:mrow><mml:msub><mml:mi>J</mml:mi><mml:mtext>Ridge</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:msubsup><mml:mi>c</mml:mi><mml:mi>j</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></disp-formula>
            over <inline-formula><mml:math id="M79" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> hourly reference measurements of pollutant <inline-formula><mml:math id="M80" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> (i.e. NO<inline-formula><mml:math id="M81" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M82" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>); <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> represents <inline-formula><mml:math id="M84" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> non-calibrated measurement signals from the low-cost sensors, representing signals for the pollutant itself as well as signals recorded for environmental variables (temperature, humidity) and other chemical species that might cause interference with the signal in question. The cost function (Eq. <xref ref-type="disp-formula" rid="Ch1.E1"/>) determines the optimization goal. Its first term is the ordinary least squares regression error, and the second term puts a penalty on too large regression<?pagebreak page5642?> coefficients and thus avoids overfitting in high-dimensional settings. Smaller (larger) values of the regularization coefficient <inline-formula><mml:math id="M85" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> put weaker (stronger) constraints on the size of the coefficients, thereby favouring overfitting (high bias). We find the value for <inline-formula><mml:math id="M86" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> through fivefold cross-validation; i.e. each data set is split into five ordered time slices and <inline-formula><mml:math id="M87" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> optimized by fitting regressions for large ranges of <inline-formula><mml:math id="M88" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> values on four of the slices at a time, and then the best <inline-formula><mml:math id="M89" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> is found by evaluating the out-of-sample prediction error on each corresponding remaining slice using the <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score. Each slice is used once for the evaluation step. Before the training procedure, all signals are scaled to unit variance and zero mean so as to ensure that all signals are weighted equally in the regression optimization, which we explain in more detail at the end of this section. Through the constraint on the regression slopes, ridge regression can handle settings with many predictors, here calibration variables, even in the context of strong collinearity in those predictors <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx33 bib1.bibx34" id="paren.22"/>. The resulting linear regression function <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mtext>Ridge</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>,
              <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M92" display="block"><mml:mrow><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mtext>Ridge</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>p</mml:mi></mml:munderover><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            provides estimates for pollutant mixing ratios <inline-formula><mml:math id="M93" display="inline"><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> at any time <inline-formula><mml:math id="M94" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, i.e. a calibrated low-cost sensor signal, based on new sensor readings <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mtext>Ridge</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> represents a calibration function because it is not just based on a regression of the pollutant signal itself against the reference but also on multiple simultaneous predictors, including those representing known interfering factors.</p>
      <p id="d1e1504">Multiple linear regression (MLR) is the simple non-regularized case of ridge regression, i.e. where <inline-formula><mml:math id="M97" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> is set to nil. MLR is therefore a good benchmark to evaluate the importance of regularization and, when compared to RFR and GPR below, of non-linearity in the relationships. As MLR does not regularize its coefficients, it is expected to increasingly lose performance in settings with many (non-linear) calibration relationships. This loss of MLR performance in high-dimensional regression spaces is related to the “curse of dimensionality” in machine learning, which expresses the observation that one requires an exponentially increasing number of samples to constrain the regression coefficients as the number of predictors is increased linearly <xref ref-type="bibr" rid="bib1.bibx1" id="paren.23"/>. We will illustrate this phenomenon for the case of our NO<inline-formula><mml:math id="M98" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor calibrations below.</p>
      <p id="d1e1526">Finally, we note that for ridge regression, as also for GPR described below, the predictors <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> must be normalized to a common range. For ridge, this is straightforward to understand as the regression coefficients, once the predictors are normalized, provide direct measures of the importance of each predictor for the overall pollutant signal. If not normalized, the coefficients will additionally weight the relative magnitude of predictor signals, which can differ by orders of magnitude (e.g. temperature at around 273 K but a measurement signal for a trace gas of the order of 0.5 amplifier units). As a result, the predictors would be penalized differently through the same <inline-formula><mml:math id="M100" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>), which could mean that certain predictors are effectively not considered in the regressions. Here, we normalize all predictors in all regressions to zero mean and unit standard deviation according to the samples included in each training data set.</p>
</sec>
<sec id="Ch1.S2.SS4.SSS2">
  <label>2.4.2</label><title>Random forest regression</title>
      <p id="d1e1557">Random forest regression (RFR) is one of the most widely used non-linear machine learning algorithms <xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx2" id="paren.24"/>, and it has already found applications in air pollution sensor calibration as well as in other aspects of atmospheric chemistry <xref ref-type="bibr" rid="bib1.bibx21 bib1.bibx33 bib1.bibx34 bib1.bibx44 bib1.bibx49 bib1.bibx25" id="paren.25"/>. It follows the idea of ensemble learning where multiple machine learning models together make more reliable predictions than the individual models. Each RFR object consists of a collection (i.e. ensemble) of graphical tree models, which split training data by learning decision rules (Fig. <xref ref-type="fig" rid="Ch1.F3"/>). Each of these decision trees consists of a sequence of nodes, which branch into multiple tree levels until the end of the tree (the “leaf” level) is reached. Each leaf node contains at least one or several samples from the training data. The average of these samples is the prediction of each tree for any measurement of predictors <inline-formula><mml:math id="M101" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> defining a new traversion of the tree to the given leaf node. In contrast to ridge regression, there is more than one tunable hyperparameter to address overfitting. One of these hyperparameters is the maximum tree depth, i.e. the maximal number of levels within each tree, as deeper trees allow for a more detailed grouping of samples. Similarly, one can set the minimum number of samples in any leaf node. Once this minimum number is reached, the node is not further split into children nodes. Both smaller tree depth and a greater number of minimum samples in leaf nodes mitigate overfitting to the training data. Other important settings are the optimization function used to define the decision rules and, for example, the number of estimators included in an ensemble, i.e. the number of trees in the forest.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e1577">Sketch of a random forest regressor. Each random forest consists of an ensemble of <inline-formula><mml:math id="M102" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> trees. For visualization purposes, the trees shown here have only four levels of equal depth, but more complex structures can be learnt. The lowest level contains the leaf nodes. Note that in real examples, branches can have different depths; i.e. the leaf nodes can occur at different levels of the tree hierarchy; see, for example, Fig. 2 in <xref ref-type="bibr" rid="bib1.bibx49" id="text.26"/>. Once the decision rules for each node and tree are learnt from training data, each tree can be presented with new sensor readings <inline-formula><mml:math id="M103" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> at a time <inline-formula><mml:math id="M104" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> to predict pollutant concentration <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The decision rules depend, inter alia, on the tree structure and random sampling through bootstrapping, which we optimize through fivefold cross-validation. Based on the values <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, each set of predictors follows routes through the trees. The training samples collected in the corresponding leaf node define the tree-specific prediction for <inline-formula><mml:math id="M107" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>. By averaging <inline-formula><mml:math id="M108" display="inline"><mml:mi>K</mml:mi></mml:math></inline-formula> tree-wise predictions, we combat tree-specific overfitting and finally obtain a more regularized random forest prediction <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p></caption>
            <?xmltex \igopts{width=426.791339pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f03.png"/>

          </fig>

      <p id="d1e1661">The RFR training process tunes the parameter thresholds for each binary decision tree node. By introducing randomness, e.g. by selecting a subset of samples from the training data set (bootstrapping), each tree provides a somewhat different data representation. This random element is used to obtain a better, averaged prediction over all trees in the ensemble, which is less prone to overfitting than individual regression trees. We here cross-validated the scikit-learn implementation of RFR <xref ref-type="bibr" rid="bib1.bibx36" id="paren.27"/> over problem-specific ranges for the minimum number of samples required to define a split and the minimum number of samples to define a leaf node. The implementation uses an optimized version of the Classification And Regression Tree (CART) algorithm, which constructs binary decision trees using the predictor and threshold that yields the largest information gain<?pagebreak page5643?> for a split at each node. The mean squared error of samples relative to their node prediction (mean) serves as optimization criterion so as to measure the quality of a split for a given possible threshold during training. Here we consider all features when defining any new best split of the data at nodes. By increasing the number of trees in the ensemble, the RFR generalization error converges towards a lower limit.  We here set the number of trees in all regression tasks to 200 as a compromise between model convergence and computational complexity <xref ref-type="bibr" rid="bib1.bibx2" id="paren.28"/>.</p>
</sec>
<sec id="Ch1.S2.SS4.SSS3">
  <label>2.4.3</label><title>Gaussian process regression</title>
      <p id="d1e1678">Gaussian process regression (GPR) is a widely used Bayesian machine learning method to estimate non-linear dependencies <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx36 bib1.bibx22 bib1.bibx8 bib1.bibx41 bib1.bibx25 bib1.bibx35 bib1.bibx27" id="paren.29"/>. In GPR, the aim is to find a distribution over possible functions that fit the data. We first define a prior distribution of possible functions that is updated according to the data using Bayes' theorem, which provides us with a posterior distribution over possible functions. The prior distribution is a Gaussian process (GP),
              <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M110" display="block"><mml:mrow><mml:mi>Y</mml:mi><mml:mo>∼</mml:mo><mml:mtext>GP</mml:mtext><mml:mfenced close=")" open="("><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            with mean <inline-formula><mml:math id="M111" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> and a covariance function or kernel <inline-formula><mml:math id="M112" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>, which describes the covariance between any two points <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. We here “standard-scale” (i.e. centre) our data so that <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, meaning our GP is entirely defined by the covariance function. Being a kernel method, the performance of GPR depends strongly on the kernel (covariance function) design as it determines the shape of the prior and posterior distributions of the Gaussian process and in particular the characteristics of the function we are able to learn from the data. Owing to the time-varying, continuous but also oscillating nature of air pollution sensor signals, we here use a sum kernel of a radial basis function (RBF) kernel, a white noise kernel, a Matérn kernel and a “Dot-Product” kernel. The RBF kernel, also known as squared exponential kernel, is defined as
              <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M116" display="block"><mml:mrow><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>-</mml:mo><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:msup><mml:mi>l</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
            It is parameterized by a length scale <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:mi>l</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M118" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> is the Euclidean distance. The length scale determines the scale of variation in the data, and it is learnt during the Bayesian update; i.e. for a shorter length scale the function is more flexible. However, it also determines the extrapolation scale of the function, meaning that any extrapolation beyond the length scale is probably unreliable. RBF kernels are particularly helpful to model smooth variations in the data. The Matérn kernel is defined by
              <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M119" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="normal">Γ</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>)</mml:mo><mml:msup><mml:mn mathvariant="normal">2</mml:mn><mml:mrow><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">ν</mml:mi></mml:mrow></mml:msqrt><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mi mathvariant="italic">ν</mml:mi></mml:msup><mml:msub><mml:mi>K</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">ν</mml:mi></mml:mrow></mml:msqrt><mml:mi>l</mml:mi></mml:mfrac></mml:mstyle><mml:mi>d</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
            where <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mi mathvariant="italic">ν</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a modified Bessel function and <inline-formula><mml:math id="M121" display="inline"><mml:mi mathvariant="normal">Γ</mml:mi></mml:math></inline-formula> the gamma function <xref ref-type="bibr" rid="bib1.bibx36" id="paren.30"/>. We here choose <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mi mathvariant="italic">ν</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.5</mml:mn></mml:mrow></mml:math></inline-formula> as the default setting for the kernel, which determines the smoothness of the function. Overall, the Matérn kernel is useful to model less-smooth variations in the data than the RBF kernel. The Dot-Product kernel is parameterized by a hyperparameter <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">0</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:math></inline-formula>,
              <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M124" display="block"><mml:mrow><mml:mi>k</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">0</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
            and we found that adding this kernel to the sum of kernels improved our results empirically. The white noise kernel simply allows for a noise level on the data as independently<?pagebreak page5644?> and identically normally distributed, specified through a variance parameter. This parameter is similar to (and will interact with) the <inline-formula><mml:math id="M125" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> noise level described below, which is, however, tested systematically through cross-validation.</p>
      <p id="d1e2090">The Python scikit-learn implementation of the algorithm used here is based on Algorithm 2.1 of <xref ref-type="bibr" rid="bib1.bibx39" id="text.31"/>. We optimized the kernel parameters in the same way as for the other regression methods through fivefold cross-validation, and we subject them to the noise <inline-formula><mml:math id="M126" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> parameter of the scikit-learn GPR regression packages <xref ref-type="bibr" rid="bib1.bibx36" id="paren.32"/>. This parameter is not to be confused with the <inline-formula><mml:math id="M127" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> regularization parameter for ridge regression and takes the role of smoothing the kernel function so as to address overfitting. It represents a value added to the diagonal of the kernel matrix during the fitting process with larger <inline-formula><mml:math id="M128" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> values corresponding to greater noise level in the measurements of the outputs. However, we note that there is some equivalency with the <inline-formula><mml:math id="M129" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> parameter in ridge as the method is effectively a form of Tikhonov regularization that is also used in ridge regression <xref ref-type="bibr" rid="bib1.bibx36" id="paren.33"/>. Both inputs and outputs to the GPR function were standard-scaled to zero mean and unit variance based on the training data. For each GPR optimization, we chose 25 optimizer restarts with different initializations of the kernel parameters, which is necessary to approximate the best possible solution to maximize the log-marginal likelihood of the fit. More background on GPR can be found in <xref ref-type="bibr" rid="bib1.bibx39" id="text.34"/>.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Cross-validation</title>
      <p id="d1e2144">For all regression models, we performed fivefold cross-validation where the data are first split into training and test sets, keeping samples ordered by time. The training data are afterwards divided into five consecutive subsets (folds) of equal length. If the training data are not divisible by five, with a residual number of samples <inline-formula><mml:math id="M130" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>, then the first <inline-formula><mml:math id="M131" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> folds will contain one surplus sample compared to the remaining folds. Each fold is used once as a validation set, while the remaining four folds are used for training. The best set of model hyperparameters or kernel functions is found according to the average generalization error on these validation sets. After the best cross-validated hyperparameters are found, we refit the regression models on the entire training data using these hyperparameter settings (e.g. the <inline-formula><mml:math id="M132" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> value for which we found the best out-of-sample performance for ridge regression).</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><?xmltex \opttitle{NO${}_{2}$ sensor calibration}?><title>NO<inline-formula><mml:math id="M133" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor calibration</title>
      <p id="d1e2194">The skill of a sensor calibration function is expected to increase with sample size, i.e. the number of measurements used in the calibration process, but will also depend on aspects of the sampling environment. For co-location measurements, there will be time-dependent fluctuations in the value ranges encountered for the predictors (e.g. low-cost sensor signals, humidity, temperature) and predictands (reference NO<inline-formula><mml:math id="M134" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, PM<inline-formula><mml:math id="M135" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>). The calibration range in turn affects the performance of the calibration function: if faced with values outside its training range, the function effectively has to perform an extrapolation rather than interpolation, i.e. the function is not well constrained outside its training domain. This limitation is particularly critical for non-linear machine learning functions <xref ref-type="bibr" rid="bib1.bibx16 bib1.bibx33 bib1.bibx49" id="paren.35"/>. Calibration performance will further vary for each device, even for sensors of the same make, due to unavoidable randomness in the sensor production process <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx6" id="paren.36"/>. To characterize these various influences, we here test the dependence of three machine learning calibration methods, as well as of MLR, on sample size and co-location period for a number of NO<inline-formula><mml:math id="M136" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensors.</p>
      <p id="d1e2230">The NO<inline-formula><mml:math id="M137" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> co-location data at CR7 is ideally suited for this purpose. Twenty-one sensor nodes of the same make were co-located with a LAQN reference during the period October to December 2018 (Table 1). We actually co-located 30 sensor sets at the site, but we excluded any sensors with less than 820 h (samples) after outlier removal from our evaluation. The remaining sensors measure sometimes overlapping but still distinct time periods, because each sensor measurement varied in its precise co-location start and end time and was also subject to sensor-specific periods of malfunction. To detect these malfunctions, and to exclude the corresponding samples, we removed outliers (evidenced by unrealistically large measurement signals) at the original time resolution of our measurements, i.e. <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> min and prior to hourly averaging. To detect outliers for removal, we used the median absolute deviation (MAD) method, also known as “robust Z-Score method”, which identifies outliers for each variable based on their univariate deviation from their training data median. Since the median is a robust statistic to outliers itself, it is a typically a better measure to identify outliers than, for example, a deviation from the mean. Accordingly, we excluded any samples <inline-formula><mml:math id="M139" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> from the training and test data where the quantity
            <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M140" display="block"><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.6745</mml:mn><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mtext>median</mml:mtext><mml:mfenced open="{" close="}"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>
          takes on values <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:math></inline-formula> for any of the predictors, where <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the training data median value of each predictor. To train and cross-validate our calibration models, we took the first 820 h measured by each sensor set and split it into 600 h for training and cross-validation, leaving 220 h to measure the final skill on an out-of-sample test set. We highlight again that the test set will cover different time intervals for different sensors, meaning that further randomness is introduced in how we measure calibration skill. However, the relationships for each of the four calibration methods are learnt from exactly the same data and their predictions are also evaluated on the<?pagebreak page5645?> same data, meaning that their robustness and performance can still be directly compared. To measure calibration skill, we used two standard metrics in the form of the <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score (coefficient of determination), defined by
            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M144" display="block"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi>i</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          and the RMSE between the reference measurements <inline-formula><mml:math id="M145" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> and our calibrated signals <inline-formula><mml:math id="M146" display="inline"><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:math></inline-formula> on the test sets. For particularly poor calibration functions, the <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score can take on infinitely negative values, whereas a value of 1 implies a perfect prediction. An <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score of 0 is equivalent to a function that predicts the correct long-term time average of the data but no fluctuations therein.</p>
      <p id="d1e2512">As discussed in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>, each of AirPublic's co-location nodes measures 15 signals (the predictors or inputs) that we consider relevant for the NO<inline-formula><mml:math id="M149" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor calibration against the LAQN reference signal for NO<inline-formula><mml:math id="M150" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> (the predictand or output). Each of the 15 inputs will potentially be systematically linearly or non-linearly correlated with the output, which allows us to learn a calibration function from the measurement data. Once we know this function, we should be able to make accurate predictions given new inputs to reproduce the LAQN reference. As we fit two linear and two non-linear algorithms, certain transformations of the inputs can be useful to facilitate the learning process. For example, a relationship between an input and the output might be an exponential dependence in the original time series so that applying a logarithmic transformation could lead to an approximately linear relationship that might be easier to learn for a linear regression function. We therefore compared three set-ups with different sets of predictors:
<list list-type="order"><list-item>
      <p id="d1e2537">using the 15 input time series as provided (label <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">15</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>);</p></list-item><list-item>
      <p id="d1e2552">adding logarithmic transformations of the predictors (<inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>); and</p></list-item><list-item>
      <p id="d1e2567">adding both logarithmic and exponential transformations of the predictors (<inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>).</p></list-item></list>
These are labelled according to their total number of predictors after adding the input transformations, i.e. 15, 30 and 45. The logarithmic and exponential transformations of each input signal <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are defined as<?xmltex \setcounter{equation}{8}?>

                <disp-formula id="Ch1.E9" specific-use="gather" content-type="subnumberedsingle"><mml:math id="M155" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E9.10"><mml:mtd><mml:mtext>9a</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>log⁡</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>log⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E9.11"><mml:mtd><mml:mtext>9b</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>exp⁡</mml:mi></mml:mrow></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M156" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mtext>max</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> is the maximum value of the predictor time series, and <inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. The latter prevents possible divisions by zero, whereas the former prevents overflow values in the function.</p><?xmltex \hack{\newpage}?>
<sec id="Ch1.S3.SS1.SSS1">
  <label>3.1.1</label><title>Comparison of regression models for all predictors</title>
      <p id="d1e2746">For a first comparison of the calibration performance of the four methods, we show <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores and RMSEs in Table 2, rows (a) to (c), averaged across all 21 sensor nodes. GPR emerges as the best-performing method for all three sets of predictor choices, reaching <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores better than 0.8 for <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. This highlights that GPR should from now on be considered an option in similar sensor calibration exercises. RFR consistently performs worse than GPR but slightly better than ridge regression, which in turn outperforms MLR in all cases, but the differences are fairly small for <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">15</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. A notable exception occurs for <inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, where the <inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score for MLR suddenly drops abruptly to around 0.2. This sudden performance loss can be understood from the aforementioned curse of dimensionality: MLR increasingly overfits the training data as the number of predictors increases; the existing sample size becomes too small to constrain the 45 regression coefficients <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx40" id="paren.37"/>. The machine learning methods can deal with this increase in dimensionality highly effectively and thus perform well throughout all three cases. Indeed, GPR and ridge regression benefit slightly from the additional predictor transformations. This robustness to regression dimensionality is a first central advantage of machine learning methods in sensor calibrations. Machine learning methods will be more reliable and will allow users to work in a higher-dimensional calibration space compared to MLR. Having said that, for 15 input features the performance of all methods appears very similar on first sight, making MLR seemingly a viable alternative to the machine learning methods. We note, however, that there is no apparent disadvantage in using machine learning methods to prevent potential dangers of overfitting depending on sample size.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T2" specific-use="star"><?xmltex \currentcnt{2}?><label>Table 2</label><caption><p id="d1e2844">Average NO<inline-formula><mml:math id="M166" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor skill depending on the selection of predictors.
Shown are average <inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores and root mean squared errors (in brackets;
units <inline-formula><mml:math id="M168" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>). Results are averaged over the 21 low-cost
sensor nodes with 600 hourly training samples each, and the evaluation is
carried out for 220 test samples each. RH stands for relative humidity, and <inline-formula><mml:math id="M169" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> stands
for temperature.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Input features</oasis:entry>
         <oasis:entry colname="col3">MLR</oasis:entry>
         <oasis:entry colname="col4">Ridge</oasis:entry>
         <oasis:entry colname="col5">RFR</oasis:entry>
         <oasis:entry colname="col6">GPR</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">(a)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">15</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.74 (6.2)</oasis:entry>
         <oasis:entry colname="col4">0.75 (6.1)</oasis:entry>
         <oasis:entry colname="col5">0.76 (5.9)</oasis:entry>
         <oasis:entry colname="col6">0.79 (5.7)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(b)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.73 (6.3)</oasis:entry>
         <oasis:entry colname="col4">0.75 (6.0)</oasis:entry>
         <oasis:entry colname="col5">0.76 (5.9)</oasis:entry>
         <oasis:entry colname="col6">0.81 (5.4)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(c)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.23 (10.6)</oasis:entry>
         <oasis:entry colname="col4">0.75 (6.0)</oasis:entry>
         <oasis:entry colname="col5">0.76 (5.9)</oasis:entry>
         <oasis:entry colname="col6">0.80 (5.5)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(d)</oasis:entry>
         <oasis:entry colname="col2">MiCS, <inline-formula><mml:math id="M173" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>, RH</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2.9</mml:mn></mml:mrow></mml:math></inline-formula> (28.3)</oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.03</mml:mn></mml:mrow></mml:math></inline-formula> (12.6)</oasis:entry>
         <oasis:entry colname="col5">0.01 (12.2)</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.12</mml:mn></mml:mrow></mml:math></inline-formula> (13.1)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(e)</oasis:entry>
         <oasis:entry colname="col2">A43F, <inline-formula><mml:math id="M177" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>, RH</oasis:entry>
         <oasis:entry colname="col3">0.25 (9.7)</oasis:entry>
         <oasis:entry colname="col4">0.22 (10.1)</oasis:entry>
         <oasis:entry colname="col5">0.44 (8.7)</oasis:entry>
         <oasis:entry colname="col6">0.47 (8.4)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(f)</oasis:entry>
         <oasis:entry colname="col2">B43F, <inline-formula><mml:math id="M178" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>, RH</oasis:entry>
         <oasis:entry colname="col3">0.20 (10.6)</oasis:entry>
         <oasis:entry colname="col4">0.39 (9.7)</oasis:entry>
         <oasis:entry colname="col5">0.43 (9.5)</oasis:entry>
         <oasis:entry colname="col6">0.49 (9.3)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(g)</oasis:entry>
         <oasis:entry colname="col2">NO/O<inline-formula><mml:math id="M179" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:math></inline-formula>/B43F/<inline-formula><mml:math id="M180" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>/RH</oasis:entry>
         <oasis:entry colname="col3">0.68 (6.9)</oasis:entry>
         <oasis:entry colname="col4">0.75 (6.2)</oasis:entry>
         <oasis:entry colname="col5">0.69 (6.8)</oasis:entry>
         <oasis:entry colname="col6">0.77 (6.0)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(h)</oasis:entry>
         <oasis:entry colname="col2">(g) + A43F</oasis:entry>
         <oasis:entry colname="col3">0.72 (6.4)</oasis:entry>
         <oasis:entry colname="col4">0.74 (6.2)</oasis:entry>
         <oasis:entry colname="col5">0.75 (6.1)</oasis:entry>
         <oasis:entry colname="col6">0.79 (5.7)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(i)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M181" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> + B43F (<inline-formula><mml:math id="M182" display="inline"><mml:mrow><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col3">0.78 (5.8)</oasis:entry>
         <oasis:entry colname="col4">0.79 (5.6)</oasis:entry>
         <oasis:entry colname="col5">0.78 (5.7)</oasis:entry>
         <oasis:entry colname="col6">0.84 (5.0)</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S3.SS1.SSS2">
  <label>3.1.2</label><title>Calibration performance depending on sample size</title>
      <p id="d1e3258">We next consider the performance dependence on sample size of the training data (Fig. <xref ref-type="fig" rid="Ch1.F4"/>). The advantages of machine learning methods become even more evident for smaller numbers of training samples, even if we consider case (b) with 30 predictors, i.e. <inline-formula><mml:math id="M183" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, for which we found that MLR performs fairly well if trained on 600 h of data. The mean <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score and RMSE (<inline-formula><mml:math id="M185" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>) quickly deteriorate for smaller sample sizes for MLR, in particular below a threshold of less than 400 h of training data. Ridge regression – its statistical learning equivalent – always outperforms MLR. Both GPR and RFR can already perform well at small samples sizes of less than 300 h. While all methods converge towards similar performance approaching 600 h of training data (Table 2), MLR is generally performing worse than ridge regression and significantly worse that RFR and GPR.</p>
      <?pagebreak page5646?><p id="d1e3304"><?xmltex \hack{\newpage}?>Further evidence for advantages of machine learning methods are provided in Fig. <xref ref-type="fig" rid="Ch1.F5"/>, showing boxplots of the <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score distributions across all 21 sensor nodes depending on sample size (300, 400, 500, 600 h) and regression method. While median sensor performances of MLR, ridge and GPR ultimately become comparable, MLR is typically found to yield a number of poor-performing calibration functions with some <inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores well below 0.6 even for 600 training hours. In contrast, the distributions are far narrower for the machine learning methods: GPR and RFR do not show a single extreme outlier even after being trained on only 400 h of data, providing strong indications that the two methods are the most reliable. After 600 h, one can effectively expect that all sensors will provide <inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula> if trained using GPR. Overall, this highlights again that machine learning methods will provide better average skill but are also expected to provide more reliable calibration functions through co-location measurements independent of sensor device and the peculiarities of the individual training and test data set.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e3355">Error metrics as a function of the number of training samples for <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, as labelled. The figure highlights the convergence of both metrics for the different regression methods as the sample size increases. Note that MLR would not converge for <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">45</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, owing to the curse of dimensionality. This tendency can also be seen here for small sample sizes, where MLR rapidly loses performance. Results are averaged over the 21 low-cost sensor nodes with 600 hourly training samples each, and the evaluation is carried out for 220 test samples each.</p></caption>
            <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f04.png"/>

          </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e3389">Node-specific <inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores depending on calibration method and training sample size, evaluated on consistent 220 h test data sets in each case (see main text). The boxes extend from the lower to the upper quartile; inset lines mark the median. The whiskers extending from the box indicate the range excluding outliers (fliers). Each circle represents the <inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score on the test set for an individual node (21 in total). For MLR, some sensor nodes remain poorly calibrated even for larger sample sizes.</p></caption>
            <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f05.png"/>

          </fig>

</sec>
<sec id="Ch1.S3.SS1.SSS3">
  <label>3.1.3</label><?xmltex \opttitle{Calibration performance depending on predictor choices and NO${}_{2}$ device}?><title>Calibration performance depending on predictor choices and NO<inline-formula><mml:math id="M195" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> device</title>
      <p id="d1e3449">Tests (a) to (c) listed in Table <xref ref-type="table" rid="Ch1.T2"/> indicate that the machine learning regressions for NO<inline-formula><mml:math id="M196" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, specifically GPR, can benefit slightly from additional logarithmic predictor transformations but that adding exponential transformations on top of these predictors does not further increase predictive skill, as measured through the <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score and RMSE. Incorporating the logarithmic transformations, we next tested the importance of various predictors to achieve a certain level of calibration skill (rows (d) to (i) in Table <xref ref-type="table" rid="Ch1.T2"/>). This provides two important insights: firstly, we test the predictive skill if we use the individual MiCS and AlphaSense NO<inline-formula><mml:math id="M198" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensors separately, i.e. if individual sensors are performing better than others in our calibration setting and if we need all sensors to obtain the best level of calibration performance. Secondly, we test if other environmental influences such as humidity and temperature significantly affect sensor performance.</p>
      <p id="d1e3485">We first tested three set-ups in which we used only the sensor signals of the two cheaper MiCS devices (d) and then set-ups with the more expensive AlphaSense A43F (e) and B43F (f) devices. Using just the MiCS devices, the <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score drops from 0.75–0.81 for the machine learning methods to around zero, meaning that hardly any of the variation in the true NO<inline-formula><mml:math id="M200" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> reference signal is captured. Using our calibration approach here, the MiCS would therefore not be sufficient to achieve a meaningful measurement performance. The picture looks slightly better, albeit still far from perfect, for the individual A43F and B43F devices for which <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores of almost 0.5 are reached using non-linear calibration methods. We note that the linear MLR and ridge methods do not achieve the same performance, but ridge outperforms MLR.<?pagebreak page5647?> The most recently developed AlphaSense sensor used in our study, B43F, is the best-performing stand-alone sensor. If we add the <inline-formula><mml:math id="M202" display="inline"><mml:mrow class="chem"><mml:mi mathvariant="normal">NO</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="normal">ozone</mml:mi></mml:mrow></mml:math></inline-formula> sensor as well as the humidity and temperature signals to the predictors – case (g) – its performance alone almost reaches the same as for the <inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> configuration. This implies that the interference with <inline-formula><mml:math id="M204" display="inline"><mml:mrow class="chem"><mml:mi mathvariant="normal">NO</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="normal">ozone</mml:mi></mml:mrow></mml:math></inline-formula>, temperature and humidity might be significant and has to be taken into account in the calibration, and if only one sensor could be chosen for the measurements, the B43F sensor would be the best choice. By further adding the A43F sensor to the predictors the predictive skill is only mildly improved (h). Finally, we note that, in this stationary sensor setting, further predictive skill can be gained by considering past measurement values. Here, we included the 1 h lagged signal of the best B43F sensor (i). This is clearly only possible if there is a delayed consistency (or autocorrelation) in the data, which here leads to the best average <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> generalization score of 0.84 for GPR and related gains in terms of the RMSE. While being an interesting feature, we will not consider such set-ups in the following, because we intend sensors to be transferable among locations, and they should only rely on live signals for the hour of measurement in question.</p>
      <p id="d1e3566">In summary, using all sensor signals in combination is a robust and skilful set-up for our NO<inline-formula><mml:math id="M206" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor calibration and is therefore a prudent choice, at least if one of the machine learning methods is used to control for the curse of dimensionality. In particular, the B43F sensor is important to consider in the calibration, but further calibration skill is gained by also considering environmental factors, the presence of interference from ozone and NO, and additional NO<inline-formula><mml:math id="M207" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> devices.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><?xmltex \opttitle{PM${}_{{10}}$ sensor calibration}?><title>PM<inline-formula><mml:math id="M208" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensor calibration</title>
      <p id="d1e3606">In the same way as for NO<inline-formula><mml:math id="M209" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>, we tested several calibration settings for the PM<inline-formula><mml:math id="M210" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors. For this purpose, we consider the measurements for the location CarPark, where we co-located three sensors (IDs 19, 25 and 26) with a higher-cost device (Table 1). However, after data cleaning, we have only 509 and 439 samples (hours) for sensors 19 and 25 available, respectively, which our NO<inline-formula><mml:math id="M211" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> analysis above indicates is too short to obtain robust statistics for training and testing the sensors. Instead we focus our analysis on sensor 26 for which there are 1314 h of measurements available. We split these data into 400 samples for training and cross-validation, leaving 914 samples for testing the sensor calibration. Below we discuss results for various calibration configurations, using the 24 predictors for PM<inline-formula><mml:math id="M212" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>) and the same four regression methods as for NO<inline-formula><mml:math id="M213" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>. The baseline case with just 24 predictors is named <inline-formula><mml:math id="M214" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">24</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, following the same nomenclature as for NO<inline-formula><mml:math id="M215" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>. <inline-formula><mml:math id="M216" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">48</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">72</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> refer to the cases with additional logarithmic and exponential transformations of the predictors according to Eqs. (<xref ref-type="disp-formula" rid="Ch1.E9.10"/>) and (<xref ref-type="disp-formula" rid="Ch1.E9.11"/>). In addition, we test the effects of environmental conditions, as expressed through relative humidity and temperature, by excluding these two variables from the calibration procedure, while using the <inline-formula><mml:math id="M218" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">48</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> set-up with the additional log-transformed predictors.</p>
      <p id="d1e3715">The results of these tests are summarized in Table <xref ref-type="table" rid="Ch1.T3"/>. For the baseline case of 24 non-transformed predictors, RFR (<inline-formula><mml:math id="M219" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.70</mml:mn></mml:mrow></mml:math></inline-formula>) is outperformed by ridge regression (<inline-formula><mml:math id="M220" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.79</mml:mn></mml:mrow></mml:math></inline-formula>) and GPR (<inline-formula><mml:math id="M221" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.79</mml:mn></mml:mrow></mml:math></inline-formula>). This is mainly the result of the fact that some of the pollution values measured with sensor 26 during the test period lie outside the range of values encountered at training stage. RFR cannot predict values beyond its training range (i.e. it cannot extrapolate to higher values) and can therefore not predict those values accurately <xref ref-type="bibr" rid="bib1.bibx49 bib1.bibx25" id="paren.38"><named-content content-type="pre">see also</named-content></xref>. Instead, RFR<?pagebreak page5648?> constantly predicts the maximum value encountered during training in those cases.</p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T3" specific-use="star"><?xmltex \currentcnt{3}?><label>Table 3</label><caption><p id="d1e3773">Sensor skill depending on the selection of predictors for PM<inline-formula><mml:math id="M222" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensor 26 at location CarPark. Shown are the <inline-formula><mml:math id="M223" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores and RMSEs (in brackets) evaluated over 914 h of test data, after training and cross-validating the algorithms on 400 h of data. RMSEs are given in <inline-formula><mml:math id="M224" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. In (d), relative humidity (RH) and temperature (<inline-formula><mml:math id="M225" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula>) are removed from the calibration variables to test their importance for the measurement skill.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:colspec colnum="5" colname="col5" align="center"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Predictors</oasis:entry>
         <oasis:entry colname="col3">MLR</oasis:entry>
         <oasis:entry colname="col4">Ridge</oasis:entry>
         <oasis:entry colname="col5">RFR</oasis:entry>
         <oasis:entry colname="col6">GPR</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">(a)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M226" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">24</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.28 (13.0)</oasis:entry>
         <oasis:entry colname="col4">0.79 (6.9)</oasis:entry>
         <oasis:entry colname="col5">0.70 (8.3)</oasis:entry>
         <oasis:entry colname="col6">0.79 (7.1)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(b)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">48</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.70 (8.4)</oasis:entry>
         <oasis:entry colname="col4">0.80 (6.8)</oasis:entry>
         <oasis:entry colname="col5">0.70 (8.3)</oasis:entry>
         <oasis:entry colname="col6">0.76 (7.4)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(c)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">72</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.47 (11.1)</oasis:entry>
         <oasis:entry colname="col4">0.80 (6.8)</oasis:entry>
         <oasis:entry colname="col5">0.70 (8.3)</oasis:entry>
         <oasis:entry colname="col6">0.79 (7.0)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(d)</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M229" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">48</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>  – RH, <inline-formula><mml:math id="M230" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">0.76 (7.6)</oasis:entry>
         <oasis:entry colname="col4">0.78 (7.2)</oasis:entry>
         <oasis:entry colname="col5">0.67 (8.8)</oasis:entry>
         <oasis:entry colname="col6">0.75 (7.6)</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e3999">However, this problem is not entirely exclusive to RFR but is inherited by all methods, with RFR only being the most prominent case. We illustrate the more general issue, which will occur in any co-location calibration setting, in Fig. <xref ref-type="fig" rid="Ch1.F6"/>. In the training data, there are not any pollution values beyond ca. 40 <inline-formula><mml:math id="M231" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, so the RFR predictions simply level off at that value. This is a serious constraint in actual field measurements where one would be particularly interested in episodes of highest pollution. We note that this effect is somewhat alleviated by using GPR and even more so by ridge regression. For the latter, this behaviour is intuitive as the linear relationships learnt by ridge will hold to a good approximation even under extrapolation to previously unseen values. However, even for ridge regression the predictions eventually deviate from the <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> line for the highest pollution levels. This aspect will be crucial to consider for any co-location calibration approach, as is also evident from the poor MLR performance, despite being another linear method. In addition, MLR sometimes predicts substantially negative values, producing an overall <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score of below 0.3, whereas the machine learning methods appear to avoid the problem of negative predictions almost entirely. In conclusion, we highlight the necessity for co-location studies to ensure that maximum pollution values encountered during training and testing/deployment are as similar as possible. Extrapolations beyond 10–20 <inline-formula><mml:math id="M234" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> appear to be unreliable even if ridge regression is used as calibration algorithm, which is the best among our four methods to combat extrapolation issues.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e4067">Calibrated PM<inline-formula><mml:math id="M235" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> values (in <inline-formula><mml:math id="M236" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>) versus the reference measurements for 900 h of test data at location CarPark for the <inline-formula><mml:math id="M237" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">24</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> predictor set-up. The ideal <inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> perfect prediction line is drawn in black. Inset values <inline-formula><mml:math id="M239" display="inline"><mml:mi>R</mml:mi></mml:math></inline-formula> are the Pearson correlation coefficients.</p></caption>
          <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f06.png"/>

        </fig>

      <p id="d1e4134">A test with additional log-transformations (<inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">48</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) of the predictors led to test score improvements for the two linear methods (Table 3), in particular for MLR (<inline-formula><mml:math id="M241" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula>) but also for ridge regression (<inline-formula><mml:math id="M242" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>). This implies that the log-transformations have helped linearize certain predictor–predictand relationships. Further exponential transformations (<inline-formula><mml:math id="M243" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">72</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) and thus also further increasing the predictor dimensionality did not lead to an improvement in calibration skill. We therefore ran one final test using the <inline-formula><mml:math id="M244" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">48</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> set-up but without relative humidity and temperature included as predictors. This test confirmed that the sensor signals indeed experience a slight interference from humidity and temperature, at least considering the machine learning regressions. Notably, this loss of skill is not observed for MLR for which the <inline-formula><mml:math id="M245" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score actually improves. A likely explanation for this behaviour is the curse of dimensionality that affects MLR more significantly than the three machine learning methods, so the reduction in collinear dimensions (given the sample size constraint) is more beneficial than the information gained by including temperature and humidity in the MLR regression.</p>
      <p id="d1e4212">In summary, we have found that ridge regression and GPR are the two most reliable and high-performing calibration methods for the PM<inline-formula><mml:math id="M246" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensor. We are able to attain very good <inline-formula><mml:math id="M247" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores <inline-formula><mml:math id="M248" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula> for all four regression methods though. An important point to highlight is the characteristics of the training domain, in particular of the pollution levels encountered during the training data measurements. If the value range is not sufficient to cover the range of interest for future measurement campaigns, then ridge regression might be the most robust choice to alleviate the underprediction of the most extreme pollution values. However, the power of extrapolation of any method is limited, so we underline the need to carefully check every training data set to see if it fulfils such crucial criteria; see also similar discussions in other calibration contexts <xref ref-type="bibr" rid="bib1.bibx16 bib1.bibx49 bib1.bibx25" id="paren.39"/>.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Site transferability</title>
      <p id="d1e4256">Finally, we aim to address the question of site transferability, i.e. how reliably a sensor calibrated through co-location can be used to measure air pollution at a different location. One of the sensor nodes (ID 19) was used for NO<inline-formula><mml:math id="M249" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> measurements at both locations, CR7 and CR9, and was also used to measure PM<inline-formula><mml:math id="M250" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> at CR9 and CarPark, allowing us to address this question for our methodology. Note that these tests also include a shift in the time of year (Table 1), which has been hypothesized to be one potentially limiting factor in site transferability.
The results of these transferability tests for PM<inline-formula><mml:math id="M251" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> (from CR9 to CarPark and vice versa) and NO<inline-formula><mml:math id="M252" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> (from CR7 to CR9 and vice versa) are shown in Figs. <xref ref-type="fig" rid="Ch1.F7"/> and <xref ref-type="fig" rid="Ch1.F8"/>, respectively.</p>
      <?pagebreak page5649?><p id="d1e4300">For PM<inline-formula><mml:math id="M253" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula>, we trained the regressions, using the <inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">24</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> predictor set-up, on 400 h of data at either location. This emulates a situation in which, according to our results above, we limit the co-location period to a minimum number of samples required to achieve reasonable performances across all four regression methods. To mitigate issues related to extrapolation (Fig. <xref ref-type="fig" rid="Ch1.F6"/>), we selected the last 400 h of the time series for location CarPark (Fig. <xref ref-type="fig" rid="Ch1.F7"/>a) and hours 600 to 1000 of the time series for location CR9 (Fig. <xref ref-type="fig" rid="Ch1.F7"/>b). This way we still emulate a possible minimal scenario of 400 consecutive hours of co-location, while also including near maximum and minimum pollution values within our training data (given the available measurement data). We note that alternative sampling approaches, such as random sampling with shuffling of the data, could lead to artificial effects at validation and testing stages because of autocorrelation effects that could inflate apparent calibration skill. The maximum pollution found within the two time slices differ only by ca. <inline-formula><mml:math id="M255" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M256" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, for which at least ridge regression should provide reasonable extrapolation performances. For the resulting predictions at location CarPark, using models trained on the CR9 data, we achieve generally very good <inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores ranging between 0.67 for RFR and 0.78 for ridge regression. The site-transferred measurement performance of sensor 19 is therefore almost as good as the one for sensor 26 at the co-location site itself (Table 3); i.e. we cannot detect any significant loss in measurement performance due to the site transfer. A surprising element is that MLR performs almost as good as ridge regression in this case, whereas it performed poorly for sensor 26, where it only achieved an <inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score of 0.28 (Table 3). This underlines our previous observation that the performance of MLR is more sensitive to the specific sensor hardware, with sometimes low performance for relatively small sample sizes (Fig. <xref ref-type="fig" rid="Ch1.F5"/>). However, our results also show again that linear methods appear to generally perform well for our PM<inline-formula><mml:math id="M259" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors, with ridge regression being the most reliable and high-performing choice overall. These results are further supported by calculations concerning the detection of only the most extreme pollution events in the time series; see also the definition of such events described in the caption of Fig. <xref ref-type="fig" rid="Ch1.F7"/>. We characterize such events in the form of their statistical recall in our sensor measurements (the fraction of extreme pollution events in the reference time series that are also identified by our calibrated sensors) and precision (the fraction of extreme events identified by our sensors that are indeed also extreme events in the reference time series) and show the results as inset numbers in Figs. <xref ref-type="fig" rid="Ch1.F7"/> and <xref ref-type="fig" rid="Ch1.F8"/>. Overall, these site-transfer PM<inline-formula><mml:math id="M260" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> results from CR9 to CarPark imply that sensors calibrated through co-location can achieve high measurement performance distant from the co-location site.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e4410">Tests of PM<inline-formula><mml:math id="M261" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensor site transfers using calibration models trained on 400 h of data. <bold>(a)</bold> Predictions for the four regression models (as labelled) and reference measurements at location CarPark, using models trained at CR9, and <bold>(b)</bold> at location CR9, using models trained on data measured at CarPark. The inset values provide the <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores for each method relative to the reference as well as the corresponding recall and precision for the detection of the strongest pollution events, which are typically of particular interest in real-life situations (here defined as events when two values within the last 3 h exceeded a threshold of 35 <inline-formula><mml:math id="M263" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>). For compactness, we only show data for times at which both reference and low-cost sensor data were available, and we label these hours as a consecutive timeline.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f07.png"/>

        </fig>

      <p id="d1e4465">However, we do find that site transferability is not always as straightforward as found for this particular case. For example, for the inverse transfer using models trained at CarPark and predicting PM<inline-formula><mml:math id="M264" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> pollution levels at CR9, we find lower <inline-formula><mml:math id="M265" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores overall. There is consistency in the sense that MLR and ridge remain the best performing methods for sensor 19 with <inline-formula><mml:math id="M266" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores of around 0.5, but the sensors now miss several significant pollution events in the time series. We note, however, that many of the most extreme pollution events are still detected, which is evident from the still relatively high precision and recall scores for all methods. These results underline that, in general at least, a good performance can be achieved with co-location calibrations but that there are also significant challenges posed by site transfers. In particular, the problem is not necessarily symmetric among sites, i.e. the skill of the method can depend on the direction of the transfer, even if the pollution levels at both sites are similar. We therefore hope that our insights and results will motivate further work in this direction, with the aim to identify possible causes of such effects. We discuss some of the possible reasons for this behaviour in Sect. 4.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><?xmltex \def\figurename{Figure}?><label>Figure 8</label><caption><p id="d1e4501">Tests of NO<inline-formula><mml:math id="M267" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor site transfers using calibration models trained on all available samples for locations CR7 and CR9. <bold>(a)</bold> Predictions for the four regression models (as labelled) and reference measurements at location CR7, using models trained on data from CR9 and <bold>(b)</bold> vice versa. The inset values provide the <inline-formula><mml:math id="M268" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores for each method relative to the reference as well as the corresponding recall and precision for the detection of the strongest pollution events, which are typically of particular interest in real-life situations. Since the two locations were subject to very different pollution ranges (note the different value ranges on the <inline-formula><mml:math id="M269" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axes), these are defined in <bold>(a)</bold> as 45 <inline-formula><mml:math id="M270" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and in <bold>(b)</bold> as 90 <inline-formula><mml:math id="M271" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, and we indicate these thresholds by grey dashed lines. An extreme pollution event occurs when the threshold is exceeded for 2 of the last 3 h. For compactness, we only show data for times at which both reference and low-cost sensor data were available, and we label these hours as a consecutive timeline.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://amt.copernicus.org/articles/14/5637/2021/amt-14-5637-2021-f08.png"/>

        </fig>

      <?pagebreak page5650?><p id="d1e4588">Similarly, we find promising results for the NO<inline-formula><mml:math id="M272" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor site transfer using the <inline-formula><mml:math id="M273" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mn mathvariant="normal">30</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> predictor set-up (Fig. <xref ref-type="fig" rid="Ch1.F8"/>). The key challenge for the sensor transfer from CR7 to CR9 is that the maximal pollution levels at the two locations differ substantially, with peak concentration being around 100 <inline-formula><mml:math id="M274" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> greater at CR9. To allow for the best possible learning opportunities for the regression algorithms, we therefore used all available samples for training, which are 1482 samples at CR9 and 829 samples at CR7. This leads to overall good performance of the non-linear RFR and GPR methods at location CR7 using models trained at CR9. As no extrapolation is necessary, these methods achieve a good performance of <inline-formula><mml:math id="M275" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores <inline-formula><mml:math id="M276" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula> and also a good balance of precision and recall. The results are, however, slightly worse than for the same site calibrations (Table 2). Ridge regression has a tendency to overpredict NO<inline-formula><mml:math id="M277" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> pollution levels in this particular case, likely because it cannot capture some non-linear effects that would have limited the prediction values. As a result, it also reproduces almost all extreme pollution events where the concentration of NO<inline-formula><mml:math id="M278" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> exceeds 45 <inline-formula><mml:math id="M279" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> (recall <inline-formula><mml:math id="M280" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.98</mml:mn></mml:mrow></mml:math></inline-formula>) but also predicts many false pollution events (precision <inline-formula><mml:math id="M281" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.49</mml:mn></mml:mrow></mml:math></inline-formula>).</p>
      <p id="d1e4711">Despite the large sample size, MLR performs poorly for both site transfers (<inline-formula><mml:math id="M282" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M283" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.31</mml:mn></mml:mrow></mml:math></inline-formula>). In particular, MLR underpredicts, sometimes providing even impossible negative pollution estimates at CR7, whereas it provides several runaway positive values at CR9 (Fig. <xref ref-type="fig" rid="Ch1.F8"/>). However, at CR9 all methods struggle with the impossible challenge of extrapolation far outside their training domain, which effectively is an extreme demonstration of the effects of an ill-considered training range (cf. Fig. <xref ref-type="fig" rid="Ch1.F6"/>). Among the machine learning methods, the effect is as expected most prominent for RFR (<inline-formula><mml:math id="M284" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.36</mml:mn></mml:mrow></mml:math></inline-formula>), which cannot predict any pollution values beyond those encountered at training stage. This is a serious limitation and means that the method scores nil on precision and recall of any extreme pollution events at CR9 where NO<inline-formula><mml:math id="M285" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> levels exceeded 90 <inline-formula><mml:math id="M286" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. GPR is slightly better at extrapolating beyond its training domain (compare also Fig. <xref ref-type="fig" rid="Ch1.F6"/>) but still not good enough to reproduce any of the extreme pollution events, giving rise to equally low precision and recall. Ridge regression, as a regularized linear method, performs best in the sense that it is able to reproduce at least a few of the extreme events (recall <inline-formula><mml:math id="M287" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>) and predicting no false extreme events (precision <inline-formula><mml:math id="M288" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1.0</mml:mn></mml:mrow></mml:math></inline-formula>), while still achieving an <inline-formula><mml:math id="M289" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> score of 0.53. Nonetheless, it is clear from the time series in Fig. <xref ref-type="fig" rid="Ch1.F8"/> that none of the regression methods works for this site transfer, simply because of the too large extrapolation range.</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Discussion and conclusions</title>
      <p id="d1e4832">We have compared four different regression methods to calibrate a number of low-cost NO<inline-formula><mml:math id="M290" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M291" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors against reference measurement signals, by means of co-location at three separate sites in London, UK. A summary of the various features of each regression algorithm is given in Table <xref ref-type="table" rid="Ch1.T4"/>. Comparing the four regression methods, our main conclusions are the following:
<list list-type="order"><list-item>
      <p id="d1e4857">For the 21 NO<inline-formula><mml:math id="M292" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> sensor nodes, Gaussian process regression (GPR) is generally performing best at the same measurement site, followed by ridge regression, random forest regression (RFR), and multiple linear regression (MLR). For a single sensor PM<inline-formula><mml:math id="M293" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> calibration, we find that ridge regression and GPR attain about the same measurement performance, with a slight edge for ridge<?pagebreak page5651?> regression. We note that in particular the relative performance of GPR differs greatly from a recent study by <xref ref-type="bibr" rid="bib1.bibx25" id="text.40"/>, likely due to our different choice of kernel design.</p></list-item><list-item>
      <p id="d1e4882">Special care must be taken of the calibration conditions, in particular if sensors are thereafter used for measurements in areas where higher pollution levels are to be expected. The linear ridge method can best mitigate the catastrophic measurement failure in such extrapolation settings, for both NO<inline-formula><mml:math id="M294" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M295" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> measurements, but also fails if measurement signals deviate by more than around 10–20 <inline-formula><mml:math id="M296" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">µ</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">m</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> from the maximum pollution level in the training data. For our NO<inline-formula><mml:math id="M297" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> measurement with site transfer, we find that the non-linear methods, GPR and RFR, can outperform ridge regression, assuming that the training pollution range encapsulates the range of values encountered at the new site. For the PM<inline-formula><mml:math id="M298" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensor calibrations and corresponding site transfers, we find that ridge regression is the highest performing and most reliable calibration algorithm overall.</p></list-item><list-item>
      <p id="d1e4941">All three machine learning methods (ridge, GPR and RFR) generally outperform or perform as least as good as MLR. The machine learning methods are also more reliable if many signals are used for calibration, or if the number of measurement samples is relatively small. MLR suffers most significantly from the curse of dimensionality in those settings and can produce highly erroneous results.</p></list-item><list-item>
      <p id="d1e4945">Under careful consideration of the calibration conditions, given expected measurement conditions, the low-cost sensors typically achieve high performances with <inline-formula><mml:math id="M299" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> scores often exceeding 0.8 on new unseen test data.</p></list-item></list></p>

<?xmltex \floatpos{t}?><table-wrap id="Ch1.T4" specific-use="star"><?xmltex \currentcnt{4}?><label>Table 4</label><caption><p id="d1e4962">Summary and qualitative evaluation of key features of each of the four regression algorithms. The robustness to the curse of dimensionality refers to how well each algorithm can handle an increase in the number of calibration variables. In addition, we summarize how well each method performs for our calibrations here if evaluated against test data sampled at the same site as the one used for training and after site transfers. The variable performance of GPR (and also RFR) for site transfers is mainly the result of its limited skill to extrapolate so that its performance will always depend strongly on how representative the training data is in terms of value ranges. We also note that robustness to the curse of dimensionality (third column) is only considered with respect to the maximum of 72 predictors considered here. In even higher dimensions more significant differences across the machine learning methods are expected to occur eventually.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="left"/>
     <oasis:colspec colnum="7" colname="col7" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Non-linear</oasis:entry>
         <oasis:entry colname="col3">Robustness to curse</oasis:entry>
         <oasis:entry namest="col4" nameend="col5" align="center" colsep="1">Same site </oasis:entry>
         <oasis:entry namest="col6" nameend="col7" align="center">Extrapolation </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3">of dimensionality</oasis:entry>
         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1">performance </oasis:entry>
         <oasis:entry rowsep="1" namest="col6" nameend="col7" align="center">(site transfer skill) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">NO<inline-formula><mml:math id="M300" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col5">PM<inline-formula><mml:math id="M301" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">NO<inline-formula><mml:math id="M302" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">PM<inline-formula><mml:math id="M303" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">MLR</oasis:entry>
         <oasis:entry colname="col2">No</oasis:entry>
         <oasis:entry colname="col3">Poor</oasis:entry>
         <oasis:entry colname="col4">Poor to good</oasis:entry>
         <oasis:entry colname="col5">Poor to good</oasis:entry>
         <oasis:entry colname="col6">Poor</oasis:entry>
         <oasis:entry colname="col7">Moderate to good</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Ridge</oasis:entry>
         <oasis:entry colname="col2">No</oasis:entry>
         <oasis:entry colname="col3">Very good</oasis:entry>
         <oasis:entry colname="col4">Very good</oasis:entry>
         <oasis:entry colname="col5">Best</oasis:entry>
         <oasis:entry colname="col6">Moderate</oasis:entry>
         <oasis:entry colname="col7">Good</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GPR</oasis:entry>
         <oasis:entry colname="col2">Yes</oasis:entry>
         <oasis:entry colname="col3">Very good</oasis:entry>
         <oasis:entry colname="col4">Best</oasis:entry>
         <oasis:entry colname="col5">Very good</oasis:entry>
         <oasis:entry colname="col6">Poor to good</oasis:entry>
         <oasis:entry colname="col7">Moderate to good</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RFR</oasis:entry>
         <oasis:entry colname="col2">Yes</oasis:entry>
         <oasis:entry colname="col3">Very good</oasis:entry>
         <oasis:entry colname="col4">Very good</oasis:entry>
         <oasis:entry colname="col5">Good</oasis:entry>
         <oasis:entry colname="col6">Poor to good</oasis:entry>
         <oasis:entry colname="col7">Poor to good</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d1e5179">On another note, we highlight that we sometimes found significant signals in our test data sets that were not reproduced by our low-cost sensor nodes (see, for example, the pollution spike at <inline-formula><mml:math id="M304" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">140</mml:mn></mml:mrow></mml:math></inline-formula> in Fig. <xref ref-type="fig" rid="Ch1.F7"/>a) even if the measured pollution value lies well within the training data range. It is hard to assign reasons to this surprising sensor behaviour, as our low-cost sensors are apparently able to capture most of the other pollution spikes well for the same data set. One<?pagebreak page5652?> possible reason is a calibration blind spot; i.e. we encounter a new type of sensor interference which we did not find in the training data. For example, this could be substantial changes in environmental conditions or, for example, PM composition, which are not captured by the calibration function. In our interpretation of results, this would represent again an extrapolation with respect to the predictors and/or the predictands. However, we think that this is unlikely, given that the behaviour is not found frequently, at least in this particular time series. Two other options are (a) imperfect co-location (e.g. we might have missed an important local pollution plume by chance) or (b) temporary sensor failures that were removed by the MAD outlier removal, i.e. that our sensors were temporarily inoperable at the time of a pollution spike that dominated the values for the given measurement hour. Interesting aspects to explore as part of future work might be to compare how sensitive site transfer performances are to the measurement principle of the NO<inline-formula><mml:math id="M305" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M306" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> devices used and effects of the seasonal cycle (e.g. winter-based calibrations to be used during summer). We hope that future measurement campaigns can provide further insights into such calibration challenges, and we hope that our study can motivate further work in this direction.</p>
      <p id="d1e5215">In conclusion, our results underline the potential of machine learning algorithms for the calibration of co-located low-cost NO<inline-formula><mml:math id="M307" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> and PM<inline-formula><mml:math id="M308" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> sensors. At the same time, we highlight several significant challenges that will always have to be considered in similar calibration processes. This includes the need for a well-adjusted calibration data set to avoid calibration failure if the algorithm needs to extrapolate to significantly higher pollution values and the role of individual choices relating to the combination of calibration variables and calibration algorithms, e.g. concerning the curse of dimensionality, predictor transformations and linearity in the predictor–predictand relationships. Recent studies indicate that the issues related to extrapolation can be mitigated through the application of hybrid models in which non-linear machine learning models are used within the training domain and a simpler linear regression approach otherwise <xref ref-type="bibr" rid="bib1.bibx16 bib1.bibx25" id="paren.41"/>. We note that in particular ridge regression could be a good compromise, which does not require a somewhat arbitrary hybrid-model definition. Having said that, we also found that even high-dimensional linear methods have ultimately limited extrapolation skill (Fig. <xref ref-type="fig" rid="Ch1.F6"/>), so the consideration of the training data pollution range remains of fundamental importance. We hope that such insights will contribute to ever less expensive and more spatially dense measurements of air pollution in the future and that our work will motivate additional measurement campaigns, testing of other calibration algorithms and further low-cost sensor development.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e5246">Analysis code will be made available on Peer Nowack's GitHub website (<uri>https://github.com/peernow/AMT2021</uri>; <ext-link xlink:href="https://doi.org/10.5281/zenodo.5215849" ext-link-type="DOI">10.5281/zenodo.5215849</ext-link>, <xref ref-type="bibr" rid="bib1.bibx32" id="altparen.42"/>) and under <uri>https://github.com/airpublic/thing_api</uri> (last access: 1 May 2019).</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e5264">Data are available from Hannah Gardiner (hannahgardiner1@hotmail.com).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e5270">PN wrote the paper and carried out the analysis, supported by LK. The study was designed and suggested by HG, LK and PN. LK, JC and HG (through AirPublic) measured and prepared the data.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e5276">The authors declare that they have no conflict of interest.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <?pagebreak page5653?><p id="d1e5282">Publisher’s note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e5288">Peer Nowack is supported through an Imperial College Research Fellowship. AirPublic is supported by ClimateKIC, DigitalCatapult and Future Cities Catapult via Organicity.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e5293">This research has been supported by the European Research Council, H2020 European Research Council (OrganiCity (grant no. 645198)).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e5299">This paper was edited by Hang Su and reviewed by Carl Malings and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><?xmltex \def\ref@label{{Bishop(2006)}}?><label>Bishop(2006)</label><?label Bishop2006?><mixed-citation>
Bishop, C. M.: Pattern recognition and machine learning, Springer  Science+Business Media, Singapore, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx2"><?xmltex \def\ref@label{{Breiman(2001)}}?><label>Breiman(2001)</label><?label Breiman2001?><mixed-citation>Breiman, L.: Random forests, Mach. Learn., 45, 5–32,  <ext-link xlink:href="https://doi.org/10.1201/9780429469275-8" ext-link-type="DOI">10.1201/9780429469275-8</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx3"><?xmltex \def\ref@label{{Breiman and Friedman(1997)}}?><label>Breiman and Friedman(1997)</label><?label Breiman1997?><mixed-citation>Breiman, L. and Friedman, J. H.: Predicting multivariate responses in multiple linear regression, J. Roy. Stat. Soc.-B, 59, 3–54, <ext-link xlink:href="https://doi.org/10.1111/1467-9868.00054" ext-link-type="DOI">10.1111/1467-9868.00054</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Casey and Hannigan(2018)}}?><label>Casey and Hannigan(2018)</label><?label Casey2018?><mixed-citation>Casey, J. G. and Hannigan, M. P.: Testing the performance of field calibration techniques for low-cost gas sensors in new deployment locations: across a county line and across Colorado, Atmos. Meas. Tech., 11, 6351–6378, <ext-link xlink:href="https://doi.org/10.5194/amt-11-6351-2018" ext-link-type="DOI">10.5194/amt-11-6351-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx5"><?xmltex \def\ref@label{{Casey et~al.(2019)n}}?><label>Casey et al.(2019)n</label><?label Casey2019?><mixed-citation>Casey, J. G., Collier-Oxandale, A., and Hannigan, M.: Performance of  artificial neural networks and linear models to quantify 4 trace gas species  in an oil and gas production region with low-cost sensors, Sensor.
Actuat. B-Chem., 283, 504–514, <ext-link xlink:href="https://doi.org/10.1016/j.snb.2018.12.049" ext-link-type="DOI">10.1016/j.snb.2018.12.049</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx6"><?xmltex \def\ref@label{{Castell et~al.(2017)}}?><label>Castell et al.(2017)</label><?label Castell2017?><mixed-citation>Castell, N., Dauge, F. R., Schneider, P., Vogt, M., Lerner, U., Fishbain, B.,  Broday, D., and Bartonova, A.: Can commercial low-cost sensor platforms  contribute to air quality monitoring and exposure estimates?, Environ.  Int., 99, 293–302, <ext-link xlink:href="https://doi.org/10.1016/j.envint.2016.12.007" ext-link-type="DOI">10.1016/j.envint.2016.12.007</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx7"><?xmltex \def\ref@label{{Cross et~al.(2017)}}?><label>Cross et al.(2017)</label><?label Cross2017?><mixed-citation>Cross, E. S., Williams, L. R., Lewis, D. K., Magoon, G. R., Onasch, T. B., Kaminsky, M. L., Worsnop, D. R., and Jayne, J. T.: Use of electrochemical sensors for measurement of air pollution: correcting interference response and validating measurements, Atmos. Meas. Tech., 10, 3575–3588, <ext-link xlink:href="https://doi.org/10.5194/amt-10-3575-2017" ext-link-type="DOI">10.5194/amt-10-3575-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx8"><?xmltex \def\ref@label{{{De Vito} et~al.(2018)}}?><label>De Vito et al.(2018)</label><?label DeVito2018?><mixed-citation>De Vito, S., Esposito, E., Salvato, M., Popoola, O., Formisano, F., Jones,  R., and Di Francia, G.: Calibrating chemical multisensory devices for real  world applications: An in-depth comparison of quantitative machine learning  approaches, Sensor. Actuat. B-Chem., 255, 1191–1210,  <ext-link xlink:href="https://doi.org/10.1016/j.snb.2017.07.155" ext-link-type="DOI">10.1016/j.snb.2017.07.155</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx9"><?xmltex \def\ref@label{{{De Vito} et~al.(2019)}}?><label>De Vito et al.(2019)</label><?label DeVito2019?><mixed-citation>De Vito, S., Esposito, E., Formisano, F., Massera, E., Auria, P. D., and Di Francia, G.: Adaptive Machine learning for Backup Air Quality Multisensor Systems continuous calibration, 2019 IEEE International Symposium on Olfaction and Electronic Nose (ISOEN), 26–29 May 2019, Fukuoka, Japan, 1–4,  <ext-link xlink:href="https://doi.org/10.1109/isoen.2019.8823250" ext-link-type="DOI">10.1109/isoen.2019.8823250</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Dormann et~al.(2013)}}?><label>Dormann et al.(2013)</label><?label Dormann2013?><mixed-citation>Dormann, C. F., Elith, J., Bacher, S., Buchmann, C., Carl, G., Carré, G., Marquéz, J. R., Gruber, B., Lafourcade, B., Leitão, P. J.,  Münkemüller, T., Mcclean, C., Osborne, P. E., Reineking, B., Schröder, B., Skidmore, A. K., Zurell, D., and Lautenbach, S.: Collinearity: A review of methods to deal with it and a simulation study  evaluating their performance, Ecography, 36, 27–46,  <ext-link xlink:href="https://doi.org/10.1111/j.1600-0587.2012.07348.x" ext-link-type="DOI">10.1111/j.1600-0587.2012.07348.x</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx11"><?xmltex \def\ref@label{{Eilenberg et~al.(2020)}}?><label>Eilenberg et al.(2020)</label><?label RoseEilenberg2020?><mixed-citation>Eilenberg, S. R., Subramanian, R., Malings, C., Hauryliuk, A., Presto, A. A.,  and Robinson, A. L.: Using a network of lower-cost monitors to identify the  influence of modifiable factors driving spatial patterns in fine particulate  matter concentrations in an urban environment, J. Expo. Sci. Env. Epid., 30, 949–961, <ext-link xlink:href="https://doi.org/10.1038/s41370-020-0255-x" ext-link-type="DOI">10.1038/s41370-020-0255-x</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx12"><?xmltex \def\ref@label{{Esposito et~al.(2016)}}?><label>Esposito et al.(2016)</label><?label Esposito2016?><mixed-citation>Esposito, E., De Vito, S., Salvato, M., Bright, V., Jones, R. L., and  Popoola, O.: Dynamic neural network architectures for on field stochastic  calibration of indicative low cost air quality sensing systems, Sensor.
Actuat. B-Chem., 231, 701–713, <ext-link xlink:href="https://doi.org/10.1016/j.snb.2016.03.038" ext-link-type="DOI">10.1016/j.snb.2016.03.038</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx13"><?xmltex \def\ref@label{{{European Environment Agency}(2019)}}?><label>European Environment Agency(2019)</label><?label EuropeanEnvironmentAgency2019?><mixed-citation>European Environment Agency: Air quality in Europe – 2019 report, available at:  <uri>http://www.eea.europa.eu/publications/air-quality-in-europe-2012</uri> (last access: 1 November 2020), 2019.</mixed-citation></ref>
      <ref id="bib1.bibx14"><?xmltex \def\ref@label{{Fang and Bate(2017)}}?><label>Fang and Bate(2017)</label><?label Fang2017a?><mixed-citation>
Fang, X. and Bate, I.: Using Multi-parameters for Calibration of Low-cost  Sensors in Urban Environment, Proceedings of the 2017 International  Conference on Embedded Wireless Systems and Networks, 20–22 February 2017, Uppsala, Sweden, 1–11, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx15"><?xmltex \def\ref@label{{Green et~al.(2009)}}?><label>Green et al.(2009)</label><?label Green2009?><mixed-citation>Green, D. C., Fuller, G. W., and Baker, T.: Development and validation of the  volatile correction model for PM<inline-formula><mml:math id="M309" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> – An empirical method for adjusting TEOM measurements for their loss of volatile particulate matter, Atmos.  Environ., 43, 2132–2141, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2009.01.024" ext-link-type="DOI">10.1016/j.atmosenv.2009.01.024</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Hagan et~al.(2018)}}?><label>Hagan et al.(2018)</label><?label Hagan2017?><mixed-citation>Hagan, D. H., Isaacman-VanWertz, G., Franklin, J. P., Wallace, L. M. M., Kocar, B. D., Heald, C. L., and Kroll, J. H.: Calibration and assessment of electrochemical air quality sensors by co-location with regulatory-grade instruments, Atmos. Meas. Tech., 11, 315–328, <ext-link xlink:href="https://doi.org/10.5194/amt-11-315-2018" ext-link-type="DOI">10.5194/amt-11-315-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx17"><?xmltex \def\ref@label{{Hagler et~al.(2018)}}?><label>Hagler et al.(2018)</label><?label Hagler2018?><mixed-citation>Hagler, G. S., Williams, R., Papapostolou, V., and Polidori, A.: Air Quality  Sensors and Data Adjustment Algorithms: When Is It No Longer a Measurement?,  Environ. Sci. Technol., 52, 5530–5531, <ext-link xlink:href="https://doi.org/10.1021/acs.est.8b01826" ext-link-type="DOI">10.1021/acs.est.8b01826</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx18"><?xmltex \def\ref@label{{Hoerl and Kennard(1970)}}?><label>Hoerl and Kennard(1970)</label><?label Hoerl1970?><mixed-citation>Hoerl, A. E. and Kennard, R. W.: Ridge Regression: Biased Estimation for  Nonorthogonal Problems, Technometrics, 12, 55–67, <ext-link xlink:href="https://doi.org/10.1080/00401706.2000.10485983" ext-link-type="DOI">10.1080/00401706.2000.10485983</ext-link>, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx19"><?xmltex \def\ref@label{{James et~al.(2013)}}?><label>James et al.(2013)</label><?label James2013?><mixed-citation>James, G., Witten, D., Hastie, T., and Tibshirani, R.: An Introduction to  Statistical Learning, Springer Science+Business Media, New York,  <ext-link xlink:href="https://doi.org/10.1007/978-1-4614-7138-7" ext-link-type="DOI">10.1007/978-1-4614-7138-7</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx20"><?xmltex \def\ref@label{{Jiao et~al.(2016)}}?><label>Jiao et al.(2016)</label><?label Jiao2016?><mixed-citation>Jiao, W., Hagler, G., Williams, R., Sharpe, R., Brown, R., Garver, D., Judge, R., Caudill, M., Rickard, J., Davis, M., Weinstock, L., Zimmer-Dauphinee, S., and Buckley, K.: Community Air Sensor Network (CAIRSENSE) project: evaluation of low-cost sensor performance in a suburban environment in the southeastern United States, Atmos. Meas. Tech., 9, 5281–5292, <ext-link xlink:href="https://doi.org/10.5194/amt-9-5281-2016" ext-link-type="DOI">10.5194/amt-9-5281-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Keller and Evans(2019)}}?><label>Keller and Evans(2019)</label><?label Keller2018?><mixed-citation>Keller, C. A. and Evans, M. J.: Application of random forest regression to<?pagebreak page5654?> the calculation of gas-phase chemistry within the GEOS-Chem chemistry model v10, Geosci. Model Dev., 12, 1209–1225, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-1209-2019" ext-link-type="DOI">10.5194/gmd-12-1209-2019</ext-link>,  2019.</mixed-citation></ref>
      <ref id="bib1.bibx22"><?xmltex \def\ref@label{{Lewis et~al.(2016)}}?><label>Lewis et al.(2016)</label><?label Lewis2016?><mixed-citation>Lewis, A. C., Lee, J. D., Edwards, P. M., Shaw, M. D., Evans, M. J., Moller,  S. J., Smith, K. R., Buckley, J. W., Ellis, M., Gillot, S. R., and White, A.:  Evaluating the performance of low cost chemical sensors for air pollution  research, Faraday Discuss., 189, 85–103, <ext-link xlink:href="https://doi.org/10.1039/c5fd00201j" ext-link-type="DOI">10.1039/c5fd00201j</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx23"><?xmltex \def\ref@label{{Lewis et~al.(2018)}}?><label>Lewis et al.(2018)</label><?label Lewis2018report?><mixed-citation>Lewis, A. C., von Schneidermesser, E., and Peltier, R. E.: Low-cost sensors  for the measurement of atmospheric composition: overview of topic and future  applications, Tech. rep.,  World Meteorological Organization, available at: <ext-link xlink:href="https://www.ccacoalition.org/en/resources/low-cost-sensors-measurement-atmospheric-composition-overview-topic-and-future">https://www.ccacoalition.org/en/resources/low-cost-sensors-measurement-atmospheric-composition-overview-topic-and-future</ext-link> (last access: 1 November 2020), 2018.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Liu et~al.(2019)}}?><label>Liu et al.(2019)</label><?label Liu2019?><mixed-citation>Liu, H. Y., Schneider, P., Haugen, R., and Vogt, M.: Performance assessment of  a low-cost PM<inline-formula><mml:math id="M310" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2.5</mml:mn></mml:msub></mml:math></inline-formula> sensor for a near four-month period in Oslo, Norway, Atmosphere, 10, 41, <ext-link xlink:href="https://doi.org/10.3390/atmos10020041" ext-link-type="DOI">10.3390/atmos10020041</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx25"><?xmltex \def\ref@label{{Malings et~al.(2019)}}?><label>Malings et al.(2019)</label><?label Malings2019?><mixed-citation>Malings, C., Tanzer, R., Hauryliuk, A., Kumar, S. P. N., Zimmerman, N., Kara, L. B., Presto, A. A., and R. Subramanian: Development of a general calibration model and long-term performance evaluation of low-cost sensors for air pollutant gas monitoring, Atmos. Meas. Tech., 12, 903–920, <ext-link xlink:href="https://doi.org/10.5194/amt-12-903-2019" ext-link-type="DOI">10.5194/amt-12-903-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx26"><?xmltex \def\ref@label{{Malings et~al.(2020)}}?><label>Malings et al.(2020)</label><?label Malings2020?><mixed-citation>Malings, C., Tanzer, R., Hauryliuk, A., Saha, P. K., Robinson, A. L., Presto,   A. A., and Subramanian, R.: Fine particle mass monitoring with low-cost sensors: Corrections and long-term performance evaluation, Aerosol Sci. Tech., 54, 160–174, <ext-link xlink:href="https://doi.org/10.1080/02786826.2019.1623863" ext-link-type="DOI">10.1080/02786826.2019.1623863</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Mansfield et~al.(2020)}}?><label>Mansfield et al.(2020)</label><?label Mansfield2020?><mixed-citation>Mansfield, L., Nowack, P., Kasoar, M., Everitt, R., Collins, W. J., and  Voulgarakis, A.: Can we predict climate change from short-term simulations   using machine learning?, npj Climate and Atmospheric Science, 3, 44,  <ext-link xlink:href="https://doi.org/10.1038/s41612-020-00148-5" ext-link-type="DOI">10.1038/s41612-020-00148-5</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx28"><?xmltex \def\ref@label{{Masson et~al.(2015)}}?><label>Masson et al.(2015)</label><?label Masson2015?><mixed-citation>Masson, N., Piedrahita, R., and Hannigan, M.: Quantification method for  electrolytic sensors in long-term monitoring of ambient air quality, Sensors, 15, 27283–27302, <ext-link xlink:href="https://doi.org/10.3390/s151027283" ext-link-type="DOI">10.3390/s151027283</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx29"><?xmltex \def\ref@label{{Mead et~al.(2013)}}?><label>Mead et al.(2013)</label><?label Mead2013?><mixed-citation>Mead, M. I., Popoola, O. A., Stewart, G. B., Landshoff, P., Calleja, M., Hayes, M., Baldovi, J. J., McLeod, M. W., Hodgson, T. F., Dicks, J., Lewis, A., Cohen, J., Baron, R., Saffell, J. R., and Jones, R. L.: The use of  electrochemical sensors for monitoring urban air quality in low-cost,  high-density networks, Atmos. Environ., 70, 186–203,  <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2012.11.060" ext-link-type="DOI">10.1016/j.atmosenv.2012.11.060</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{Moltchanov et~al.(2015)}}?><label>Moltchanov et al.(2015)</label><?label Moltchanov2015?><mixed-citation>Moltchanov, S., Levy, I., Etzion, Y., Lerner, U., Broday, D. M., and Fishbain, B.: On the feasibility of measuring urban air pollution by wireless distributed sensor networks, Sci. Total Environ., 502, 537–547, <ext-link xlink:href="https://doi.org/10.1016/j.scitotenv.2014.09.059" ext-link-type="DOI">10.1016/j.scitotenv.2014.09.059</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx31"><?xmltex \def\ref@label{{Munir et~al.(2019)}}?><label>Munir et al.(2019)</label><?label Munir2019?><mixed-citation>Munir, S., Mayfield, M., Coca, D., Jubb, S. A., and Osammor, O.: Analysing the performance of low-cost air quality sensors, their drivers, relative benefits and calibration in cities – a case study in Sheffield, Environ. Monit. Assess., 191, 94, <ext-link xlink:href="https://doi.org/10.1007/s10661-019-7231-8" ext-link-type="DOI">10.1007/s10661-019-7231-8</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx32"><?xmltex \def\ref@label{{Nowack and Konstantinovskiy(2021)}}?><label>Nowack and Konstantinovskiy(2021)</label><?label NowackKonstantinovskiy2021?><mixed-citation>Nowack, P. and Konstantinovskiy, L.: Code in support of Nowack et al. (2021) in Atmospheric Measurement Techniques (Version 2), Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.5215849" ext-link-type="DOI">10.5281/zenodo.5215849</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx33"><?xmltex \def\ref@label{{Nowack et~al.(2018)}}?><label>Nowack et al.(2018)</label><?label Nowack2018h?><mixed-citation>Nowack, P., Braesicke, P., Haigh, J., Abraham, N. L., Pyle, J., and  Voulgarakis, A.: Using machine learning to build temperature-based ozone  parameterizations for climate sensitivity simulations, Environ. Res. Lett., 13, 104016, <ext-link xlink:href="https://doi.org/10.1088/1748-9326/aae2be" ext-link-type="DOI">10.1088/1748-9326/aae2be</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx34"><?xmltex \def\ref@label{{Nowack et~al.(2019)}}?><label>Nowack et al.(2019)</label><?label Nowack2019a?><mixed-citation>
Nowack, P., Ong, Q. Y. E., Braesicke, P., Haigh, J. D., Luke, A., Pyle, J., and Voulgarakis, A.: Machine learning parameterizations for ozone: climate model transferability, in: Conference Proceedings of the 9th International   Workshop on Climate Informatics, 2–4 October 2019, Paris, France, 263–268, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx35"><?xmltex \def\ref@label{{Nowack et~al.(2020)}}?><label>Nowack et al.(2020)</label><?label Nowack2020?><mixed-citation>Nowack, P., Runge, J., Eyring, V., and Haigh, J. D.: Causal networks for  climate model evaluation and constrained projections, Nat. Commun., 11, 1415, <ext-link xlink:href="https://doi.org/10.1038/s41467-020-15195-y" ext-link-type="DOI">10.1038/s41467-020-15195-y</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx36"><?xmltex \def\ref@label{{Pedregosa et~al.(2011)}}?><label>Pedregosa et al.(2011)</label><?label Pedregosa2011?><mixed-citation>
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,  O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J.,  Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E.:  Scikit-learn: Machine learning in Python, J. Mach. Learn. Res., 12, 2825–2830, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx37"><?xmltex \def\ref@label{{Popoola et~al.(2016)}}?><label>Popoola et al.(2016)</label><?label Popoola2016?><mixed-citation>Popoola, O. A., Stewart, G. B., Mead, M. I., and Jones, R. L.: Development of  a baseline-temperature correction methodology for electrochemical sensors and  its implications for long-term stability, Atmos. Environ., 147, 330–343, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2016.10.024" ext-link-type="DOI">10.1016/j.atmosenv.2016.10.024</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx38"><?xmltex \def\ref@label{{Rai and Kumar(2018)}}?><label>Rai and Kumar(2018)</label><?label Rai2017?><mixed-citation>Rai, A. C. and Kumar, P.: Summary of air quality sensors and recommendations  for application, Ref. Ares, p. 65, available at: <ext-link xlink:href="https://www.iscapeproject.eu/wp-content/uploads/2017/09/iSCAPE_D1.5_Summary-of-air-quality-sensors-and-recommendations-for-application.pdf">https://www.iscapeproject.eu/wp-content/uploads/2017/09/iSCAPE_D1.5_Summary-of-air-quality-sensors-and-recommendations-for-application.pdf</ext-link> (last access: 1 November 2020), 2018.</mixed-citation></ref>
      <ref id="bib1.bibx39"><?xmltex \def\ref@label{{Rasmussen and Williams(2006)}}?><label>Rasmussen and Williams(2006)</label><?label Rasmussen2006?><mixed-citation>
Rasmussen, C. E. and Williams, C. K. I.: Gaussian Processes for Machine  Learning, MIT Press, Cambridge, Massachusetts, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx40"><?xmltex \def\ref@label{{Runge et~al.(2012)}}?><label>Runge et al.(2012)</label><?label Runge2012?><mixed-citation>Runge, J., Heitzig, J., Petoukhov, V., and Kurths, J.: Escaping the curse of  dimensionality in estimating multivariate transfer entropy, Phys. Rev.  Lett., 108, 258701, <ext-link xlink:href="https://doi.org/10.1103/PhysRevLett.108.258701" ext-link-type="DOI">10.1103/PhysRevLett.108.258701</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx41"><?xmltex \def\ref@label{{Runge et~al.(2019)}}?><label>Runge et al.(2019)</label><?label Runge2019a?><mixed-citation>Runge, J., Nowack, P., Kretschmer, M., Flaxman, S., and Sejdinovic, D.: Detecting and quantifying causal associations in large nonlinear time series  datasets, Science Advances, 5, eaau4996, <ext-link xlink:href="https://doi.org/10.1126/sciadv.aau4996" ext-link-type="DOI">10.1126/sciadv.aau4996</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx42"><?xmltex \def\ref@label{{Sadighi et~al.(2018)}}?><label>Sadighi et al.(2018)</label><?label Sadighi2018?><mixed-citation>Sadighi, K., Coffey, E., Polidori, A., Feenstra, B., Lv, Q., Henze, D. K., and Hannigan, M.: Intra-urban spatial variability of surface ozone in Riverside, CA: viability and validation of low-cost sensors, Atmos. Meas. Tech., 11, 1777–1792, <ext-link xlink:href="https://doi.org/10.5194/amt-11-1777-2018" ext-link-type="DOI">10.5194/amt-11-1777-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx43"><?xmltex \def\ref@label{{Sayahi et~al.(2020)}}?><label>Sayahi et al.(2020)</label><?label Sayahi2020?><mixed-citation>Sayahi, T., Garff, A., Quah, T., Lê, K., Becnel, T., Powell, K. M.,  Gaillardon, P. E., Butterfield, A. E., and Kelly, K. E.: Long-term  calibration models to estimate ozone concentrations with a metal oxide  sensor, Environ. Pollut., 267, 115363, <ext-link xlink:href="https://doi.org/10.1016/j.envpol.2020.115363" ext-link-type="DOI">10.1016/j.envpol.2020.115363</ext-link>,  2020.</mixed-citation></ref>
      <ref id="bib1.bibx44"><?xmltex \def\ref@label{{Sherwen et~al.(2019)}}?><label>Sherwen et al.(2019)</label><?label Sherwen2019?><mixed-citation>Sherwen, T., Chance, R. J., Tinel, L., Ellis, D., Evans, M. J., and Carpenter, L. J.: A machine-learning-based global sea-surface iodide distribution, Earth Syst. Sci. Data, 11, 1239–1262, <ext-link xlink:href="https://doi.org/10.5194/essd-11-1239-2019" ext-link-type="DOI">10.5194/essd-11-1239-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx45"><?xmltex \def\ref@label{{Spinelle et~al.(2015)}}?><label>Spinelle et al.(2015)</label><?label Spinelle2015?><mixed-citation>Spinelle, L., Gerboles, M., Villani, M. G., Aleixandre, M., and Bonavitacola,  F.: Field calibration of a cluster of low-cost available sensors for air  quality monitoring. Part A: Ozone and nitrogen dioxide, Sensor. Actuat. B-Chem., 215, 249–257, <ext-link xlink:href="https://doi.org/10.1016/j.snb.2015.03.031" ext-link-type="DOI">10.1016/j.snb.2015.03.031</ext-link>, 2015.</mixed-citation></ref>
      <?pagebreak page5655?><ref id="bib1.bibx46"><?xmltex \def\ref@label{{Spinelle et~al.(2017)}}?><label>Spinelle et al.(2017)</label><?label Spinelle2017?><mixed-citation>Spinelle, L., Gerboles, M., Villani, M. G., Aleixandre, M., and Bonavitacola,  F.: Field calibration of a cluster of low-cost commercially available  sensors for air quality monitoring. Part B: NO, CO and CO2, Sensor. Actuat. B-Chem., 238, 706–715, <ext-link xlink:href="https://doi.org/10.1016/j.snb.2016.07.036" ext-link-type="DOI">10.1016/j.snb.2016.07.036</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx47"><?xmltex \def\ref@label{{Tanzer et~al.(2019)}}?><label>Tanzer et al.(2019)</label><?label Tanzer2019?><mixed-citation>Tanzer, R., Malings, C., Hauryliuk, A., Subramanian, R., and Presto, A. A.: Demonstration of a low-cost multi-pollutant network to quantify intra-urban  spatial variations in air pollutant source impacts and to evaluate  environmental justice, Int. J. Environ. Res. Pub. He., 16, 2523, <ext-link xlink:href="https://doi.org/10.3390/ijerph16142523" ext-link-type="DOI">10.3390/ijerph16142523</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx48"><?xmltex \def\ref@label{{Vikram et~al.(2019)}}?><label>Vikram et al.(2019)</label><?label Vikram2019?><mixed-citation>Vikram, S., Collier-Oxandale, A., Ostertag, M. H., Menarini, M., Chermak, C., Dasgupta, S., Rosing, T., Hannigan, M., and Griswold, W. G.: Evaluating and improving the reliability of gas-phase sensor system calibrations across new locations for ambient measurements and personal exposure monitoring, Atmos. Meas. Tech., 12, 4211–4239, <ext-link xlink:href="https://doi.org/10.5194/amt-12-4211-2019" ext-link-type="DOI">10.5194/amt-12-4211-2019</ext-link>, 2019.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx49"><?xmltex \def\ref@label{{Zimmerman et~al.(2018)}}?><label>Zimmerman et al.(2018)</label><?label Zimmerman2018?><mixed-citation>Zimmerman, N., Presto, A. A., Kumar, S. P. N., Gu, J., Hauryliuk, A., Robinson, E. S., Robinson, A. L., and R. Subramanian: A machine learning calibration model using random forests to improve sensor performance for lower-cost air quality monitoring, Atmos. Meas. Tech., 11, 291–313, <ext-link xlink:href="https://doi.org/10.5194/amt-11-291-2018" ext-link-type="DOI">10.5194/amt-11-291-2018</ext-link>, 2018.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Machine learning calibration of low-cost NO<sub>2</sub> and PM<sub>10</sub> sensors: non-linear algorithms and their impact on site transferability</article-title-html>
<abstract-html><p>Low-cost air pollution sensors often fail to attain sufficient performance compared with state-of-the-art measurement stations, and they typically require expensive laboratory-based calibration procedures. A repeatedly proposed strategy to overcome these limitations is calibration through co-location with public measurement stations. Here we test the idea of using machine learning algorithms for such calibration tasks using hourly-averaged co-location data for nitrogen dioxide (NO<sub>2</sub>) and particulate matter of particle sizes smaller than 10&thinsp;µm (PM<sub>10</sub>) at three different locations in the urban area of London, UK. We compare the performance of ridge regression, a linear statistical learning algorithm, to two non-linear algorithms in the form of random forest regression (RFR) and Gaussian process regression (GPR). We further benchmark the performance of all three machine learning methods relative to the more common multiple linear regression (MLR). We obtain very good out-of-sample <i>R</i><sup>2</sup> scores (coefficient of determination)  &gt; 0.7, frequently exceeding 0.8, for the machine learning calibrated low-cost sensors. In contrast, the performance of MLR is more dependent on random variations in the sensor hardware and co-located signals, and it is also more sensitive to the length of the co-location period. We find that, subject to certain conditions, GPR is typically the best-performing method in our calibration setting, followed by ridge regression and RFR. We also highlight several key limitations of the machine learning methods, which will be crucial to consider in any co-location calibration. In particular, all methods are fundamentally limited in how well they can reproduce pollution levels that lie outside those encountered at training stage. We find, however, that the linear ridge regression outperforms the non-linear methods in extrapolation settings. GPR can allow for a small degree of extrapolation, whereas RFR can only predict values within the training range. This algorithm-dependent ability to extrapolate is one of the key limiting factors when the calibrated sensors are deployed away from the co-location site itself. Consequently, we find that ridge regression is often performing as good as or even better than GPR after sensor relocation. Our results highlight the potential of co-location approaches paired with machine learning calibration techniques to reduce costs of air pollution measurements, subject to careful consideration of the co-location training conditions, the choice of calibration variables and the features of the calibration algorithm.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Bishop(2006)</label><mixed-citation>
Bishop, C. M.: Pattern recognition and machine learning, Springer  Science+Business Media, Singapore, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Breiman(2001)</label><mixed-citation>
Breiman, L.: Random forests, Mach. Learn., 45, 5–32,  <a href="https://doi.org/10.1201/9780429469275-8" target="_blank">https://doi.org/10.1201/9780429469275-8</a>, 2001.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Breiman and Friedman(1997)</label><mixed-citation>
Breiman, L. and Friedman, J. H.: Predicting multivariate responses in multiple linear regression, J. Roy. Stat. Soc.-B, 59, 3–54, <a href="https://doi.org/10.1111/1467-9868.00054" target="_blank">https://doi.org/10.1111/1467-9868.00054</a>, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Casey and Hannigan(2018)</label><mixed-citation>
Casey, J. G. and Hannigan, M. P.: Testing the performance of field calibration techniques for low-cost gas sensors in new deployment locations: across a county line and across Colorado, Atmos. Meas. Tech., 11, 6351–6378, <a href="https://doi.org/10.5194/amt-11-6351-2018" target="_blank">https://doi.org/10.5194/amt-11-6351-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Casey et al.(2019)n</label><mixed-citation>
Casey, J. G., Collier-Oxandale, A., and Hannigan, M.: Performance of  artificial neural networks and linear models to quantify 4 trace gas species  in an oil and gas production region with low-cost sensors, Sensor.
Actuat. B-Chem., 283, 504–514, <a href="https://doi.org/10.1016/j.snb.2018.12.049" target="_blank">https://doi.org/10.1016/j.snb.2018.12.049</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Castell et al.(2017)</label><mixed-citation>
Castell, N., Dauge, F. R., Schneider, P., Vogt, M., Lerner, U., Fishbain, B.,  Broday, D., and Bartonova, A.: Can commercial low-cost sensor platforms  contribute to air quality monitoring and exposure estimates?, Environ.  Int., 99, 293–302, <a href="https://doi.org/10.1016/j.envint.2016.12.007" target="_blank">https://doi.org/10.1016/j.envint.2016.12.007</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Cross et al.(2017)</label><mixed-citation>
Cross, E. S., Williams, L. R., Lewis, D. K., Magoon, G. R., Onasch, T. B., Kaminsky, M. L., Worsnop, D. R., and Jayne, J. T.: Use of electrochemical sensors for measurement of air pollution: correcting interference response and validating measurements, Atmos. Meas. Tech., 10, 3575–3588, <a href="https://doi.org/10.5194/amt-10-3575-2017" target="_blank">https://doi.org/10.5194/amt-10-3575-2017</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>De Vito et al.(2018)</label><mixed-citation>
De Vito, S., Esposito, E., Salvato, M., Popoola, O., Formisano, F., Jones,  R., and Di Francia, G.: Calibrating chemical multisensory devices for real  world applications: An in-depth comparison of quantitative machine learning  approaches, Sensor. Actuat. B-Chem., 255, 1191–1210,  <a href="https://doi.org/10.1016/j.snb.2017.07.155" target="_blank">https://doi.org/10.1016/j.snb.2017.07.155</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>De Vito et al.(2019)</label><mixed-citation>
De Vito, S., Esposito, E., Formisano, F., Massera, E., Auria, P. D., and Di Francia, G.: Adaptive Machine learning for Backup Air Quality Multisensor Systems continuous calibration, 2019 IEEE International Symposium on Olfaction and Electronic Nose (ISOEN), 26–29 May 2019, Fukuoka, Japan, 1–4,  <a href="https://doi.org/10.1109/isoen.2019.8823250" target="_blank">https://doi.org/10.1109/isoen.2019.8823250</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Dormann et al.(2013)</label><mixed-citation>
Dormann, C. F., Elith, J., Bacher, S., Buchmann, C., Carl, G., Carré, G., Marquéz, J. R., Gruber, B., Lafourcade, B., Leitão, P. J.,  Münkemüller, T., Mcclean, C., Osborne, P. E., Reineking, B., Schröder, B., Skidmore, A. K., Zurell, D., and Lautenbach, S.: Collinearity: A review of methods to deal with it and a simulation study  evaluating their performance, Ecography, 36, 27–46,  <a href="https://doi.org/10.1111/j.1600-0587.2012.07348.x" target="_blank">https://doi.org/10.1111/j.1600-0587.2012.07348.x</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Eilenberg et al.(2020)</label><mixed-citation>
Eilenberg, S. R., Subramanian, R., Malings, C., Hauryliuk, A., Presto, A. A.,  and Robinson, A. L.: Using a network of lower-cost monitors to identify the  influence of modifiable factors driving spatial patterns in fine particulate  matter concentrations in an urban environment, J. Expo. Sci. Env. Epid., 30, 949–961, <a href="https://doi.org/10.1038/s41370-020-0255-x" target="_blank">https://doi.org/10.1038/s41370-020-0255-x</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Esposito et al.(2016)</label><mixed-citation>
Esposito, E., De Vito, S., Salvato, M., Bright, V., Jones, R. L., and  Popoola, O.: Dynamic neural network architectures for on field stochastic  calibration of indicative low cost air quality sensing systems, Sensor.
Actuat. B-Chem., 231, 701–713, <a href="https://doi.org/10.1016/j.snb.2016.03.038" target="_blank">https://doi.org/10.1016/j.snb.2016.03.038</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>European Environment Agency(2019)</label><mixed-citation>
European Environment Agency: Air quality in Europe – 2019 report, available at:  <a href="http://www.eea.europa.eu/publications/air-quality-in-europe-2012" target="_blank"/> (last access: 1 November 2020), 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Fang and Bate(2017)</label><mixed-citation>
Fang, X. and Bate, I.: Using Multi-parameters for Calibration of Low-cost  Sensors in Urban Environment, Proceedings of the 2017 International  Conference on Embedded Wireless Systems and Networks, 20–22 February 2017, Uppsala, Sweden, 1–11, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Green et al.(2009)</label><mixed-citation>
Green, D. C., Fuller, G. W., and Baker, T.: Development and validation of the  volatile correction model for PM<sub>10</sub> – An empirical method for adjusting TEOM measurements for their loss of volatile particulate matter, Atmos.  Environ., 43, 2132–2141, <a href="https://doi.org/10.1016/j.atmosenv.2009.01.024" target="_blank">https://doi.org/10.1016/j.atmosenv.2009.01.024</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Hagan et al.(2018)</label><mixed-citation>
Hagan, D. H., Isaacman-VanWertz, G., Franklin, J. P., Wallace, L. M. M., Kocar, B. D., Heald, C. L., and Kroll, J. H.: Calibration and assessment of electrochemical air quality sensors by co-location with regulatory-grade instruments, Atmos. Meas. Tech., 11, 315–328, <a href="https://doi.org/10.5194/amt-11-315-2018" target="_blank">https://doi.org/10.5194/amt-11-315-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Hagler et al.(2018)</label><mixed-citation>
Hagler, G. S., Williams, R., Papapostolou, V., and Polidori, A.: Air Quality  Sensors and Data Adjustment Algorithms: When Is It No Longer a Measurement?,  Environ. Sci. Technol., 52, 5530–5531, <a href="https://doi.org/10.1021/acs.est.8b01826" target="_blank">https://doi.org/10.1021/acs.est.8b01826</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Hoerl and Kennard(1970)</label><mixed-citation>
Hoerl, A. E. and Kennard, R. W.: Ridge Regression: Biased Estimation for  Nonorthogonal Problems, Technometrics, 12, 55–67, <a href="https://doi.org/10.1080/00401706.2000.10485983" target="_blank">https://doi.org/10.1080/00401706.2000.10485983</a>, 1970.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>James et al.(2013)</label><mixed-citation>
James, G., Witten, D., Hastie, T., and Tibshirani, R.: An Introduction to  Statistical Learning, Springer Science+Business Media, New York,  <a href="https://doi.org/10.1007/978-1-4614-7138-7" target="_blank">https://doi.org/10.1007/978-1-4614-7138-7</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Jiao et al.(2016)</label><mixed-citation>
Jiao, W., Hagler, G., Williams, R., Sharpe, R., Brown, R., Garver, D., Judge, R., Caudill, M., Rickard, J., Davis, M., Weinstock, L., Zimmer-Dauphinee, S., and Buckley, K.: Community Air Sensor Network (CAIRSENSE) project: evaluation of low-cost sensor performance in a suburban environment in the southeastern United States, Atmos. Meas. Tech., 9, 5281–5292, <a href="https://doi.org/10.5194/amt-9-5281-2016" target="_blank">https://doi.org/10.5194/amt-9-5281-2016</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Keller and Evans(2019)</label><mixed-citation>
Keller, C. A. and Evans, M. J.: Application of random forest regression to the calculation of gas-phase chemistry within the GEOS-Chem chemistry model v10, Geosci. Model Dev., 12, 1209–1225, <a href="https://doi.org/10.5194/gmd-12-1209-2019" target="_blank">https://doi.org/10.5194/gmd-12-1209-2019</a>,  2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Lewis et al.(2016)</label><mixed-citation>
Lewis, A. C., Lee, J. D., Edwards, P. M., Shaw, M. D., Evans, M. J., Moller,  S. J., Smith, K. R., Buckley, J. W., Ellis, M., Gillot, S. R., and White, A.:  Evaluating the performance of low cost chemical sensors for air pollution  research, Faraday Discuss., 189, 85–103, <a href="https://doi.org/10.1039/c5fd00201j" target="_blank">https://doi.org/10.1039/c5fd00201j</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Lewis et al.(2018)</label><mixed-citation>
Lewis, A. C., von Schneidermesser, E., and Peltier, R. E.: Low-cost sensors  for the measurement of atmospheric composition: overview of topic and future  applications, Tech. rep.,  World Meteorological Organization, available at: <a href="https://www.ccacoalition.org/en/resources/low-cost-sensors-measurement-atmospheric-composition-overview-topic-and-future" target="_blank">https://www.ccacoalition.org/en/resources/low-cost-sensors-measurement-atmospheric-composition-overview-topic-and-future</a> (last access: 1 November 2020), 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Liu et al.(2019)</label><mixed-citation>
Liu, H. Y., Schneider, P., Haugen, R., and Vogt, M.: Performance assessment of  a low-cost PM<sub>2.5</sub> sensor for a near four-month period in Oslo, Norway, Atmosphere, 10, 41, <a href="https://doi.org/10.3390/atmos10020041" target="_blank">https://doi.org/10.3390/atmos10020041</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Malings et al.(2019)</label><mixed-citation>
Malings, C., Tanzer, R., Hauryliuk, A., Kumar, S. P. N., Zimmerman, N., Kara, L. B., Presto, A. A., and R. Subramanian: Development of a general calibration model and long-term performance evaluation of low-cost sensors for air pollutant gas monitoring, Atmos. Meas. Tech., 12, 903–920, <a href="https://doi.org/10.5194/amt-12-903-2019" target="_blank">https://doi.org/10.5194/amt-12-903-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Malings et al.(2020)</label><mixed-citation>
Malings, C., Tanzer, R., Hauryliuk, A., Saha, P. K., Robinson, A. L., Presto,   A. A., and Subramanian, R.: Fine particle mass monitoring with low-cost sensors: Corrections and long-term performance evaluation, Aerosol Sci. Tech., 54, 160–174, <a href="https://doi.org/10.1080/02786826.2019.1623863" target="_blank">https://doi.org/10.1080/02786826.2019.1623863</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Mansfield et al.(2020)</label><mixed-citation>
Mansfield, L., Nowack, P., Kasoar, M., Everitt, R., Collins, W. J., and  Voulgarakis, A.: Can we predict climate change from short-term simulations   using machine learning?, npj Climate and Atmospheric Science, 3, 44,  <a href="https://doi.org/10.1038/s41612-020-00148-5" target="_blank">https://doi.org/10.1038/s41612-020-00148-5</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Masson et al.(2015)</label><mixed-citation>
Masson, N., Piedrahita, R., and Hannigan, M.: Quantification method for  electrolytic sensors in long-term monitoring of ambient air quality, Sensors, 15, 27283–27302, <a href="https://doi.org/10.3390/s151027283" target="_blank">https://doi.org/10.3390/s151027283</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Mead et al.(2013)</label><mixed-citation>
Mead, M. I., Popoola, O. A., Stewart, G. B., Landshoff, P., Calleja, M., Hayes, M., Baldovi, J. J., McLeod, M. W., Hodgson, T. F., Dicks, J., Lewis, A., Cohen, J., Baron, R., Saffell, J. R., and Jones, R. L.: The use of  electrochemical sensors for monitoring urban air quality in low-cost,  high-density networks, Atmos. Environ., 70, 186–203,  <a href="https://doi.org/10.1016/j.atmosenv.2012.11.060" target="_blank">https://doi.org/10.1016/j.atmosenv.2012.11.060</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Moltchanov et al.(2015)</label><mixed-citation>
Moltchanov, S., Levy, I., Etzion, Y., Lerner, U., Broday, D. M., and Fishbain, B.: On the feasibility of measuring urban air pollution by wireless distributed sensor networks, Sci. Total Environ., 502, 537–547, <a href="https://doi.org/10.1016/j.scitotenv.2014.09.059" target="_blank">https://doi.org/10.1016/j.scitotenv.2014.09.059</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Munir et al.(2019)</label><mixed-citation>
Munir, S., Mayfield, M., Coca, D., Jubb, S. A., and Osammor, O.: Analysing the performance of low-cost air quality sensors, their drivers, relative benefits and calibration in cities – a case study in Sheffield, Environ. Monit. Assess., 191, 94, <a href="https://doi.org/10.1007/s10661-019-7231-8" target="_blank">https://doi.org/10.1007/s10661-019-7231-8</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Nowack and Konstantinovskiy(2021)</label><mixed-citation>
Nowack, P. and Konstantinovskiy, L.: Code in support of Nowack et al. (2021) in Atmospheric Measurement Techniques (Version 2), Zenodo [code], <a href="https://doi.org/10.5281/zenodo.5215849" target="_blank">https://doi.org/10.5281/zenodo.5215849</a>, 2021.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Nowack et al.(2018)</label><mixed-citation>
Nowack, P., Braesicke, P., Haigh, J., Abraham, N. L., Pyle, J., and  Voulgarakis, A.: Using machine learning to build temperature-based ozone  parameterizations for climate sensitivity simulations, Environ. Res. Lett., 13, 104016, <a href="https://doi.org/10.1088/1748-9326/aae2be" target="_blank">https://doi.org/10.1088/1748-9326/aae2be</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Nowack et al.(2019)</label><mixed-citation>
Nowack, P., Ong, Q. Y. E., Braesicke, P., Haigh, J. D., Luke, A., Pyle, J., and Voulgarakis, A.: Machine learning parameterizations for ozone: climate model transferability, in: Conference Proceedings of the 9th International   Workshop on Climate Informatics, 2–4 October 2019, Paris, France, 263–268, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Nowack et al.(2020)</label><mixed-citation>
Nowack, P., Runge, J., Eyring, V., and Haigh, J. D.: Causal networks for  climate model evaluation and constrained projections, Nat. Commun., 11, 1415, <a href="https://doi.org/10.1038/s41467-020-15195-y" target="_blank">https://doi.org/10.1038/s41467-020-15195-y</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Pedregosa et al.(2011)</label><mixed-citation>
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,  O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J.,  Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E.:  Scikit-learn: Machine learning in Python, J. Mach. Learn. Res., 12, 2825–2830, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Popoola et al.(2016)</label><mixed-citation>
Popoola, O. A., Stewart, G. B., Mead, M. I., and Jones, R. L.: Development of  a baseline-temperature correction methodology for electrochemical sensors and  its implications for long-term stability, Atmos. Environ., 147, 330–343, <a href="https://doi.org/10.1016/j.atmosenv.2016.10.024" target="_blank">https://doi.org/10.1016/j.atmosenv.2016.10.024</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Rai and Kumar(2018)</label><mixed-citation>
Rai, A. C. and Kumar, P.: Summary of air quality sensors and recommendations  for application, Ref. Ares, p. 65, available at: <a href="https://www.iscapeproject.eu/wp-content/uploads/2017/09/iSCAPE_D1.5_Summary-of-air-quality-sensors-and-recommendations-for-application.pdf" target="_blank">https://www.iscapeproject.eu/wp-content/uploads/2017/09/iSCAPE_D1.5_Summary-of-air-quality-sensors-and-recommendations-for-application.pdf</a> (last access: 1 November 2020), 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Rasmussen and Williams(2006)</label><mixed-citation>
Rasmussen, C. E. and Williams, C. K. I.: Gaussian Processes for Machine  Learning, MIT Press, Cambridge, Massachusetts, 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Runge et al.(2012)</label><mixed-citation>
Runge, J., Heitzig, J., Petoukhov, V., and Kurths, J.: Escaping the curse of  dimensionality in estimating multivariate transfer entropy, Phys. Rev.  Lett., 108, 258701, <a href="https://doi.org/10.1103/PhysRevLett.108.258701" target="_blank">https://doi.org/10.1103/PhysRevLett.108.258701</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Runge et al.(2019)</label><mixed-citation>
Runge, J., Nowack, P., Kretschmer, M., Flaxman, S., and Sejdinovic, D.: Detecting and quantifying causal associations in large nonlinear time series  datasets, Science Advances, 5, eaau4996, <a href="https://doi.org/10.1126/sciadv.aau4996" target="_blank">https://doi.org/10.1126/sciadv.aau4996</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Sadighi et al.(2018)</label><mixed-citation>
Sadighi, K., Coffey, E., Polidori, A., Feenstra, B., Lv, Q., Henze, D. K., and Hannigan, M.: Intra-urban spatial variability of surface ozone in Riverside, CA: viability and validation of low-cost sensors, Atmos. Meas. Tech., 11, 1777–1792, <a href="https://doi.org/10.5194/amt-11-1777-2018" target="_blank">https://doi.org/10.5194/amt-11-1777-2018</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Sayahi et al.(2020)</label><mixed-citation>
Sayahi, T., Garff, A., Quah, T., Lê, K., Becnel, T., Powell, K. M.,  Gaillardon, P. E., Butterfield, A. E., and Kelly, K. E.: Long-term  calibration models to estimate ozone concentrations with a metal oxide  sensor, Environ. Pollut., 267, 115363, <a href="https://doi.org/10.1016/j.envpol.2020.115363" target="_blank">https://doi.org/10.1016/j.envpol.2020.115363</a>,  2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Sherwen et al.(2019)</label><mixed-citation>
Sherwen, T., Chance, R. J., Tinel, L., Ellis, D., Evans, M. J., and Carpenter, L. J.: A machine-learning-based global sea-surface iodide distribution, Earth Syst. Sci. Data, 11, 1239–1262, <a href="https://doi.org/10.5194/essd-11-1239-2019" target="_blank">https://doi.org/10.5194/essd-11-1239-2019</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Spinelle et al.(2015)</label><mixed-citation>
Spinelle, L., Gerboles, M., Villani, M. G., Aleixandre, M., and Bonavitacola,  F.: Field calibration of a cluster of low-cost available sensors for air  quality monitoring. Part A: Ozone and nitrogen dioxide, Sensor. Actuat. B-Chem., 215, 249–257, <a href="https://doi.org/10.1016/j.snb.2015.03.031" target="_blank">https://doi.org/10.1016/j.snb.2015.03.031</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Spinelle et al.(2017)</label><mixed-citation>
Spinelle, L., Gerboles, M., Villani, M. G., Aleixandre, M., and Bonavitacola,  F.: Field calibration of a cluster of low-cost commercially available  sensors for air quality monitoring. Part B: NO, CO and CO2, Sensor. Actuat. B-Chem., 238, 706–715, <a href="https://doi.org/10.1016/j.snb.2016.07.036" target="_blank">https://doi.org/10.1016/j.snb.2016.07.036</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Tanzer et al.(2019)</label><mixed-citation>
Tanzer, R., Malings, C., Hauryliuk, A., Subramanian, R., and Presto, A. A.: Demonstration of a low-cost multi-pollutant network to quantify intra-urban  spatial variations in air pollutant source impacts and to evaluate  environmental justice, Int. J. Environ. Res. Pub. He., 16, 2523, <a href="https://doi.org/10.3390/ijerph16142523" target="_blank">https://doi.org/10.3390/ijerph16142523</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Vikram et al.(2019)</label><mixed-citation>
Vikram, S., Collier-Oxandale, A., Ostertag, M. H., Menarini, M., Chermak, C., Dasgupta, S., Rosing, T., Hannigan, M., and Griswold, W. G.: Evaluating and improving the reliability of gas-phase sensor system calibrations across new locations for ambient measurements and personal exposure monitoring, Atmos. Meas. Tech., 12, 4211–4239, <a href="https://doi.org/10.5194/amt-12-4211-2019" target="_blank">https://doi.org/10.5194/amt-12-4211-2019</a>, 2019.

</mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Zimmerman et al.(2018)</label><mixed-citation>
Zimmerman, N., Presto, A. A., Kumar, S. P. N., Gu, J., Hauryliuk, A., Robinson, E. S., Robinson, A. L., and R. Subramanian: A machine learning calibration model using random forests to improve sensor performance for lower-cost air quality monitoring, Atmos. Meas. Tech., 11, 291–313, <a href="https://doi.org/10.5194/amt-11-291-2018" target="_blank">https://doi.org/10.5194/amt-11-291-2018</a>, 2018.
</mixed-citation></ref-html>--></article>
