<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article"><?xmltex \makeatother\@nolinetrue\makeatletter?>
  <front>
    <journal-meta><journal-id journal-id-type="publisher">AMT</journal-id><journal-title-group>
    <journal-title>Atmospheric Measurement Techniques</journal-title>
    <abbrev-journal-title abbrev-type="publisher">AMT</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Atmos. Meas. Tech.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1867-8548</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/amt-16-3085-2023</article-id><title-group><article-title>A data-driven persistence test for robust (probabilistic)<?xmltex \hack{\break}?> quality control of measured environmental time series:<?xmltex \hack{\break}?> constant value episodes</article-title><alt-title>Probabilistic quality control of measured environmental time series</alt-title>
      </title-group><?xmltex \runningtitle{Probabilistic quality control of measured environmental time series}?><?xmltex \runningauthor{N. Kaffashzadeh}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name><surname>Kaffashzadeh</surname><given-names>Najmeh</given-names></name>
          <email>najmeh.kaffashzadeh@gmail.com</email>
        <ext-link>https://orcid.org/0000-0001-5835-9604</ext-link></contrib>
        <aff id="aff1"><institution>Institute of Geophysics, University of Tehran, Tehran, Iran</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Najmeh Kaffashzadeh (najmeh.kaffashzadeh@gmail.com)</corresp></author-notes><pub-date><day>21</day><month>June</month><year>2023</year></pub-date>
      
      <volume>16</volume>
      <issue>12</issue>
      <fpage>3085</fpage><lpage>3100</lpage>
      <history>
        <date date-type="received"><day>13</day><month>November</month><year>2022</year></date>
           <date date-type="rev-request"><day>18</day><month>January</month><year>2023</year></date>
           <date date-type="rev-recd"><day>25</day><month>April</month><year>2023</year></date>
           <date date-type="accepted"><day>16</day><month>May</month><year>2023</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2023 Najmeh Kaffashzadeh</copyright-statement>
        <copyright-year>2023</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023.html">This article is available from https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023.html</self-uri><self-uri xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023.pdf">The full text article is available as a PDF file from https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d1e82">Robust quality control is a prerequisite and an essential
component in any data application. That is especially important for time
series of environmental observations such as air quality due to their
dynamic and irreversible nature. One of the common issues in these data is
constant value episodes (CVEs), where a set of consecutive data values
remains constant over a given period. Although CVEs are often considered to be an indicator of sensor failure or other measurement errors and are removed during quality control procedures, there are situations when CVEs reflect natural environmental phenomena, and they should not be removed from the data or analysis. Assessing whether the CVEs are erroneous data or valid observations is a challenge. As there are no formal procedures established for this, their classification is based on subjective judgment and is therefore uncertain and irreproducible. This paper presents a novel test procedure, i.e., constant value test, to estimate the probability of CVEs being valid data. The theoretical foundation of this test is based on
statistical characteristics and probability theory and takes into account
the numerical precision of the data values. The test is a data-driven
(parametric) approach, which makes it usable for time series analysis in
different environmental research domains, as long as serial dependency is
given and the data distribution is not too different from Gaussian. The
robustness of the test was demonstrated with sensitivity studies using
synthetic data with different distributions. Example applications to
measured air temperature and ozone mixing ratio data confirm the versatility
of the test.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>European Research Council</funding-source>
<award-id>ERC-2017-ADG 787576</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

      <?xmltex \hack{\newpage}?>
<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e96">Millions of sensors monitor the environment every day, and their data are
used in many applications such as trend analysis (Fang et al., 2013; Mills
et al., 2016, 2018; Chang et al., 2017; Fleming et al., 2018; Lefohn et al., 2018) and forecasts (Gardner, 1999; Zhang et al., 2012; Debry and Mallet, 2014; Zhou et al., 2019) to provide important information on global challenges
such as climate change, air quality, soil degradation, etc. The measurement
process can be interpreted as sampling from a true distribution of
atmospheric state variables, for example, temperature or air pollutant
concentration, at a given location. Each measured value is an estimation of
“truth” that has been obtained through a set of data samples (Grant and
Leavenworth, 1996). A common feature of many environmental time series is
the fact that the true distribution changes with time. This makes such
measurements irreproducible.</p>
      <p id="d1e99">Measured data can be contaminated by various errors such as systematic,
random, non-representative and gross errors (Gandin, 1988; Steinacker et
al., 2011). These errors can arise from poor sensor calibration, long-term
sensor drift, noise, non-resolvable processes by an observational network,
and mistakes during data processing, decoding or transmission. Some of
these errors arise from unpredictable natural phenomena such as floods,
fire, frost and animal activities (Campbell et al., 2013) that cannot be
documented in every detail. Although many efforts are devoted to developing
advanced analytical tools and methods, these errors can have deleterious
effects on the statistical analyses. For instance, outliers, i.e., values
far outside of the norm for a variable or population, can increase the error
variance or reduce<?pagebreak page3086?> the power of statistical tests (Osborne and Overbay,
2004). Specifically, constant value episodes (CVEs) can decrease the
normality when the assumption of a normal distribution must be satisfied,
for example, in linear regression. Thus, even the most sophisticated
statistical model can be vulnerable against unknown and potentially erroneous
data. If such errors in the data are not identified by applying quality
control (QC) procedures, the information obtained from the data will be
misleading, and the results from scientific data analyses can be unreliable
and biased. Therefore, robust QC procedures are an essential component in
the data production chain and a requirement for having a more reliable
quantification of trend or other statistical analysis.</p>
      <p id="d1e102">Many research initiatives and environmental monitoring programs have thus
established standards and guidelines for QC procedures. Most of them rely on
visual screening of data, and therefore personal inspection, and on manual
elimination of erroneous values based on empirical knowledge and
investigator experiences. Several advanced tools such as GCE (Scully-Allison
et al., 2018), CoTeDe (Castelão, 2016), and AutoQC (Good et al., 2022) and
comprehensive user manuals such as QARTOD (Bushnell et al., 2019) and WMO-AWS
(Zahumenský, 2004) have been developed with precise rules to overcome this
subjectivity. However, their application is often limited to a few variables
or specific data sets, for example, from limited geographic regions with
relatively homogenous conditions. This, in turn, can be problematic if one
wants to assemble global data sets of various environmental variables. For
example, in the Tropospheric Ozone Assessment Report (TOAR), a global
database with ground-level ozone measurements at more than 10 000 locations
around the world, was built with data from more than 30 different
contributors (Schultz et al., 2017). Different QC procedures at these
agencies and sites led to increased uncertainty in the assessment. At this
scale of data, manual inspection methods are not only error prone but also
impractical. It is therefore desirable to develop a more generic, robust and
data-driven approach for the QC of environmental monitoring time series.</p>
      <p id="d1e105">The focus of this study is to develop a QC test for CVEs as the first
element for such data-driven QC. CVEs are a common feature in air quality
time series and other environmental data sets. As an example, in a specific
35-year-long ozone time series with hourly sampling, CVEs with a length of 2
occurred 20 313 times. Therefore, about 6.7 % of the data values are
CVEs, meaning that such incidents are expected to occur naturally about 16 times per 10 d in the hourly data. The CVEs with a longer length, e.g.,
3, 4 and 5, occur 6190, 2887 and 1681 times, respectively, and so the
proportion of these incidents are 4.85, 2.26 and 1.31 for 10 d hourly data time series. While they can be detected through a persistence test, a
qualified judgment whether such data are erroneous or not is a difficult
undertaking. If CVEs are excluded from the data (Horsburgh et al., 2015;
Gudmundsson et al., 2018), the results of the analysis, such as model–data
comparisons (Bey et al., 2001; Horowitz et al., 2003; Dawson et al., 2008;
Emmons et al., 2010; Lamarque et al., 2012; Rasmussen et al., 2012; Tilmes et
al., 2012; Im et al., 2015; Schnell et al., 2015; Lyapina et al., 2016;
Sofen et al., 2016), can become biased. That can be an issue in (re)analysis
products (Inness et al., 2019; Hersbach et al., 2020), where assimilation
processes reduce misfits between observations and their modeled values. If
the models correctly capture CVE events, excluding the CVEs will lead to
a type I error. On the other hand, if CVEs originating from instrument
malfunctions are included in the analysis, that will raise type I and type
II errors and likely raise unreliable results.</p>
      <p id="d1e109">This study presents a new (QC) test procedure, i.e., constant value test
(CVT), which estimates the probability of a CVE representing valid data.
Data users can select a threshold of an acceptable probability depending on
their scientific study or data analysis task. The CVT is entirely
data-driven and makes very few assumptions about the properties of the
underlying values' distribution and probability density function (Gaussian).
Currently, the method is valid for data with a Gaussian frequency
distribution. Possible extensions of the method are discussed in the
conclusions section. In principle, it is possible to use the technique of
statistical simulations to examine how the CVE probabilities change for
non-Gaussian distributions. However, this is beyond the scope of this paper.
Due to its generality, the test is applicable for a wide variety of
environmental variables with a serial dependency (autocorrelation). The
article structure is as follows: the method (CVT) is described in Sect. 2.
In Sect. 3, the approach is evaluated using synthetic data for demonstration
purposes. The results of three real test cases are discussed in Sect. 4. And,
finally, conclusions are given in Sect. 5.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methodology</title>
      <p id="d1e120">Before describing the proposed method, we briefly summarize some issues with
existing methods. In existing QC frameworks, the persistence test is
typically defined based on the minimum expected variability, but this
requires prior knowledge about the true statistical distribution of the
measurements. For example, Zahumenský (2004) has defined that air
temperature measurements shall be flagged as “doubtful or suspected
value” if the measured variable varies by less than 0.1 K over 60 min. Such a priori assumptions may lead to false data labeling when environmental conditions are exceptionally stable and the true data variability is reduced for some period of time. For instance, temperature variation of 0.1 K can occur in the morning when radiative forcing is small, e.g., on a foggy day in autumn. In measurements of air pollutant concentrations, longer periods of zero values can be found if the measured concentrations are below the instrument detection limit or if chemical conversion leads to a complete removal of a species. For example,
ground-level ozone concentrations at urban sites remain zero<?pagebreak page3087?> for several hours if there is a high level of nitrogen oxide emission.</p>
      <p id="d1e123">The assessment of CVEs will also have to depend on the numerical precision
or resolution (res), which is the number of significant digits with which an observation is recorded (Chapman, 2005). For example, historical
measurements of ground-level ozone (Azusa station) in the EPA Air Quality
System (AQS) in the 1980s were reported with a resolution of 8 ppb (parts per
billion). Another pollutant in the EPA AQS database for which reporting
precision has changed over time since 1980 is carbon monoxide at the Fresno
station (California state). So, it is not uncommon to find episodes of
several hours when all measurements are reported as the same value, and it
would be implausible to remove all of them as “erroneous measurements”.</p>
      <p id="d1e126">The CVT takes these considerations into account and provides a data-driven
approach with very few a priori assumptions. It consists of two main
procedures: first, CVEs need to be found and the length of the episodes must
be recorded, then, in the second step, the probability of each CVE being a
period of valid data with low variability is estimated. While the first
procedure can be simply implemented by taking the differences of consecutive
values, a possible complication arises if the time series contains missing
data or if the data were irregularly sampled. While the software
accompanying this paper has a provision to deal with missing data, we ignore
the second issue for the purpose of this paper and require that the time
series has been sampled at regular intervals. The following method
description focuses on the estimation of the likelihood that two or more
constant values occur in reality and are thus not necessarily resulting from
measurement or data processing errors.</p>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Statistical background</title>
      <p id="d1e136">To describe the joint process of a given time series, we assume such a
stochastic process can be represented as a multivariate Gaussian
distribution (Tong, 1990; Rencher, 2002). Let <inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> be a
series of random variables; the joint probability density function of a
multivariate Gaussian distribution, <inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mi mathvariant="script">N</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mi mathvariant="normal">Σ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, can be written as
            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M3" display="block"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi mathvariant="normal">Σ</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfenced></mml:mrow><mml:msqrt><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:msup><mml:mo>)</mml:mo><mml:mi>k</mml:mi></mml:msup><mml:mi mathvariant="normal">|</mml:mi><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi mathvariant="normal">|</mml:mi></mml:mrow></mml:msqrt></mml:mfrac></mml:mstyle><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          Here, <inline-formula><mml:math id="M4" display="inline"><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:math></inline-formula> is an <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> mean vector and <inline-formula><mml:math id="M6" display="inline"><mml:mi mathvariant="bold">Σ</mml:mi></mml:math></inline-formula> is an <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>×</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:math></inline-formula> positive definite covariance matrix. In the stationary case, without loss of generality, <inline-formula><mml:math id="M8" display="inline"><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:math></inline-formula> can be assumed to be a constant, and <inline-formula><mml:math id="M9" display="inline"><mml:mi mathvariant="bold">Σ</mml:mi></mml:math></inline-formula> can be represented as multiplication of a finite
constant variance <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and a (auto)correlation matrix <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mfenced close="}" open="{"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>;</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>, with <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> if <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula> (diagonal) and <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>≤</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mo>)</mml:mo><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> if <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mi>i</mml:mi><mml:mo>≠</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:math></inline-formula> (off-diagonal) for a given time series.</p>
      <p id="d1e457">Long range approximation of an environmental time series is generally
unnecessary and computationally expensive (e.g., Wincek and Reinsel, 1986;
Guttorp et al., 1994; Niu, 1996; Fioletov and Shepherd, 2003; Kumar and De
Ridder, 2010). Here we use an assumption that an environmental time series is
auto-correlated and can be approximated by an autoregressive (AR(1)) process
(Tiao et al., 1990; Weatherhead at al., 1998, 2000; Reinsel et al., 2002).
The definition of an AR(1) process, the <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, i.e., data value at time
<inline-formula><mml:math id="M17" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>, can be written as
            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M18" display="block"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mtext>const</mml:mtext><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
          Here, <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ε</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is a white noise, and const is an offset. With the assumption of the AR(1) process, the correlation matrix can be approximated by one parameter, <inline-formula><mml:math id="M20" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula>, since <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mtext>Corr</mml:mtext><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mrow><mml:mfenced close="|" open="|"><mml:mi>h</mml:mi></mml:mfenced></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> (the correlation between any two points is only dependent on the time interval, <inline-formula><mml:math id="M22" display="inline"><mml:mi>h</mml:mi></mml:math></inline-formula>); thus, the stochastic process can be governed by three parameters, i.e., <inline-formula><mml:math id="M23" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M25" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula>.</p>
      <p id="d1e603">The general likelihood of an AR(1) process can be approximated using the
first-order Markov property as
            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M26" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:msubsup><mml:mo>∏</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mi>p</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the density of initial state, which is not critical in this study, because the focus is placed on the probability of a consecutive state that is identical to the previous value, i.e., the second term, and <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula> represents the probability distribution of <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> depending only on <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. The above equation is a general form without a distributional assumption. To derive the explicit
form for the Gaussian case, we start from a univariate and a bivariate
probability density function:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M31" display="block"><mml:mtable rowspacing="5.690551pt" displaystyle="true"><mml:mlabeledtr id="Ch1.E4"><mml:mtd><mml:mtext>4</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi>f</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced open="[" close="]"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E5"><mml:mtd><mml:mtext>5</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>f</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mo mathsize="2.5em">(</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo mathsize="2.5em">[</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathsize="2.5em">]</mml:mo><mml:mo mathsize="2.5em">)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            Then the conditional probability distribution of <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> given <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:math></inline-formula>
can be derived by the Bayes' theorem and written as (see Appendix A)
            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M34" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>∼</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mi>N</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
          where <inline-formula><mml:math id="M35" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> is an arbitrary constant. The implication of such a formulation is that the resulting probability is also a function of <inline-formula><mml:math id="M36" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula>: if the statistical model parameters <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are fixed, a shorter distance of <inline-formula><mml:math id="M38" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula> from the mean, <inline-formula><mml:math id="M39" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, will result in a relatively higher probability density than those are far away.</p><?xmltex \hack{\newpage}?>
</sec>
<?pagebreak page3088?><sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Constant value episode (CVE) probability</title>
      <p id="d1e1209">The estimation of the CVT probability consists of the following two steps.</p>
      <p id="d1e1212"><italic>Step 1.</italic> <italic>Deriving a joint probability density.</italic> For a series of (dependent) events, <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>≤</mml:mo><mml:mi>k</mml:mi><mml:mo>≤</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:math></inline-formula>, the joint density of probability can be described through a product of multiple conditional probabilities as
            <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M42" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>p</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>∩</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>∩</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:msubsup><mml:mo>∏</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mi>p</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msubsup><mml:mo>∩</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi>A</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo><mml:msubsup><mml:mo>∏</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mi>p</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          The first equality yields from the chain rule of the joint distribution (Schum, 2001); the second equality is a special case of an AR(1) process.</p>
      <p id="d1e1392"><italic>Step 2.</italic> <italic>Imposing a distributional assumption to the joint probability distribution.</italic> From Eq. (6), the probability of consecutive values in a
series with Gaussian probability density can be determined by
            <disp-formula id="Ch1.E8" content-type="numbered"><label>8</label><mml:math id="M43" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mtext>CVE</mml:mtext><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>≠</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">res</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">res</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced open="[" close="]"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mfenced close=")" open="("><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          The integral reflects the fact that digital data are recorded with finite
numerical precision. Then, according to the property of an AR(1) process, the
probability of a CVE with a length of <inline-formula><mml:math id="M44" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> can be calculated through <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mtext>CVE</mml:mtext><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> raising to the power of <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> as
            <disp-formula id="Ch1.E9" content-type="numbered"><label>9</label><mml:math id="M47" display="block"><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi>P</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mtext>CVE</mml:mtext><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>≠</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mo mathsize="2.5em">(</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">res</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">res</mml:mi><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><?xmltex \hack{\hbox\bgroup\fontsize{9.5}{9.5}\selectfont$\displaystyle}?><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced open="[" close="]"><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mfenced close=")" open="("><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mfenced><mml:msup><mml:mo mathsize="2.5em">)</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          Since this equation is designed for a constant event, so the marginal
probability remains a constant for each CVE. To diminish the influence of
CVE on <inline-formula><mml:math id="M48" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, they were excluded first, then the <inline-formula><mml:math id="M49" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M50" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and
<inline-formula><mml:math id="M51" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> were calculated.</p>
      <p id="d1e1816">For non-normal cases, the explicit parameterization of a non-independent
joint distribution is difficult to derive due to mathematical challenges and
often does not have a closed form. The non-parametric alternative is to use
empirical distribution (Epanechnikov, 1969; Waterman and Whiteman, 1978) or
kernel distribution (Hwang et al., 1994; Duong and Hazelton, 2005), but this
approach is not desirable for database management at this stage because it
is difficult to develop a unified framework that is adequate for all
situations. Besides, the empirical distribution estimates a probability
without taking into account auto-correlation, i.e., independent of the
adjacent data points.</p>
      <p id="d1e1820">The AR(1) assumption can be relaxed by increasing the order of
autocorrelation without too much complexity. For example, for an AR(2)
process, one could specify the covariance matrix in Eq. (1) as
            <disp-formula id="Ch1.E10" content-type="numbered"><label>10</label><mml:math id="M52" display="block"><mml:mrow><mml:mi mathvariant="normal">Σ</mml:mi><mml:mo>=</mml:mo><mml:mfenced close="|" open="|"><mml:mtable class="matrix" columnalign="center center center" framespacing="0em"><mml:mtr><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mfenced></mml:mrow></mml:math></disp-formula>
          and modify Eq. (7) in step 1 as
            <disp-formula id="Ch1.E11" content-type="numbered"><label>11</label><mml:math id="M53" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mi>p</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>∩</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>∩</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mfenced><mml:mi>p</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msub><mml:mi>A</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:mfenced><mml:msubsup><mml:mo>∏</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mi>p</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">|</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
          then update the conditional probability parameterized by <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> in step 2. The more general extension of the autoregressive model is out of the scope of this study and can be referred to in Box et al. (2015).</p>
      <p id="d1e2078">For the variables with extra incidences of zero such as nitrogen oxides (NO,
NO<inline-formula><mml:math id="M55" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula>) and ozone, the lower interval of the integration in Eq. (9) was
changed from <inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo></mml:mrow></mml:math></inline-formula> res to 0. Note that in reality “zero” values in measurements may actually be recorded as small positive or negative numbers. This detail is ignored in the following because there is no universally applicable correction available. Some data sets may require a linear or non-linear bias correction, while for other data sets a simple cutoff, e.g., set to zero if <inline-formula><mml:math id="M57" display="inline"><mml:mo>|</mml:mo></mml:math></inline-formula> value <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mo>&lt;</mml:mo></mml:mrow></mml:math></inline-formula> threshold, may be more appropriate.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Model sensitivity test</title>
      <p id="d1e2126">The <inline-formula><mml:math id="M59" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> in Eq. (9) is affected by the parameters <inline-formula><mml:math id="M60" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M61" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M62" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M63" display="inline"><mml:mi>c</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M64" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> and res. A simulation study was developed to evaluate the sensitivity of <inline-formula><mml:math id="M65" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> to each parameter. Several experiments were conducted by generating a synthetic data series to demonstrate the influence of each parameter. For each experiment, the CVT was performed over a range of possible values.</p>
      <p id="d1e2179">A set of first-order autoregressive, AR(1), time series with hourly time
steps and a length of 240 values (10 d) was generated using Eq. (2) and a
random noise generator. As a reference case (ref), we set <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>. The numerical precision was defined as 0.01. Four sets of CVEs with the same length (<inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>) were added to this time
series. The distance of the CVE from the mean, i.e., <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></inline-formula>, was given as
0, 1, 2 and 3<inline-formula><mml:math id="M71" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> (see Fig. 1). In this figure, four CVEs are illustrated with a color code, i.e., red, blue, cyan and black, which are
shown with boxes. The <inline-formula><mml:math id="M72" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> varies from <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mn mathvariant="normal">7.67</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for the first
CVE to <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.77</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for the fourth (last) CVE. As stated in
Sect. 2.1, the value of <inline-formula><mml:math id="M75" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> decreases as <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></inline-formula> increases. CVEs which are
further away from the mean are less likely to occur in nature.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><?xmltex \def\figurename{Figure}?><label>Figure 1</label><caption><p id="d1e2314">A synthetic AR(1) time series with Gaussian data distribution and
four arbitrarily selected CVEs of length <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>, with <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, 4, 8, and 12, respectively. The CVEs are shown using a color code, i.e., red, blue, cyan and black. The numerical precision (res) is chosen as 0.01.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f01.png"/>

      </fig>

      <?pagebreak page3089?><p id="d1e2388">To assess the effect of <inline-formula><mml:math id="M82" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> on <inline-formula><mml:math id="M83" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>, a set of values ranging from 2 to 10 were
selected for the <inline-formula><mml:math id="M84" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. All other parameters were fixed as in the baseline time
series. As expected from Eq. (9), the <inline-formula><mml:math id="M85" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> decreases exponentially with <inline-formula><mml:math id="M86" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>  (Fig. B1a). Note that the slope of this exponential decrease depends on <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></inline-formula>. The larger the <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></inline-formula>, the larger would be the slope. That is in agreement with Fig. 1, where the <inline-formula><mml:math id="M89" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> decreases as the CVE gets further from the mean. However, the probability of finding two consecutive data points with the same value is about <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">300</mml:mn></mml:mrow></mml:math></inline-formula>, i.e., in a year-long time series such incidents are expected to occur naturally about once per year if the sampling resolution is daily and about 25 times if the sampling resolution is hourly.</p>
      <p id="d1e2470">To investigate the non-linear influence of <inline-formula><mml:math id="M91" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> on <inline-formula><mml:math id="M92" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> in Eq. (9), a range of values, i.e., 0.1, 0.2, 0.3, 0.4, 0.5, 1, 2, 3, 4, 5, 10 and 20, were set as <inline-formula><mml:math id="M93" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>, while other parameters remained unchanged. In this scenario, the <inline-formula><mml:math id="M94" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> changes from <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.22</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for the smallest <inline-formula><mml:math id="M96" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> to
<inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:mn mathvariant="normal">8.93</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> for the largest one (Fig. B1b). By using Eq. (9), it thus becomes possible to estimate likelihoods for naturally occurring CVEs for data sets with different variability, in contrast to classical approaches, which use a fixed variability threshold.</p>
      <p id="d1e2545">The most interesting parameter to consider in the CVT is the lag-1
auto-correlation (<inline-formula><mml:math id="M98" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula>). A sensitivity experiment with several
additional time series was performed to assess the sensitivity of <inline-formula><mml:math id="M99" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> with
respect to <inline-formula><mml:math id="M100" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> (Fig. B1c). In this figure, <inline-formula><mml:math id="M101" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> ranges from <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mn mathvariant="normal">1.23</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.5</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. The larger the <inline-formula><mml:math id="M104" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> (i.e., stronger persistence), the larger would be the probability of naturally occurring CVEs. The estimated probability is very
sensitive to <inline-formula><mml:math id="M105" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> as it approaches 1. At the limit value of 1, Eq. (9) is undefined. If <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, the time series only consists of noise, so it is less probable to get any CVEs.</p>
      <p id="d1e2639">Another parameter influencing <inline-formula><mml:math id="M107" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> is the data digital resolution (res) or
precision, where the data have been recorded in a fixed numerical precision
(number of decimals) or as integers with possible rounding to the nearest
multiple of 5, 10, etc. This parameter is shown in Eq. (9), where the
resulting probability is integrated over the range of values from <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo></mml:mrow></mml:math></inline-formula> res<inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>+</mml:mo></mml:mrow></mml:math></inline-formula> res<inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:mo>/</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d1e2689">To investigate the sensitivity of the <inline-formula><mml:math id="M112" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> to the res parameter, the baseline time series was resampled by using several resolutions, i.e., 0.0001, 0.0002, 0.0005, 0.001, 0.002, 0.005, 0.01, 0.02, 0.05, 0.1, 0.2, 0.5, 1, 2 and 5. As shown in Fig. B2a for the example of res <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>, larger res leads to additional CVEs, and it becomes harder to distinguish valid episodes from erroneous incidents. But here, to isolate influence of res on <inline-formula><mml:math id="M114" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>, first the data were truncated to a new resolution, then the CVEs were added to the data. The CVT results are shown in Fig. B2b, in which the <inline-formula><mml:math id="M115" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> changes from
<inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.77</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">11</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:mn mathvariant="normal">7.57</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>. This shows that by increasing the res, the <inline-formula><mml:math id="M118" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> increases, meaning that if the data are recorded in a coarse resolution, there is a higher chance to count those data as valid data.</p>
      <p id="d1e2767">An experiment with several scaling factors, i.e., fc <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, 0.2, 0.5, 1, 2, 5 and 10, was performed to check the robustness of the CVT to the
different data transformations. In this experiment, the CVEs were added
first; then the scaling, i.e., <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:mi>x</mml:mi><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo><mml:mo>×</mml:mo></mml:mrow></mml:math></inline-formula> fc, was applied; and the data were truncated to a new numerical resolution given by res <inline-formula><mml:math id="M121" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> fc. Scaling changes other parameters such as <inline-formula><mml:math id="M122" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula> or <inline-formula><mml:math id="M123" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>, except <inline-formula><mml:math id="M124" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula>, which remains invariant. Figure B1d shows the robustness of the CVT output (<inline-formula><mml:math id="M125" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>) with scaling. It is important to note that Eq. (9) is robust to the other data transformation such as normalization and standardization (see Appendix C).</p>
      <p id="d1e2833">A combined sensitivity analysis was performed to illustrate the effect of
the parameters <inline-formula><mml:math id="M126" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M127" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> and res in Eq. (9), i.e., the
conditional probability for two consecutive values was evaluated over a
range of conditions (<inline-formula><mml:math id="M128" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M129" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> from 0.01 to 0.99, and res of
0.01, 0.1 and 0.5), with <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>-</mml:mo><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>. The results are shown in Fig. 2 and can be interpreted as an upper limit for <inline-formula><mml:math id="M131" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> that two successive values are valid data because <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>-</mml:mo><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula> represents the maximum of the Gaussian distribution in Eq. (9). Using the chain rule from Eq. (11), these results can easily be extrapolated to longer CVEs. As Fig. 2 shows, the probability of finding two valid consecutive data points with the same value decreases rather quickly with increasing standard deviation <inline-formula><mml:math id="M133" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>. The <inline-formula><mml:math id="M134" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> has limited influence up to values of around 0.7. Above this threshold, the likelihood of a two-value CVE increases drastically. A coarser numerical resolution makes it more likely to encounter constant values in reality. At res similar to <inline-formula><mml:math id="M135" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>, the length, <inline-formula><mml:math id="M136" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, of the<?pagebreak page3090?> CVE will have to be much larger than 2 to reliably classify it as erroneous. In practical applications, one would generally set a threshold for the acceptable probability first. The information provided in Fig. 2 can then help to identify typical parameters of the time series, where this threshold will be reached.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><?xmltex \def\figurename{Figure}?><label>Figure 2</label><caption><p id="d1e2934">Conditional probabilities to find a measured value <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> given <inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> for three different numerical resolutions, i.e., <bold>(a)</bold> res <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.01</mml:mn></mml:mrow></mml:math></inline-formula>, <bold>(b)</bold> res <inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula> and <bold>(c)</bold> res <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>. In this figure, the <inline-formula><mml:math id="M142" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M143" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> range from 0.01 to 0.99.</p></caption>
        <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f02.png"/>

      </fig>

</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results and discussion</title>
      <p id="d1e3033">Two data time series were retrieved from the Tropospheric Ozone Assessment
Report (TOAR) database (Schultz et al., 2017) to illustrate the practical
use of the CVT. This database holds in situ measured data time series for
ground-level ozone in hourly time resolutions. We selected the time series of
ozone mixing ratio at the Azusa station (34<inline-formula><mml:math id="M144" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>8<inline-formula><mml:math id="M145" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> N,
117<inline-formula><mml:math id="M146" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>55<inline-formula><mml:math id="M147" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> W) in California that has data from the 1980s,
when the data were recorded with a resolution of 8 ppb. Besides this, the TOAR
database contains data for meteorological variables at some stations. We
selected one temperature time series at the Cape Grim station, Tasmania
(40<inline-formula><mml:math id="M148" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>68<inline-formula><mml:math id="M149" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> S, 144<inline-formula><mml:math id="M150" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>69<inline-formula><mml:math id="M151" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> E). This station is
located at an altitude of 94 m directly on the coast, and it is a Southern
Hemisphere background site with an extensive record back into 1980. The
station primarily measures air which has passed over the Southern Ocean for
several days. So, temperature variations at this site are often of small
amplitude. Data series of carbon monoxide at the Fresno station
(36.78<inline-formula><mml:math id="M152" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 119.77<inline-formula><mml:math id="M153" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W) were obtained from the EPA AQS
database. This data was reported with a precision of 1 ppm in 1980 and later
in 2022 with a higher precision of 0.001 ppm. The precision changes might
have arisen from the method detection limits (0.5 and 0.001), measurement
methods (instrumental non-dispersive infrared and instrumental gas filter
correlation Teledyne API 300 EU) or method types (non-FRM and FRM; Federal Reference Method) detailed
in their data files.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Temperature</title>
      <p id="d1e3134">Temperature is one of the key variables relevant to air quality research.
For example, temperature is often used as a primary predictor for
smog-related air quality. For demonstration of the CVT in a real data
situation, 10 d of a temperature time series were selected. The <inline-formula><mml:math id="M154" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>,
<inline-formula><mml:math id="M155" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M156" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> of the selected 10 d time series are 12.55,
1.59 and 0.94, respectively. The recorded numerical resolution of the data
is 0.01. The time series along with the probability, <inline-formula><mml:math id="M157" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>, of each value being a valid observation is shown in Fig. 3. Altogether, 18 CVEs are visible in
Fig. 3, 15 of them with <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, 2 with <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> and 1 with <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><?xmltex \def\figurename{Figure}?><label>Figure 3</label><caption><p id="d1e3204">Temperature time series at the Cape Grim station (40<inline-formula><mml:math id="M161" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>68<inline-formula><mml:math id="M162" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> S, 144<inline-formula><mml:math id="M163" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>69<inline-formula><mml:math id="M164" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> E) from 10 to 20 January 1983. Black and blue lines show the temperature value (<inline-formula><mml:math id="M165" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C) and its associated probability, <inline-formula><mml:math id="M166" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>, in Eq. (9), respectively. In this figure, the time is shown in UTC. The <inline-formula><mml:math id="M167" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> is not affected by the unit conversion, i.e., degrees Celsius (<inline-formula><mml:math id="M168" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C) to kelvin (K). The data were retrieved from the TOAR database.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f03.png"/>

        </fig>

      <p id="d1e3282">The CVEs occur at more or less regular times in the early morning, e.g., 04,
05 and nighttime hours, e.g., 10, 21, 22 and 23 (see Fig. 4). That can be
because of the local meteorological phenomena at this site, where the
temperature has little variance. Therefore, these CVEs are less likely to be
erroneous data.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4"><?xmltex \currentcnt{4}?><?xmltex \def\figurename{Figure}?><label>Figure 4</label><caption><p id="d1e3288">The number of CVEs occurring for the different hours in a day,
i.e., <inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:mfenced close="}" open="{"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mn mathvariant="normal">23</mml:mn></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>, for the temperature time series shown in Fig. 3.</p></caption>
          <?xmltex \igopts{width=142.26378pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f04.png"/>

        </fig>

      <p id="d1e3315">The probabilities estimated by the CVT are above 0.2 in most cases, which
means that if the CVEs were to be flagged as erroneous data, one would err
in one out of five cases and throw out the valid measurements. The CVE on
18 January yields the lowest probability (0.008), in line with the
expectation of the human data analyst because it is a sparse CVE with four
consecutive values (<inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>). This example illustrates that it will generally
be impossible to define a universal threshold for <inline-formula><mml:math id="M171" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> but that instead it depends on the use case. For example, in a data quality control workflow at the originating institution, one may decide to rule out data with <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>&lt;</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> but have a data curator cross-check the measurements with larger
<inline-formula><mml:math id="M173" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>. In contrast, when these data are integrated in a larger analysis
consisting of many stations, one might apply the CVT to rule out data with
<inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>&lt;</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> or even <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:mi>P</mml:mi><mml:mo>&lt;</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> to increase the statistical robustness of the analysis.</p>
      <p id="d1e3399">Other criteria for selecting a threshold for <inline-formula><mml:math id="M176" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> could be climate regions. In
the polar regions, the diurnal cycle of the temperature in summer could be quite
high, but coastal sites in that area with a dense fog might have morning
periods when the temperature is rather constant. The first shows a larger
<inline-formula><mml:math id="M177" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> than the latter, so the <inline-formula><mml:math id="M178" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> will be less in the polar than the
coastal sites, assuming all other parameters are constant (as shown in Fig. B1b). One may adopt a smaller threshold for <inline-formula><mml:math id="M179" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> in polar than in coastal sites. Or for the same climatological region with constant temperature
values at night or in the day, when the diurnal cycle reaches maximum or
minimum, the CVT would give CVEs a lower probability, as they are further
from the mean (larger <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:math></inline-formula>). So, the <inline-formula><mml:math id="M181" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> of the CVEs at extrema can be less than the CVEs with the same <inline-formula><mml:math id="M182" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> in this series.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Ozone</title>
      <p id="d1e3465">Ozone near the ground is an air pollutant that is detrimental to human
health and vegetation growth. Ozone measurement techniques have evolved over
time, and it can therefore be challenging to assess the data quality of a
decade-long monitoring data set, such as that from the Azusa station in
California, U.S. (34<inline-formula><mml:math id="M183" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>8<inline-formula><mml:math id="M184" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> N, 117<inline-formula><mml:math id="M185" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>55<inline-formula><mml:math id="M186" display="inline"><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup></mml:math></inline-formula> W), that contains a relatively long data record from 1980 to 2016.</p>
      <?pagebreak page3091?><p id="d1e3504">Figure 5 shows a 10 d example from this measurement series for the year
1990, with the <inline-formula><mml:math id="M187" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M188" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M189" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> of 16.55, 17.32 and 0.79, respectively. During the early period, the data were reported in a low
resolution, here an interval of 8 ppb. As a consequence, the time series
contains many CVEs, and most of them are probably valid. In contrast, for the
year 2012 when the data are recorded in a higher data resolution, i.e., 1 ppb, the number of the CVE is small (see Fig. D1). As mentioned in the
introduction, urban ozone time series often show very low values
(effectively zero), which are, however, recorded as small positive or negative values, here <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula> ppb. Figure 5 shows the probabilities between
<inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:mn mathvariant="normal">3.12</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M192" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula> for these episodes, which have values of 2 ppb. There are also three CVEs, with large <inline-formula><mml:math id="M193" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> (<inline-formula><mml:math id="M194" display="inline"><mml:mrow><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>) and very low
ozone mixing ratios of 2 ppb, which are shown with red circles in Fig. 5.
This illustrates the issue of zero-bounded data mentioned in the
methodology. The CVT can recognize such cases, and the associated
probabilities are <inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:mn mathvariant="normal">3.12</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.22</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.48</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, for the CVE1, CVE2 and CVE3, respectively. That
would prevent such (valid) values from being flagged or filtered as
erroneous data, in contrast to the second part of the time series in Fig. 6
(for the year 2011), which exhibits sparse occurrence of episodes, i.e., 21 CVEs where 17, 2, 1 and 1 CVEs with the <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, 4, 7 and 9, respectively. In most cases (17 episodes), the CVEs consist of only two consecutive values (<inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>). The estimated probability for these cases is between <inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.15</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:mn mathvariant="normal">9.9</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula> (Fig. 6). One episode during 18 November 2011 consists of nine constant values of 2 ppb. The estimated <inline-formula><mml:math id="M202" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> for that incident is <inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:mn mathvariant="normal">4.6</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, and this episode would indeed raise the suspicions of trained data analysts because such a pattern in the data would require a rather special explanation (see Fig. D3).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><?xmltex \def\figurename{Figure}?><label>Figure 5</label><caption><p id="d1e3723">Time series of the ozone mixing ratio at the Azusa station, California, from 10 to 20 November 1990 (black) and the CVT test results (blue). During this period, the data were recorded in intervals of 8 ppb, i.e., res <inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>, so that valid CVEs are frequent. In total, this time series contains 45 CVEs as 27, 6, 3, 3, 1, 1, 1 and 1 episode, with the <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, 3, 4, 5, 6, 8, 9, and 11, respectively. The red circles (or ovals) highlight three examples of zero-ozone incidents (here 2 ppb) with a large length (<inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>) in this series. The cyan circles highlight the probability of the respective CVEs. The orange circle highlights a CVE with a length of 4 that contain a gap of missing data points.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f05.png"/>

        </fig>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><?xmltex \def\figurename{Figure}?><label>Figure 6</label><caption><p id="d1e3769">As Fig. 5 but from 10 to 20 November 2011, when the data were recorded with a numerical resolution of 1 ppb, i.e., res <inline-formula><mml:math id="M207" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. The red circle shows one example of missing data points in the data time series. The <inline-formula><mml:math id="M208" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M209" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M210" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> of the data in this figure are 19.9, 10.73 and 0.84, respectively.</p></caption>
          <?xmltex \igopts{width=369.885827pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f06.png"/>

        </fig>

      <p id="d1e3809">Figure 5 also illustrates the problem with missing data values that was
mentioned in the beginning of Sect. 2. On 18 November, there is a gap in the time series where the data point has been excluded, and the
values to the left and right<?pagebreak page3092?> of this episode are identical. If these values
were not treated correctly, they would be counted as a CVE episode with a
length of 8 and probability of <inline-formula><mml:math id="M211" display="inline"><mml:mrow><mml:mn mathvariant="normal">2.58</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, which is shown with an orange circle in Fig. 5. Although such incidents could raise suspicions, they are not (and should not be) detected by the CVT. An independent test needs to be designed for such situations.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Carbon monoxide</title>
      <p id="d1e3838">Exposure to elevated carbon monoxide harms the human body, in particular
those who suffer from heart diseases. This air pollutant also affects some
greenhouse gases, e.g., carbon dioxide and ozone, which are linked to
climate change and global warming. A 10 d example of the measured carbon
monoxide at the Fresno station is shown in Fig. 7. Despite the high precision
of the data for the year 2022 (res <inline-formula><mml:math id="M212" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.001</mml:mn></mml:mrow></mml:math></inline-formula>, see Fig. D4), data were recorded with a resolution of 1 ppm in 1980. These data contain fewer CVEs but with a larger <inline-formula><mml:math id="M213" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> (19 CVEs with <inline-formula><mml:math id="M214" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mn mathvariant="normal">34</mml:mn></mml:mrow></mml:math></inline-formula>) in comparison to the ozone series in Fig. 5. That could be associated with a longer lifetime of carbon
monoxide than that of ozone. This reflects that most of the CVEs in the
carbon monoxide series are valid. The CVT discerns this and estimates a
larger <inline-formula><mml:math id="M215" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> for this data, in which the smallest <inline-formula><mml:math id="M216" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> is 0.001 for the CVEs, with <inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">14</mml:mn></mml:mrow></mml:math></inline-formula> and values of 0 ppm.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><?xmltex \def\figurename{Figure}?><label>Figure 7</label><caption><p id="d1e3903">Time series of carbon monoxide at the Fresno station, California,
from 1 to 11 January 1980 (black) and the CVT test results (blue). During this period, the data were recorded in intervals of 1 ppm, i.e., res <inline-formula><mml:math id="M218" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, so that valid CVEs are frequent. In total, this time series contains 19 CVEs as 1, 1, 1, 1, 2, 2, 1, 2, 1, 1, 1, 3 and 2 episodes with the <inline-formula><mml:math id="M219" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">34</mml:mn></mml:mrow></mml:math></inline-formula>, 27, 21, 18, 15, 14, 12, 11, 10, 5, 4, 3 and 2, respectively. The <inline-formula><mml:math id="M220" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M221" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M222" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> of the data in this figure are 0.79, 0.45 and 0.65, respectively.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f07.png"/>

        </fig>

</sec>
</sec>
<?pagebreak page3093?><sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d1e3964">Environmental time series are valuable and essential data sources for
scientific assessment of air quality and climate change. One of the issues
in these data is the occurrence of the constant value episodes (CVEs). These
episodes are often considered to be indicative of sensors' malfunctions or
other measurement errors and are excluded from the data via quality control
(QC) procedures. However, these episodes can be due to the natural
environmental phenomena, and they are indeed valid observations. Thus,
distinguishing whether the CVEs are erroneous or valid data is accompanied by large uncertainty.</p>
      <p id="d1e3967">This study presented a theoretical concept and evaluation for a data-driven
constant value test (CVT), which takes into account the typical evolution of
environmental state variables such as air temperature, ozone mixing ratio
or carbon monoxide as time series with serial dependence. Based on the
calculus of a marginal, joint and conditional Gaussian probability density,
one can estimate the probability of constant value episodes (CVEs) of length
<inline-formula><mml:math id="M223" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> to occur in reality and use this information to flag data as potentially
erroneous. The threshold for such flagging needs to be selected by the data
analyst. Together with the batch size for processing pieces of the time
series (in our examples, the full length of the depicted data was used; for
practical applications on longer time series, we recommend sample sizes in
the order of 100), these are the only a priori parameters needed. Examples
with synthetic and real data demonstrate that the CVT captures many aspects
which a trained data analyst would consider in the QC of such time series.
But as a data-driven approach, it will reveal data inconsistencies (here,
CVEs due to measurement or data processing errors) in automated data
processing workflows, and it may assist manual data quality control by
making it possible to provide a fine-grained warning to the data analyst
that something may be wrong with the measurements based on a probabilistic
score.</p>
      <p id="d1e3977">The test first detects CVEs by testing for zero difference. Then, it evaluates the distribution parameters mean (<inline-formula><mml:math id="M224" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>), standard deviation
(<inline-formula><mml:math id="M225" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula>) and lag-1 auto-correlation (<inline-formula><mml:math id="M226" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula>), as well as the numerical
resolution of the data in user-defined portions (batches) of the time
series. Given these parameters, the conditional probability for two
consecutive identical values is computed and integrated over the interval
given by the numerical resolution of the recorded data. Using the chain rule
for the non-independent conditional probability, this probability can easily
be scaled to arbitrary lengths of CVEs.</p>
      <p id="d1e4001">The novelty of this approach is its foundation in statistical theory and the
concept of estimating the probability of a data sample to occur naturally.
This distinguishes the method from classical approaches where more or less
arbitrary thresholds need to be defined prior to testing. Such pre-defined
thresholds can be dangerous if conditions change, for example, when the same
thresholds are applied to data from different world regions, climatic zones
or seasons. The method is robust against such changes, and its application
requires little background knowledge about the specific data set under
investigation. The method is therefore well suited for having robust and
automated QC systems, for example, in smart sensor networks, where human
intervention is not feasible.</p>
</sec>

      
      </body>
    <back><app-group>

<?pagebreak page3094?><app id="App1.Ch1.S1">
  <?xmltex \currentcnt{A}?><label>Appendix A</label><title/>
      <p id="d1e4014">The inference of conditional probability of bivariate normal distribution
          <disp-formula id="App1.Ch1.S1.Ex1"><mml:math id="M227" display="block"><mml:mtable class="split" rowspacing="0.2ex 0.2ex 0.2ex 0.2ex 0.2ex 4.267913pt" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mi>f</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><?xmltex \hack{\hbox\bgroup\fontsize{7.7}{7.7}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><mml:mfenced close="]" open="["><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced open="[" close="]"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><?xmltex \hack{\hbox\bgroup\fontsize{8.5}{8.5}\selectfont$\displaystyle}?><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mo mathsize="2.5em">(</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo mathsize="1.1em">(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo mathsize="1.1em">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathsize="2.5em">[</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><?xmltex \hack{$\egroup}?></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathsize="2.5em">]</mml:mo><mml:mo mathsize="2.5em">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:msqrt><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="italic">π</mml:mi><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>exp⁡</mml:mi><mml:mo mathsize="2.5em">(</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mfenced open="(" close=")"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathsize="2.5em">[</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>+</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced></mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathsize="2.5em">]</mml:mo><mml:mo mathsize="2.5em">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mo>∼</mml:mo><mml:mi>N</mml:mi><mml:mo>(</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∅</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi></mml:mrow></mml:mfenced><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mi mathvariant="normal">∅</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:msup><mml:mi mathvariant="italic">σ</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mtext>given</mml:mtext><mml:mspace linebreak="nobreak" width="0.25em"/><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</app>

<app id="App1.Ch1.S2">
  <?xmltex \currentcnt{B}?><label>Appendix B</label><title/>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S2.F8"><?xmltex \currentcnt{B1}?><?xmltex \def\figurename{Figure}?><label>Figure B1</label><caption><p id="d1e4733">Sensitivity of <inline-formula><mml:math id="M228" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> to the <bold>(a)</bold> CVEs length, i.e., <inline-formula><mml:math id="M229" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>, 3, 4, 5, 6, 7, 8, 9 and 10. Other parameters are fixed as <inline-formula><mml:math id="M230" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M231" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, 4, 8 and 12. <bold>(b)</bold>
Standard deviation, i.e., <inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, 0.2, 0.3, 0.4, 0.5, 1, 2, 3,
4, 5, 10 and 20. Other parameters are fixed as <inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>,
<inline-formula><mml:math id="M237" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, 4, 8 and 12. <bold>(c)</bold> Lag-1
autocorrelation, i.e., <inline-formula><mml:math id="M239" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.91, 0.92, 0.93, 0.94, 0.5, 0.96, 0.97, 0.98 and 0.99. Other parameters are fixed as <inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M241" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M242" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M243" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, 4, 8 and 12. <bold>(d)</bold> Sensitivity of <inline-formula><mml:math id="M244" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> to scaling factor, i.e., fc <inline-formula><mml:math id="M245" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>, 0.2, 0.5, 1, 2, 5 and 10. Other parameters are fixed as <inline-formula><mml:math id="M246" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M247" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula>. The same color codes are applied as those in Fig. 1.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=284.527559pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f08.png"/>

      </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S2.F9"><?xmltex \currentcnt{B2}?><?xmltex \def\figurename{Figure}?><label>Figure B2</label><caption><p id="d1e5003"><bold>(a)</bold> The modified time series (res <inline-formula><mml:math id="M248" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula>) where ref time series were
resampled with rounding to the nearest of five. That includes more CVEs than
the ref in Fig. 1. <bold>(b)</bold> Sensitivity of <inline-formula><mml:math id="M249" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> to the digital numerical precision, i.e., res <inline-formula><mml:math id="M250" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.0001</mml:mn></mml:mrow></mml:math></inline-formula>, 0.0002, 0.0005, 0.001, 0.002, 0.005, 0.01, 0.02, 0.05, 0.1, 0.2, 0.5, 1, 2 and 5. Other parameters are fixed as <inline-formula><mml:math id="M251" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M252" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">4</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M253" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M254" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M255" display="inline"><mml:mrow><mml:mi>c</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>, 4, 8 and 12. The same color codes are applied as those in Fig. 1.</p></caption>
        <?xmltex \igopts{width=241.848425pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f09.png"/>

      </fig>

</app>

<?pagebreak page3095?><app id="App1.Ch1.S3">
  <?xmltex \currentcnt{C}?><label>Appendix C</label><title/>
      <p id="d1e5116">If the data are normalized, i.e., <inline-formula><mml:math id="M256" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mo>min⁡</mml:mo></mml:msub><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mo>max⁡</mml:mo></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mo>min⁡</mml:mo></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></p>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S3.F10"><?xmltex \currentcnt{C1}?><?xmltex \def\figurename{Figure}?><label>Figure C1</label><caption><p id="d1e5157">As Fig. 1 but the data time series are normalized, <inline-formula><mml:math id="M257" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula>,
<inline-formula><mml:math id="M258" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.15</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M259" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula> and res <inline-formula><mml:math id="M260" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.004</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f10.png"/>

      </fig>

      <p id="d1e5214">If the data are standardized, i.e., <inline-formula><mml:math id="M261" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi mathvariant="italic">σ</mml:mi></mml:mrow></mml:math></inline-formula></p>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S3.F11"><?xmltex \currentcnt{C2}?><?xmltex \def\figurename{Figure}?><label>Figure C2</label><caption><p id="d1e5239">As Fig. 1 but the data time series are standardized, <inline-formula><mml:math id="M262" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.07</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M263" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.94</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M264" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula> and res <inline-formula><mml:math id="M265" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.002</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f11.png"/>

      </fig>

<?xmltex \hack{\clearpage}?>
</app>

<?pagebreak page3096?><app id="App1.Ch1.S4">
  <?xmltex \currentcnt{D}?><label>Appendix D</label><title/>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S4.F12"><?xmltex \currentcnt{D1}?><?xmltex \def\figurename{Figure}?><label>Figure D1</label><caption><p id="d1e5309">Time series of the ozone mixing ratio at the Azusa station, California, from 10 to 20 November 2011. During this period, the data were recorded in intervals of 1 ppb, i.e., res <inline-formula><mml:math id="M266" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. <inline-formula><mml:math id="M267" display="inline"><mml:mrow><mml:mi mathvariant="italic">μ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">19.9</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M268" display="inline"><mml:mrow><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">10.73</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M269" display="inline"><mml:mrow><mml:mi mathvariant="normal">∅</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.84</mml:mn></mml:mrow></mml:math></inline-formula>.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f12.png"/>

      </fig>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S4.F13"><?xmltex \currentcnt{D2}?><?xmltex \def\figurename{Figure}?><label>Figure D2</label><caption><p id="d1e5368">As Fig. 6, but the missing values are not treated. So, the orange circle shows two CVEs, which have been merged to one incident with a longer length (<inline-formula><mml:math id="M270" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">8</mml:mn></mml:mrow></mml:math></inline-formula>).</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f13.png"/>

      </fig>

      <?xmltex \floatpos{h!}?><fig id="App1.Ch1.S4.F14"><?xmltex \currentcnt{D3}?><?xmltex \def\figurename{Figure}?><label>Figure D3</label><caption><p id="d1e5394">Number of CVEs (<inline-formula><mml:math id="M271" display="inline"><mml:mrow><mml:mo>∑</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:math></inline-formula>) of different length, i.e., <inline-formula><mml:math id="M272" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mfenced close="}" open="{"><mml:mrow><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">9</mml:mn></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>, for the ozone time series of the year 2011 (shown in Fig. 6).</p></caption>
        <?xmltex \igopts{width=142.26378pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f14.png"/>

      </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S4.F15"><?xmltex \currentcnt{D4}?><?xmltex \def\figurename{Figure}?><label>Figure D4</label><caption><p id="d1e5438">As Fig. 7 but from 1 to 11 January 2022, when the data were recorded with a numerical resolution of 0.001 ppm, i.e., res <inline-formula><mml:math id="M273" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.001</mml:mn></mml:mrow></mml:math></inline-formula>. This series shows three CVEs with the length of 2, i.e., <inline-formula><mml:math id="M274" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:math></inline-formula>. The <inline-formula><mml:math id="M275" display="inline"><mml:mi mathvariant="italic">μ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M276" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M277" display="inline"><mml:mi mathvariant="normal">∅</mml:mi></mml:math></inline-formula> of the data in this figure are 0.62, 0.4 and 0.88, respectively.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://amt.copernicus.org/articles/16/3085/2023/amt-16-3085-2023-f15.png"/>

      </fig>

</app>
  </app-group><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e5496">The Python 3.7 code of the methodology is
available in Kaffashzadeh (2023, <ext-link xlink:href="https://doi.org/10.5281/zenodo.7951896" ext-link-type="DOI">10.5281/zenodo.7951896</ext-link>).</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e5505">The TOAR data infrastructure is available in Schröder et al. (2021, <ext-link xlink:href="https://doi.org/10.34730/4d9a287dec0b42f1aa6d244de8f19eb3" ext-link-type="DOI">10.34730/4d9a287dec0b42f1aa6d244de8f19eb3</ext-link>).</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e5514">The author has declared that there are no competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d1e5520">Publisher’s note: Copernicus Publications remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e5526">The scientific and technical support, various comments, and suggestions by
Martin G. Schultz have greatly improved this paper. The Australian Bureau of Meteorology, for providing the temperature time series data from Cape Grim, and the U.S. EPA AQS, for providing the ozone time series at Azusa and carbon monoxides data at Fresno, are appreciated. The author acknowledge
the constructive comments from the editor and the two anonymous referees.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e5531">This research has been supported by the European Research Council (ERC-2017-ADG, grant agreement no. 787576).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e5538">This paper was edited by Steffen Beirle and reviewed by two anonymous referees.</p>
  </notes><?xmltex \hack{\newpage}?><?xmltex \hack{\vspace*{70mm}}?><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><?label 1?><mixed-citation>Bey, I., Jacob, D. J., Yantosca, R. M., Logan, J. A., Field, B. D., Fiore,
A. M., Li, Q., Liu, H. Y., Mickley, L. J., and Schultz, M. G.: Global
modeling of tropospheric chemistry with assimilated meteorology: Model
description and evaluation, J. Geophys. Res.-Atmos., 106, 23073–23095, <ext-link xlink:href="https://doi.org/10.1029/2001JD000807" ext-link-type="DOI">10.1029/2001JD000807</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><?label 1?><mixed-citation>
Box, G. E. P., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M.: Time series
analysis: forecasting and control, 5th edn., John Wiley &amp; Sons,
Inc, Hoboken, New Jersey, 712 pp., ISBN: 978-1-118-67502-1, 2015.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><?label 1?><mixed-citation>Bushnell, M., Waldmann, C., Seitz, S., Buckley, E., Tamburri, M., Hermes, J., Heslop, E., and Lara-Lopez, A.: Quality Assurance of Oceanographic Observations: Standards and Guidance Adopted by an International Partnership, Frontiers in Marine Science, 6, 706, <ext-link xlink:href="https://doi.org/10.3389/fmars.2019.00706" ext-link-type="DOI">10.3389/fmars.2019.00706</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><?label 1?><mixed-citation>Campbell, J. L., Rustad, L. E., Porter, J. H., Taylor, J. R., Dereszynski,
E. W., Shanley, J. B., Gries, C., Henshaw, D. L., Martin, M. E., Sheldon, W.
M., and Boose, E. R.: Quantity is Nothing without Quality: Automated QA/QC
for Streaming Environmental Sensor Data, BioScience, 63, 574–585, <ext-link xlink:href="https://doi.org/10.1525/bio.2013.63.7.10" ext-link-type="DOI">10.1525/bio.2013.63.7.10</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><?label 1?><mixed-citation>Castelão, G. P.: A Flexible System for Automatic Quality Control of
Oceanographic Data, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.1503.02714" ext-link-type="DOI">10.48550/arXiv.1503.02714</ext-link>, 17 November 2016.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><?label 1?><mixed-citation>Chang, K.-L., Petropavlovskikh, I., Copper, O. R., Schultz, M. G., and Wang,
T.: Regional trend analysis of surface ozone observations from monitoring
networks in eastern North America, Europe and East Asia, Elementa: Science of the Anthropocene, 5, 50, <ext-link xlink:href="https://doi.org/10.1525/elementa.243" ext-link-type="DOI">10.1525/elementa.243</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><?label 1?><mixed-citation>Chapman, A. D.: Principles of Data Quality, Global Biodiversity Information Facility (GBIF) Secretariat, <ext-link xlink:href="https://doi.org/10.15468/DOC.JRGG-A190" ext-link-type="DOI">10.15468/DOC.JRGG-A190</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><?label 1?><mixed-citation>Dawson, J. P., Racherla, P. N., Lynn, B. H., Adams, P. J., and Pandis, S. N.:
Simulating present-day and future air quality as climate changes: Model evaluation, Atmos. Environ., 42, 4551–4566, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2008.01.058" ext-link-type="DOI">10.1016/j.atmosenv.2008.01.058</ext-link>, 2008.</mixed-citation></ref>
      <?pagebreak page3098?><ref id="bib1.bib9"><label>9</label><?label 1?><mixed-citation>Debry, E. and Mallet, V.: Ensemble forecasting with machine learning
algorithms for ozone, nitrogen dioxide and PM<inline-formula><mml:math id="M278" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">10</mml:mn></mml:msub></mml:math></inline-formula> on the Prev'Air platform, Atmos. Environ., 91, 71–84, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2014.03.049" ext-link-type="DOI">10.1016/j.atmosenv.2014.03.049</ext-link>,
2014.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><?label 1?><mixed-citation>Duong, T. and Hazelton, M. L.: Cross-validation Bandwidth Matrices for
Multivariate Kernel Density Estimation, Scand. J. Stat., 32, 485–506, <ext-link xlink:href="https://doi.org/10.1111/j.1467-9469.2005.00445.x" ext-link-type="DOI">10.1111/j.1467-9469.2005.00445.x</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><?label 1?><mixed-citation>Emmons, L. K., Walters, S., Hess, P. G., Lamarque, J.-F., Pfister, G. G., Fillmore, D., Granier, C., Guenther, A., Kinnison, D., Laepple, T., Orlando, J., Tie, X., Tyndall, G., Wiedinmyer, C., Baughcum, S. L., and Kloster, S.: Description and evaluation of the Model for Ozone and Related chemical Tracers, version 4 (MOZART-4), Geosci. Model Dev., 3, 43–67, <ext-link xlink:href="https://doi.org/10.5194/gmd-3-43-2010" ext-link-type="DOI">10.5194/gmd-3-43-2010</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><?label 1?><mixed-citation>Epanechnikov, V. A.: Non-Parametric Estimation of a Multivariate Probability
Density, Theor. Probab. Appl.+, 14, 153–158, <ext-link xlink:href="https://doi.org/10.1137/1114019" ext-link-type="DOI">10.1137/1114019</ext-link>, 1969.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><?label 1?><mixed-citation>Fang, Y., Naik, V., Horowitz, L. W., and Mauzerall, D. L.: Air pollution and associated human mortality: the role of air pollutant emissions, climate change and methane concentration increases from the preindustrial period to present, Atmos. Chem. Phys., 13, 1377–1394, <ext-link xlink:href="https://doi.org/10.5194/acp-13-1377-2013" ext-link-type="DOI">10.5194/acp-13-1377-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><?label 1?><mixed-citation>Fioletov, V. E. and Shepherd, T. G.: Seasonal persistence of midlatitude
total ozone anomalies: PERSISTENCE OF OZONE ANOMALIES, Geophys. Res.
Lett., 30, 1417, <ext-link xlink:href="https://doi.org/10.1029/2002GL016739" ext-link-type="DOI">10.1029/2002GL016739</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><?label 1?><mixed-citation>Fleming, Z. L., Doherty, R. M., Von Schneidemesser, E., Malley, C. S.,
Cooper, O. R., Pinto, J. P., Colette, A., Xu, X., Simpson, D., Schultz, M.
G., Lefohn, A. S., Hamad, S., Moolla, R., Solberg, S., and Feng, Z.:
Tropospheric Ozone Assessment Report: Present-day ozone distribution and
trends relevant to human health, Elementa: Science of the Anthropocene, 6, 12, <ext-link xlink:href="https://doi.org/10.1525/elementa.273" ext-link-type="DOI">10.1525/elementa.273</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><?label 1?><mixed-citation>Gandin, L. S.: Complex Quality Control of Meteorological Observations,
Mon. Weather Rev., 116, 1137–1156, <ext-link xlink:href="https://doi.org/10.1175/1520-0493(1988)116&lt;1137:CQCOMO&gt;2.0.CO;2" ext-link-type="DOI">10.1175/1520-0493(1988)116&lt;1137:CQCOMO&gt;2.0.CO;2</ext-link>, 1988.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><?label 1?><mixed-citation>Gardner, M.: Neural network modelling and prediction of hourly NO<inline-formula><mml:math id="M279" display="inline"><mml:msub><mml:mi/><mml:mi>x</mml:mi></mml:msub></mml:math></inline-formula> and NO<inline-formula><mml:math id="M280" display="inline"><mml:msub><mml:mi/><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:math></inline-formula> concentrations in urban air in London, Atmos. Environ., 33, 709–719, <ext-link xlink:href="https://doi.org/10.1016/S1352-2310(98)00230-1" ext-link-type="DOI">10.1016/S1352-2310(98)00230-1</ext-link>, 1999.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><?label 1?><mixed-citation>Good, S., Mills, B., and Castelao, G.: AutoQC: Automatic quality control
analysis for the international quality controlled ocean database, Zenodo
[code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.5832003" ext-link-type="DOI">10.5281/zenodo.5832003</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><?label 1?><mixed-citation>
Grant, E. L. and Leavenworth, R. S.: Statistical Quality Control, MacGraw Hill, New-York, NY, 764 pp., ISBN 10: 0078443547, ISBN 13: 9780078443541, 1996.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><?label 1?><mixed-citation>Gudmundsson, L., Do, H. X., Leonard, M., and Westra, S.: The Global Streamflow Indices and Metadata Archive (GSIM) – Part 2: Quality control, time-series indices and homogeneity assessment, Earth Syst. Sci. Data, 10, 787–804, <ext-link xlink:href="https://doi.org/10.5194/essd-10-787-2018" ext-link-type="DOI">10.5194/essd-10-787-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><?label 1?><mixed-citation>Guttorp, P., Meiring, W., and Sampson, P. D.: A space-time analysis of
ground-level ozone data, Environmetrics, 5, 241–254, <ext-link xlink:href="https://doi.org/10.1002/env.3170050305" ext-link-type="DOI">10.1002/env.3170050305</ext-link>, 1994.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><?label 1?><mixed-citation>Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A.,
Nicolas, J., Peubey, C., Radu, R., Bonavita, M., Dee, D., Dragani, R.,
Flemming, J., Forbes, R., Geer, A., Hogan, R. J., Janisková, H. M.,
Keeley, S., Laloyaux, P., Cristina, P. L., and Thépaut, J.: The ERA5
global reanalysis, 1999–2049, Q. J. Roy. Meteor. Soc., 146, 1999–2049, <ext-link xlink:href="https://doi.org/10.1002/qj.3803" ext-link-type="DOI">10.1002/qj.3803</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><?label 1?><mixed-citation>Horowitz, L. W., Walters, S., Mauzerall, D. L., Emmons, L. K., Rasch, P. J.,
Granier, C., Tie, X., Lamarque, J.-F., Schultz, M. G., Tyndall, G. S.,
Orlando, J. J., and Brasseur, G. P.: A global simulation of tropospheric ozone and related tracers: Description and evaluation of MOZART, version 2, J. Geophys. Res., 108, 4784, <ext-link xlink:href="https://doi.org/10.1029/2002JD002853" ext-link-type="DOI">10.1029/2002JD002853</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><?label 1?><mixed-citation>Horsburgh, J. S., Reeder, S. L., Jones, A. S., and Meline, J.: Open source
software for visualization and quality control of continuous hydrologic and
water quality sensor data, Environ. Modell. Softw., 70, 32–44, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2015.04.002" ext-link-type="DOI">10.1016/j.envsoft.2015.04.002</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><?label 1?><mixed-citation>Hwang, J.-N., Lay, S.-R., and Lippman, A.: Nonparametric multivariate density estimation: a comparative study, IEEE T. Signal Proces., 42, 2795–2810, <ext-link xlink:href="https://doi.org/10.1109/78.324744" ext-link-type="DOI">10.1109/78.324744</ext-link>, 1994.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><?label 1?><mixed-citation>Im, U., Bianconi, R., Solazzo, E., Kioutsioukis, I., Badia, A., Balzarini,
A., Baró, R., Bellasio, R., Brunner, D., Chemel, C., Curci, G., Flemming, J., Forkel, R., Giordano, L., Jiménez-Guerrero, P., Hirtl, M., Hodzic, A., Honzak, L., Jorba, O., Knote, C., Kuenen, J. J. P., Makar, P. A., Manders-Groot, A., Neal, L., Pérez, J. L., Pirovano, G., Pouliot, G., San Jose, R., Savage, N., Schroder, W., Sokhi, R. S., Syrakov, D., Torian, A., Tuccella, P., Werhahn, J., Wolke, R., Yahya, K., Zabkar, R., Zhang, Y., Zhang, J., Hogrefe, C., and Galmarini, S.: Evaluation of operational on-line-coupled regional air quality models over Europe and North America in the context of AQMEII phase 2. Part I: Ozone, Atmos. Environ., 115, 404–420, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2014.09.042" ext-link-type="DOI">10.1016/j.atmosenv.2014.09.042</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><?label 1?><mixed-citation>Inness, A., Ades, M., Agustí-Panareda, A., Barré, J., Benedictow, A., Blechschmidt, A.-M., Dominguez, J. J., Engelen, R., Eskes, H., Flemming, J., Huijnen, V., Jones, L., Kipling, Z., Massart, S., Parrington, M., Peuch, V.-H., Razinger, M., Remy, S., Schulz, M., and Suttie, M.: The CAMS reanalysis of atmospheric composition, Atmos. Chem. Phys., 19, 3515–3556, <ext-link xlink:href="https://doi.org/10.5194/acp-19-3515-2019" ext-link-type="DOI">10.5194/acp-19-3515-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><?label 1?><mixed-citation>Kaffashzadeh, N.: A statistical data-driven test for probabilistic data
quality control, Version v1.0.0, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.7951896" ext-link-type="DOI">10.5281/zenodo.7951896</ext-link>,
2023.</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><?label 1?><mixed-citation>Kumar, U. and De Ridder, K.: GARCH modelling in association with FFT–ARIMA
to forecast ozone episodes, Atmos. Environ., 44, 4252–4265,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2010.06.055" ext-link-type="DOI">10.1016/j.atmosenv.2010.06.055</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib30"><label>30</label><?label 1?><mixed-citation>Lamarque, J.-F., Emmons, L. K., Hess, P. G., Kinnison, D. E., Tilmes, S., Vitt, F., Heald, C. L., Holland, E. A., Lauritzen, P. H., Neu, J., Orlando, J. J., Rasch, P. J., and Tyndall, G. K.: CAM-chem: description and evaluation of interactive atmospheric chemistry in the Community Earth System Model, Geosci. Model Dev., 5, 369–411, <ext-link xlink:href="https://doi.org/10.5194/gmd-5-369-2012" ext-link-type="DOI">10.5194/gmd-5-369-2012</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><?label 1?><mixed-citation>Lefohn, A. S., Malley, C. S., Smith, L., Wells, B., Hazucha, M., Simon, H.,
Naik, V., Mills, G., Schultz, M. G., Paoletti, E., De Marco, A., Xu, X.,
Zhang, L., Wang, T., Neufeld, H. S., Musselman, R. C., Tarasick, D., Brauer,
M., Feng, Z., Tang, H., Kobayashi, K., Sicard, P., Solberg, S., and Gerosa,
G.: Tropospheric ozone assessment report: Global ozon<?pagebreak page3099?>e metrics for climate
change, human health, and crop/ecosystem research, Elementa: Science of the Anthropocene, 6, 28, <ext-link xlink:href="https://doi.org/10.1525/elementa.279" ext-link-type="DOI">10.1525/elementa.279</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><?label 1?><mixed-citation>Lyapina, O., Schultz, M. G., and Hense, A.: Cluster analysis of European surface ozone observations for evaluation of MACC reanalysis data, Atmos. Chem. Phys., 16, 6863–6881, <ext-link xlink:href="https://doi.org/10.5194/acp-16-6863-2016" ext-link-type="DOI">10.5194/acp-16-6863-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><?label 1?><mixed-citation>Mills, G., Harmens, H., Wagg, S., Sharps, K., Hayes, F., Fowler, D., Sutton,
M., and Davies, B.: Ozone impacts on vegetation in a nitrogen enriched and
changing climate, Environ. Pollut., 208, 898–908, <ext-link xlink:href="https://doi.org/10.1016/j.envpol.2015.09.038" ext-link-type="DOI">10.1016/j.envpol.2015.09.038</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib34"><label>34</label><?label 1?><mixed-citation>Mills, G., Pleijel, H., Malley, C. S., Sinha, B., Cooper, O. R., Schultz, M.
G., Neufeld, H. S., Simpson, D., Sharps, K., Feng, Z., Gerosa, G., Harmens,
H., Kobayashi, K., Saxena, P., Paoletti, E., Sinha, V., and Xu, X.:
Tropospheric Ozone Assessment Report: Present-day tropospheric ozone
distribution and trends relevant to vegetation, Elementa: Science of the Anthropocene, 6, 47, <ext-link xlink:href="https://doi.org/10.1525/elementa.302" ext-link-type="DOI">10.1525/elementa.302</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><?label 1?><mixed-citation>Niu, X.-F.: Nonlinear Additive Models for Environmental Time Series, with
Applications to Ground-Level Ozone Data Analysis, J. Am. Stat. Assoc., 91, 1310–1321, <ext-link xlink:href="https://doi.org/10.1080/01621459.1996.10477000" ext-link-type="DOI">10.1080/01621459.1996.10477000</ext-link>, 1996.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><?label 1?><mixed-citation>Osborne, J. W. and Overbay, A.: The Power of Outliers (and Why Researchers
Should Always Check for Them), Practical Assessment, Research, and
Evaluation, 9, 6, <ext-link xlink:href="https://doi.org/10.7275/qf69-7k43" ext-link-type="DOI">10.7275/qf69-7k43</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><?label 1?><mixed-citation>Rasmussen, D. J., Fiore, A. M., Naik, V., Horowitz, L. W., McGinnis, S. J.,
and Schultz, M. G.: Surface ozone-temperature relationships in the eastern
US: A monthly climatology for evaluating chemistry-climate models,
Atmos. Environ., 47, 142–153, <ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2011.11.021" ext-link-type="DOI">10.1016/j.atmosenv.2011.11.021</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><?label 1?><mixed-citation>Reinsel, G. C., Weatherhead, E., Tiao, G. C., Miller, A. J., Nagatani, R.
M., Wuebbles, D. J., and Flynn, L. E.: On detection of turnaround and
recovery in trend for ozone, J. Geophys. Res.-Atmos., 107, ACH 1-1–ACH 1-12,
<ext-link xlink:href="https://doi.org/10.1029/2001JD000500" ext-link-type="DOI">10.1029/2001JD000500</ext-link>, 2002.</mixed-citation></ref>
      <ref id="bib1.bib39"><label>39</label><?label 1?><mixed-citation>Rencher, A. C.: Methods of multivariate analysis, 2nd edn., J. Wiley, New
York, <ext-link xlink:href="https://doi.org/10.1002/0471271357" ext-link-type="DOI">10.1002/0471271357</ext-link>, 2002.</mixed-citation></ref>
      <ref id="bib1.bib40"><label>40</label><?label 1?><mixed-citation>Schnell, J. L., Prather, M. J., Josse, B., Naik, V., Horowitz, L. W., Cameron-Smith, P., Bergmann, D., Zeng, G., Plummer, D. A., Sudo, K., Nagashima, T., Shindell, D. T., Faluvegi, G., and Strode, S. A.: Use of North American and European air quality networks to evaluate global chemistry–climate modeling of surface ozone, Atmos. Chem. Phys., 15, 10581–10596, <ext-link xlink:href="https://doi.org/10.5194/acp-15-10581-2015" ext-link-type="DOI">10.5194/acp-15-10581-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib41"><label>41</label><?label 1?><mixed-citation>Schröder, S., Schultz, M. G., Selke, N., Sun, J., Ahring, J., Mozaffari, A., Romberg, M., Epp, E., Lensing, M., Apweiler, S., Leufen, L. H., Betancourt, C., Hagemeier, B., and Rajveer, S.: TOAR Data Infrastructure, Version v1.0, FZ-Juelich B2SHARE [data set], <ext-link xlink:href="https://doi.org/10.34730/4d9a287dec0b42f1aa6d244de8f19eb3" ext-link-type="DOI">10.34730/4d9a287dec0b42f1aa6d244de8f19eb3</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib42"><label>42</label><?label 1?><mixed-citation>Schultz, M. G., Schröder, S., Lyapina, O., Cooper, O., Galbally, I.,
Petropavlovskikh, I., Von Schneidemesser, E., Tanimoto, H., Elshorbany, Y.,
Naja, M., Seguel, R., Dauert, U., Eckhardt, P., Feigenspahn, S., Fiebig, M.,
Hjellbrekke, A.-G., Hong, Y.-D., Christian Kjeld, P., Koide, H., Lear, G.,
Tarasick, D., Ueno, M., Wallasch, M., Baumgardner, D., Chuang, M.-T.,
Gillett, R., Lee, M., Molloy, S., Moolla, R., Wang, T., Sharps, K., Adame,
J. A., Ancellet, G., Apadula, F., Artaxo, P., Barlasina, M., Bogucka, M.,
Bonasoni, P., Chang, L., Colomb, A., Cuevas, E., Cupeiro, M., Degorska, A.,
Ding, A., Fröhlich, M., Frolova, M., Gadhavi, H., Gheusi, F., Gilge, S.,
Gonzalez, M. Y., Gros, V., Hamad, S. H., Helmig, D., Henriques, D.,
Hermansen, O., Holla, R., Huber, J., Im, U., Jaffe, D. A., Komala, N.,
Kubistin, D., Lam, K.-S., Laurila, T., Lee, H., Levy, I., Mazzoleni, C.,
Mazzoleni, L., McClure-Begley, A., Mohamad, M., Murovic, M., Navarro-Comas,
M., Nicodim, F., Parrish, D., Read, K. A., Reid, N., Ries, L., Saxena, P.,
Schwab, J. J., Scorgie, Y., Senik, I., Simmonds, P., Sinha, V., Skorokhod,
A., Spain, G., Spangl, W., Spoor, R., Springston, S. R., Steer, K.,
Steinbacher, M., Suharguniyawan, E., Torre, P., Trickl, T., Weili, L.,
Weller, R., Xu, X., Xue, L., and Zhiqiang, M.: Tropospheric Ozone Assessment
Report: Database and Metrics Data of Global Surface Ozone Observations, Elementa: Science of the Anthropocene, 5, 58, <ext-link xlink:href="https://doi.org/10.1525/elementa.244" ext-link-type="DOI">10.1525/elementa.244</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib43"><label>43</label><?label 1?><mixed-citation>
Schum, D. A.: The evidential foundations of probabilistic reasoning,
Northwestern University Press, Evanston, Ill., ISBN-13: 978-0810118218, 2001.</mixed-citation></ref>
      <ref id="bib1.bib44"><label>44</label><?label 1?><mixed-citation>Scully-Allison, C., Le, V., Fritzinger, E., Strachan, S., Harris, F. C., and
Dascalu, S. M.: Near Real-time Autonomous Quality Control for Streaming
Environmental Sensor Data, Procedia Comput. Sci., 126, 1656–1665,
<ext-link xlink:href="https://doi.org/10.1016/j.procs.2018.08.139" ext-link-type="DOI">10.1016/j.procs.2018.08.139</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib45"><label>45</label><?label 1?><mixed-citation>Sofen, E. D., Bowdalo, D., Evans, M. J., Apadula, F., Bonasoni, P., Cupeiro, M., Ellul, R., Galbally, I. E., Girgzdiene, R., Luppo, S., Mimouni, M., Nahas, A. C., Saliba, M., and Tørseth, K.: Gridded global surface ozone metrics for atmospheric chemistry model evaluation, Earth Syst. Sci. Data, 8, 41–59, <ext-link xlink:href="https://doi.org/10.5194/essd-8-41-2016" ext-link-type="DOI">10.5194/essd-8-41-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib46"><label>46</label><?label 1?><mixed-citation>Steinacker, R., Mayer, D., and Steiner, A.: Data Quality Control Based on
Self-Consistency, Mon. Weather Rev., 139, 3974–3991,
<ext-link xlink:href="https://doi.org/10.1175/MWR-D-10-05024.1" ext-link-type="DOI">10.1175/MWR-D-10-05024.1</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib47"><label>47</label><?label 1?><mixed-citation>Tiao, G. C., Reinsel, G. C., Xu, D., Pedrick, J. H., Zhu, X., Miller, A. J.,
DeLuisi, J. J., Mateer, C. L., and Wuebbles, D. J.: Effects of autocorrelation and temporal sampling schemes on estimates of trend and
spatial correlation, J. Geophys. Res., 95, 20507, <ext-link xlink:href="https://doi.org/10.1029/JD095iD12p20507" ext-link-type="DOI">10.1029/JD095iD12p20507</ext-link>, 1990.</mixed-citation></ref>
      <ref id="bib1.bib48"><label>48</label><?label 1?><mixed-citation>Tilmes, S., Lamarque, J.-F., Emmons, L. K., Conley, A., Schultz, M. G., Saunois, M., Thouret, V., Thompson, A. M., Oltmans, S. J., Johnson, B., and Tarasick, D.: Technical Note: Ozonesonde climatology between 1995 and 2011: description, evaluation and applications, Atmos. Chem. Phys., 12, 7475–7497, <ext-link xlink:href="https://doi.org/10.5194/acp-12-7475-2012" ext-link-type="DOI">10.5194/acp-12-7475-2012</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib49"><label>49</label><?label 1?><mixed-citation>
Tong, Y. L.: The multivariate normal distribution, Springer-Verlag, New
York, ISBN: 978-1-4613-9655-0, 1990.</mixed-citation></ref>
      <ref id="bib1.bib50"><label>50</label><?label 1?><mixed-citation>Waterman, M. S. and Whiteman, D. E.: Estimation of probability densities by
empirical density functions, International Journal of Mathematical Education in Science and Technology, 9, 127–137, <ext-link xlink:href="https://doi.org/10.1080/0020739780090201" ext-link-type="DOI">10.1080/0020739780090201</ext-link>, 1978.</mixed-citation></ref>
      <ref id="bib1.bib51"><label>51</label><?label 1?><mixed-citation>Weatherhead, E. C., Reinsel, G. C., Tiao, G. C., Meng, X.-L., Choi, D.,
Cheang, W.-K., Keller, T., DeLuisi, J., Wuebbles, D. J., Kerr, J. B.,
Miller, A. J., Oltmans, S. J., and Frederick, J. E.: Factors affecting the
detection of trends: Statistical considerations and applications to
environmental data, J. Geophys. Res.-Atmos., 103, 17149–17161, <ext-link xlink:href="https://doi.org/10.1029/98JD00995" ext-link-type="DOI">10.1029/98JD00995</ext-link>, 1998.</mixed-citation></ref>
      <?pagebreak page3100?><ref id="bib1.bib52"><label>52</label><?label 1?><mixed-citation>Weatherhead, E. C., Reinsel, G. C., Tiao, G. C., Jackman, C. H., Bishop, L.,
Frith, S. M. H., DeLuisi, J., Keller, T., Oltmans, S. J., Fleming, E. L.,
Wuebbles, D. J., Kerr, J. B., Miller, A. J., Herman, J., McPeters, R.,
Nagatani, R. M., and Frederick, J. E.: Detecting the recovery of total column
ozone, J. Geophys. Res.-Atmos., 105, 22201–22210, <ext-link xlink:href="https://doi.org/10.1029/2000JD900063" ext-link-type="DOI">10.1029/2000JD900063</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bib53"><label>53</label><?label 1?><mixed-citation>Wincek, M. A. and Reinsel, G. C.: An Exact Maximum Likelihood Estimation
Procedure for Regression-<italic>ARMA</italic> Time Series Models with Possibly
Nonconsecutive Data, J. Roy. Stat. Soc. B Met., 48, 303–313, <ext-link xlink:href="https://doi.org/10.1111/j.2517-6161.1986.tb01414.x" ext-link-type="DOI">10.1111/j.2517-6161.1986.tb01414.x</ext-link>, 1986.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bib54"><label>54</label><?label 1?><mixed-citation>Zahumenský, I.: Guidelines on Quality Control Procedures for Data from
Automatic Weather Stations, World Meteorological Organization, 11 pp.,
<uri>https://www.researchgate.net/publication/228826920_Guidelines_on_Quality_Control_Procedures_for_Data_from_Automatic_Weather_Stations</uri> (last access: 15 June 2023), 2004.</mixed-citation></ref>
      <ref id="bib1.bib55"><label>55</label><?label 1?><mixed-citation>Zhang, Y., Bocquet, M., Mallet, V., Seigneur, C., and Baklanov, A.: Real-time
air quality forecasting, part II: State of the science, current research
needs, and future prospects, Atmos. Environ., 60, 656–676,
<ext-link xlink:href="https://doi.org/10.1016/j.atmosenv.2012.02.041" ext-link-type="DOI">10.1016/j.atmosenv.2012.02.041</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib56"><label>56</label><?label 1?><mixed-citation>Zhou, Y., Chang, F.-J., Chang, L.-C., Kao, I.-F., and Wang, Y.-S.: Explore a
deep learning multi-output neural network for regional multi-step-ahead air
quality forecasts, J. Clean. Prod., 209, 134–145,
<ext-link xlink:href="https://doi.org/10.1016/j.jclepro.2018.10.243" ext-link-type="DOI">10.1016/j.jclepro.2018.10.243</ext-link>, 2019.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>A data-driven persistence test for robust (probabilistic) quality control of measured environmental time series: constant value episodes</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
      
Bey, I., Jacob, D. J., Yantosca, R. M., Logan, J. A., Field, B. D., Fiore,
A. M., Li, Q., Liu, H. Y., Mickley, L. J., and Schultz, M. G.: Global
modeling of tropospheric chemistry with assimilated meteorology: Model
description and evaluation, J. Geophys. Res.-Atmos., 106, 23073–23095, <a href="https://doi.org/10.1029/2001JD000807" target="_blank">https://doi.org/10.1029/2001JD000807</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
      
Box, G. E. P., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M.: Time series
analysis: forecasting and control, 5th edn., John Wiley &amp; Sons,
Inc, Hoboken, New Jersey, 712 pp., ISBN:&thinsp;978-1-118-67502-1, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
      
Bushnell, M., Waldmann, C., Seitz, S., Buckley, E., Tamburri, M., Hermes, J., Heslop, E., and Lara-Lopez, A.: Quality Assurance of Oceanographic Observations: Standards and Guidance Adopted by an International Partnership, Frontiers in Marine Science, 6, 706, <a href="https://doi.org/10.3389/fmars.2019.00706" target="_blank">https://doi.org/10.3389/fmars.2019.00706</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
      
Campbell, J. L., Rustad, L. E., Porter, J. H., Taylor, J. R., Dereszynski,
E. W., Shanley, J. B., Gries, C., Henshaw, D. L., Martin, M. E., Sheldon, W.
M., and Boose, E. R.: Quantity is Nothing without Quality: Automated QA/QC
for Streaming Environmental Sensor Data, BioScience, 63, 574–585, <a href="https://doi.org/10.1525/bio.2013.63.7.10" target="_blank">https://doi.org/10.1525/bio.2013.63.7.10</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
      
Castelão, G. P.: A Flexible System for Automatic Quality Control of
Oceanographic Data, arXiv [preprint], <a href="https://doi.org/10.48550/arXiv.1503.02714" target="_blank">https://doi.org/10.48550/arXiv.1503.02714</a>, 17&thinsp;November&thinsp;2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
      
Chang, K.-L., Petropavlovskikh, I., Copper, O. R., Schultz, M. G., and Wang,
T.: Regional trend analysis of surface ozone observations from monitoring
networks in eastern North America, Europe and East Asia, Elementa: Science of the Anthropocene, 5, 50, <a href="https://doi.org/10.1525/elementa.243" target="_blank">https://doi.org/10.1525/elementa.243</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
      
Chapman, A. D.: Principles of Data Quality, Global Biodiversity Information Facility (GBIF) Secretariat, <a href="https://doi.org/10.15468/DOC.JRGG-A190" target="_blank">https://doi.org/10.15468/DOC.JRGG-A190</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
      
Dawson, J. P., Racherla, P. N., Lynn, B. H., Adams, P. J., and Pandis, S. N.:
Simulating present-day and future air quality as climate changes: Model evaluation, Atmos. Environ., 42, 4551–4566, <a href="https://doi.org/10.1016/j.atmosenv.2008.01.058" target="_blank">https://doi.org/10.1016/j.atmosenv.2008.01.058</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
      
Debry, E. and Mallet, V.: Ensemble forecasting with machine learning
algorithms for ozone, nitrogen dioxide and PM<sub>10</sub> on the Prev'Air platform, Atmos. Environ., 91, 71–84, <a href="https://doi.org/10.1016/j.atmosenv.2014.03.049" target="_blank">https://doi.org/10.1016/j.atmosenv.2014.03.049</a>,
2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
      
Duong, T. and Hazelton, M. L.: Cross-validation Bandwidth Matrices for
Multivariate Kernel Density Estimation, Scand. J. Stat., 32, 485–506, <a href="https://doi.org/10.1111/j.1467-9469.2005.00445.x" target="_blank">https://doi.org/10.1111/j.1467-9469.2005.00445.x</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
      
Emmons, L. K., Walters, S., Hess, P. G., Lamarque, J.-F., Pfister, G. G., Fillmore, D., Granier, C., Guenther, A., Kinnison, D., Laepple, T., Orlando, J., Tie, X., Tyndall, G., Wiedinmyer, C., Baughcum, S. L., and Kloster, S.: Description and evaluation of the Model for Ozone and Related chemical Tracers, version 4 (MOZART-4), Geosci. Model Dev., 3, 43–67, <a href="https://doi.org/10.5194/gmd-3-43-2010" target="_blank">https://doi.org/10.5194/gmd-3-43-2010</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
      
Epanechnikov, V. A.: Non-Parametric Estimation of a Multivariate Probability
Density, Theor. Probab. Appl.+, 14, 153–158, <a href="https://doi.org/10.1137/1114019" target="_blank">https://doi.org/10.1137/1114019</a>, 1969.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
      
Fang, Y., Naik, V., Horowitz, L. W., and Mauzerall, D. L.: Air pollution and associated human mortality: the role of air pollutant emissions, climate change and methane concentration increases from the preindustrial period to present, Atmos. Chem. Phys., 13, 1377–1394, <a href="https://doi.org/10.5194/acp-13-1377-2013" target="_blank">https://doi.org/10.5194/acp-13-1377-2013</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
      
Fioletov, V. E. and Shepherd, T. G.: Seasonal persistence of midlatitude
total ozone anomalies: PERSISTENCE OF OZONE ANOMALIES, Geophys. Res.
Lett., 30, 1417, <a href="https://doi.org/10.1029/2002GL016739" target="_blank">https://doi.org/10.1029/2002GL016739</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
      
Fleming, Z. L., Doherty, R. M., Von Schneidemesser, E., Malley, C. S.,
Cooper, O. R., Pinto, J. P., Colette, A., Xu, X., Simpson, D., Schultz, M.
G., Lefohn, A. S., Hamad, S., Moolla, R., Solberg, S., and Feng, Z.:
Tropospheric Ozone Assessment Report: Present-day ozone distribution and
trends relevant to human health, Elementa: Science of the Anthropocene, 6, 12, <a href="https://doi.org/10.1525/elementa.273" target="_blank">https://doi.org/10.1525/elementa.273</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
      
Gandin, L. S.: Complex Quality Control of Meteorological Observations,
Mon. Weather Rev., 116, 1137–1156, <a href="https://doi.org/10.1175/1520-0493(1988)116&lt;1137:CQCOMO&gt;2.0.CO;2" target="_blank">https://doi.org/10.1175/1520-0493(1988)116&lt;1137:CQCOMO&gt;2.0.CO;2</a>, 1988.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
      
Gardner, M.: Neural network modelling and prediction of hourly NO<sub><i>x</i></sub> and NO<sub>2</sub> concentrations in urban air in London, Atmos. Environ., 33, 709–719, <a href="https://doi.org/10.1016/S1352-2310(98)00230-1" target="_blank">https://doi.org/10.1016/S1352-2310(98)00230-1</a>, 1999.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
      
Good, S., Mills, B., and Castelao, G.: AutoQC: Automatic quality control
analysis for the international quality controlled ocean database, Zenodo
[code], <a href="https://doi.org/10.5281/zenodo.5832003" target="_blank">https://doi.org/10.5281/zenodo.5832003</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
      
Grant, E. L. and Leavenworth, R. S.: Statistical Quality Control, MacGraw Hill, New-York, NY, 764 pp., ISBN&thinsp;10:&thinsp;0078443547, ISBN&thinsp;13:&thinsp;9780078443541, 1996.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
      
Gudmundsson, L., Do, H. X., Leonard, M., and Westra, S.: The Global Streamflow Indices and Metadata Archive (GSIM) – Part 2: Quality control, time-series indices and homogeneity assessment, Earth Syst. Sci. Data, 10, 787–804, <a href="https://doi.org/10.5194/essd-10-787-2018" target="_blank">https://doi.org/10.5194/essd-10-787-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
      
Guttorp, P., Meiring, W., and Sampson, P. D.: A space-time analysis of
ground-level ozone data, Environmetrics, 5, 241–254, <a href="https://doi.org/10.1002/env.3170050305" target="_blank">https://doi.org/10.1002/env.3170050305</a>, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
      
Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A.,
Nicolas, J., Peubey, C., Radu, R., Bonavita, M., Dee, D., Dragani, R.,
Flemming, J., Forbes, R., Geer, A., Hogan, R. J., Janisková, H. M.,
Keeley, S., Laloyaux, P., Cristina, P. L., and Thépaut, J.: The ERA5
global reanalysis, 1999–2049, Q. J. Roy. Meteor. Soc., 146, 1999–2049, <a href="https://doi.org/10.1002/qj.3803" target="_blank">https://doi.org/10.1002/qj.3803</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
      
Horowitz, L. W., Walters, S., Mauzerall, D. L., Emmons, L. K., Rasch, P. J.,
Granier, C., Tie, X., Lamarque, J.-F., Schultz, M. G., Tyndall, G. S.,
Orlando, J. J., and Brasseur, G. P.: A global simulation of tropospheric ozone and related tracers: Description and evaluation of MOZART, version 2, J. Geophys. Res., 108, 4784, <a href="https://doi.org/10.1029/2002JD002853" target="_blank">https://doi.org/10.1029/2002JD002853</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
      
Horsburgh, J. S., Reeder, S. L., Jones, A. S., and Meline, J.: Open source
software for visualization and quality control of continuous hydrologic and
water quality sensor data, Environ. Modell. Softw., 70, 32–44, <a href="https://doi.org/10.1016/j.envsoft.2015.04.002" target="_blank">https://doi.org/10.1016/j.envsoft.2015.04.002</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
      
Hwang, J.-N., Lay, S.-R., and Lippman, A.: Nonparametric multivariate density estimation: a comparative study, IEEE T. Signal Proces., 42, 2795–2810, <a href="https://doi.org/10.1109/78.324744" target="_blank">https://doi.org/10.1109/78.324744</a>, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
      
Im, U., Bianconi, R., Solazzo, E., Kioutsioukis, I., Badia, A., Balzarini,
A., Baró, R., Bellasio, R., Brunner, D., Chemel, C., Curci, G., Flemming, J., Forkel, R., Giordano, L., Jiménez-Guerrero, P., Hirtl, M., Hodzic, A., Honzak, L., Jorba, O., Knote, C., Kuenen, J. J. P., Makar, P. A., Manders-Groot, A., Neal, L., Pérez, J. L., Pirovano, G., Pouliot, G., San Jose, R., Savage, N., Schroder, W., Sokhi, R. S., Syrakov, D., Torian, A., Tuccella, P., Werhahn, J., Wolke, R., Yahya, K., Zabkar, R., Zhang, Y., Zhang, J., Hogrefe, C., and Galmarini, S.: Evaluation of operational on-line-coupled regional air quality models over Europe and North America in the context of AQMEII phase 2. Part I: Ozone, Atmos. Environ., 115, 404–420, <a href="https://doi.org/10.1016/j.atmosenv.2014.09.042" target="_blank">https://doi.org/10.1016/j.atmosenv.2014.09.042</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
      
Inness, A., Ades, M., Agustí-Panareda, A., Barré, J., Benedictow, A., Blechschmidt, A.-M., Dominguez, J. J., Engelen, R., Eskes, H., Flemming, J., Huijnen, V., Jones, L., Kipling, Z., Massart, S., Parrington, M., Peuch, V.-H., Razinger, M., Remy, S., Schulz, M., and Suttie, M.: The CAMS reanalysis of atmospheric composition, Atmos. Chem. Phys., 19, 3515–3556, <a href="https://doi.org/10.5194/acp-19-3515-2019" target="_blank">https://doi.org/10.5194/acp-19-3515-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
      
Kaffashzadeh, N.: A statistical data-driven test for probabilistic data
quality control, Version v1.0.0, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.7951896" target="_blank">https://doi.org/10.5281/zenodo.7951896</a>,
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
      
Kumar, U. and De Ridder, K.: GARCH modelling in association with FFT–ARIMA
to forecast ozone episodes, Atmos. Environ., 44, 4252–4265,
<a href="https://doi.org/10.1016/j.atmosenv.2010.06.055" target="_blank">https://doi.org/10.1016/j.atmosenv.2010.06.055</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
      
Lamarque, J.-F., Emmons, L. K., Hess, P. G., Kinnison, D. E., Tilmes, S., Vitt, F., Heald, C. L., Holland, E. A., Lauritzen, P. H., Neu, J., Orlando, J. J., Rasch, P. J., and Tyndall, G. K.: CAM-chem: description and evaluation of interactive atmospheric chemistry in the Community Earth System Model, Geosci. Model Dev., 5, 369–411, <a href="https://doi.org/10.5194/gmd-5-369-2012" target="_blank">https://doi.org/10.5194/gmd-5-369-2012</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
      
Lefohn, A. S., Malley, C. S., Smith, L., Wells, B., Hazucha, M., Simon, H.,
Naik, V., Mills, G., Schultz, M. G., Paoletti, E., De Marco, A., Xu, X.,
Zhang, L., Wang, T., Neufeld, H. S., Musselman, R. C., Tarasick, D., Brauer,
M., Feng, Z., Tang, H., Kobayashi, K., Sicard, P., Solberg, S., and Gerosa,
G.: Tropospheric ozone assessment report: Global ozone metrics for climate
change, human health, and crop/ecosystem research, Elementa: Science of the Anthropocene, 6, 28, <a href="https://doi.org/10.1525/elementa.279" target="_blank">https://doi.org/10.1525/elementa.279</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
      
Lyapina, O., Schultz, M. G., and Hense, A.: Cluster analysis of European surface ozone observations for evaluation of MACC reanalysis data, Atmos. Chem. Phys., 16, 6863–6881, <a href="https://doi.org/10.5194/acp-16-6863-2016" target="_blank">https://doi.org/10.5194/acp-16-6863-2016</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
      
Mills, G., Harmens, H., Wagg, S., Sharps, K., Hayes, F., Fowler, D., Sutton,
M., and Davies, B.: Ozone impacts on vegetation in a nitrogen enriched and
changing climate, Environ. Pollut., 208, 898–908, <a href="https://doi.org/10.1016/j.envpol.2015.09.038" target="_blank">https://doi.org/10.1016/j.envpol.2015.09.038</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
      
Mills, G., Pleijel, H., Malley, C. S., Sinha, B., Cooper, O. R., Schultz, M.
G., Neufeld, H. S., Simpson, D., Sharps, K., Feng, Z., Gerosa, G., Harmens,
H., Kobayashi, K., Saxena, P., Paoletti, E., Sinha, V., and Xu, X.:
Tropospheric Ozone Assessment Report: Present-day tropospheric ozone
distribution and trends relevant to vegetation, Elementa: Science of the Anthropocene, 6, 47, <a href="https://doi.org/10.1525/elementa.302" target="_blank">https://doi.org/10.1525/elementa.302</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
      
Niu, X.-F.: Nonlinear Additive Models for Environmental Time Series, with
Applications to Ground-Level Ozone Data Analysis, J. Am. Stat. Assoc., 91, 1310–1321, <a href="https://doi.org/10.1080/01621459.1996.10477000" target="_blank">https://doi.org/10.1080/01621459.1996.10477000</a>, 1996.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
      
Osborne, J. W. and Overbay, A.: The Power of Outliers (and Why Researchers
Should Always Check for Them), Practical Assessment, Research, and
Evaluation, 9, 6, <a href="https://doi.org/10.7275/qf69-7k43" target="_blank">https://doi.org/10.7275/qf69-7k43</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
      
Rasmussen, D. J., Fiore, A. M., Naik, V., Horowitz, L. W., McGinnis, S. J.,
and Schultz, M. G.: Surface ozone-temperature relationships in the eastern
US: A monthly climatology for evaluating chemistry-climate models,
Atmos. Environ., 47, 142–153, <a href="https://doi.org/10.1016/j.atmosenv.2011.11.021" target="_blank">https://doi.org/10.1016/j.atmosenv.2011.11.021</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
      
Reinsel, G. C., Weatherhead, E., Tiao, G. C., Miller, A. J., Nagatani, R.
M., Wuebbles, D. J., and Flynn, L. E.: On detection of turnaround and
recovery in trend for ozone, J. Geophys. Res.-Atmos., 107, ACH 1-1–ACH 1-12,
<a href="https://doi.org/10.1029/2001JD000500" target="_blank">https://doi.org/10.1029/2001JD000500</a>, 2002.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>39</label><mixed-citation>
      
Rencher, A. C.: Methods of multivariate analysis, 2nd edn., J. Wiley, New
York, <a href="https://doi.org/10.1002/0471271357" target="_blank">https://doi.org/10.1002/0471271357</a>, 2002.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>40</label><mixed-citation>
      
Schnell, J. L., Prather, M. J., Josse, B., Naik, V., Horowitz, L. W., Cameron-Smith, P., Bergmann, D., Zeng, G., Plummer, D. A., Sudo, K., Nagashima, T., Shindell, D. T., Faluvegi, G., and Strode, S. A.: Use of North American and European air quality networks to evaluate global chemistry–climate modeling of surface ozone, Atmos. Chem. Phys., 15, 10581–10596, <a href="https://doi.org/10.5194/acp-15-10581-2015" target="_blank">https://doi.org/10.5194/acp-15-10581-2015</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>41</label><mixed-citation>
      
Schröder, S., Schultz, M. G., Selke, N., Sun, J., Ahring, J., Mozaffari, A., Romberg, M., Epp, E., Lensing, M., Apweiler, S., Leufen, L. H., Betancourt, C., Hagemeier, B., and Rajveer, S.: TOAR Data Infrastructure, Version v1.0, FZ-Juelich B2SHARE [data set], <a href="https://doi.org/10.34730/4d9a287dec0b42f1aa6d244de8f19eb3" target="_blank">https://doi.org/10.34730/4d9a287dec0b42f1aa6d244de8f19eb3</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>42</label><mixed-citation>
      
Schultz, M. G., Schröder, S., Lyapina, O., Cooper, O., Galbally, I.,
Petropavlovskikh, I., Von Schneidemesser, E., Tanimoto, H., Elshorbany, Y.,
Naja, M., Seguel, R., Dauert, U., Eckhardt, P., Feigenspahn, S., Fiebig, M.,
Hjellbrekke, A.-G., Hong, Y.-D., Christian Kjeld, P., Koide, H., Lear, G.,
Tarasick, D., Ueno, M., Wallasch, M., Baumgardner, D., Chuang, M.-T.,
Gillett, R., Lee, M., Molloy, S., Moolla, R., Wang, T., Sharps, K., Adame,
J. A., Ancellet, G., Apadula, F., Artaxo, P., Barlasina, M., Bogucka, M.,
Bonasoni, P., Chang, L., Colomb, A., Cuevas, E., Cupeiro, M., Degorska, A.,
Ding, A., Fröhlich, M., Frolova, M., Gadhavi, H., Gheusi, F., Gilge, S.,
Gonzalez, M. Y., Gros, V., Hamad, S. H., Helmig, D., Henriques, D.,
Hermansen, O., Holla, R., Huber, J., Im, U., Jaffe, D. A., Komala, N.,
Kubistin, D., Lam, K.-S., Laurila, T., Lee, H., Levy, I., Mazzoleni, C.,
Mazzoleni, L., McClure-Begley, A., Mohamad, M., Murovic, M., Navarro-Comas,
M., Nicodim, F., Parrish, D., Read, K. A., Reid, N., Ries, L., Saxena, P.,
Schwab, J. J., Scorgie, Y., Senik, I., Simmonds, P., Sinha, V., Skorokhod,
A., Spain, G., Spangl, W., Spoor, R., Springston, S. R., Steer, K.,
Steinbacher, M., Suharguniyawan, E., Torre, P., Trickl, T., Weili, L.,
Weller, R., Xu, X., Xue, L., and Zhiqiang, M.: Tropospheric Ozone Assessment
Report: Database and Metrics Data of Global Surface Ozone Observations, Elementa: Science of the Anthropocene, 5, 58, <a href="https://doi.org/10.1525/elementa.244" target="_blank">https://doi.org/10.1525/elementa.244</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>43</label><mixed-citation>
      
Schum, D. A.: The evidential foundations of probabilistic reasoning,
Northwestern University Press, Evanston, Ill., ISBN-13:&thinsp;978-0810118218, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>44</label><mixed-citation>
      
Scully-Allison, C., Le, V., Fritzinger, E., Strachan, S., Harris, F. C., and
Dascalu, S. M.: Near Real-time Autonomous Quality Control for Streaming
Environmental Sensor Data, Procedia Comput. Sci., 126, 1656–1665,
<a href="https://doi.org/10.1016/j.procs.2018.08.139" target="_blank">https://doi.org/10.1016/j.procs.2018.08.139</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>45</label><mixed-citation>
      
Sofen, E. D., Bowdalo, D., Evans, M. J., Apadula, F., Bonasoni, P., Cupeiro, M., Ellul, R., Galbally, I. E., Girgzdiene, R., Luppo, S., Mimouni, M., Nahas, A. C., Saliba, M., and Tørseth, K.: Gridded global surface ozone metrics for atmospheric chemistry model evaluation, Earth Syst. Sci. Data, 8, 41–59, <a href="https://doi.org/10.5194/essd-8-41-2016" target="_blank">https://doi.org/10.5194/essd-8-41-2016</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>46</label><mixed-citation>
      
Steinacker, R., Mayer, D., and Steiner, A.: Data Quality Control Based on
Self-Consistency, Mon. Weather Rev., 139, 3974–3991,
<a href="https://doi.org/10.1175/MWR-D-10-05024.1" target="_blank">https://doi.org/10.1175/MWR-D-10-05024.1</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>47</label><mixed-citation>
      
Tiao, G. C., Reinsel, G. C., Xu, D., Pedrick, J. H., Zhu, X., Miller, A. J.,
DeLuisi, J. J., Mateer, C. L., and Wuebbles, D. J.: Effects of autocorrelation and temporal sampling schemes on estimates of trend and
spatial correlation, J. Geophys. Res., 95, 20507, <a href="https://doi.org/10.1029/JD095iD12p20507" target="_blank">https://doi.org/10.1029/JD095iD12p20507</a>, 1990.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>48</label><mixed-citation>
      
Tilmes, S., Lamarque, J.-F., Emmons, L. K., Conley, A., Schultz, M. G., Saunois, M., Thouret, V., Thompson, A. M., Oltmans, S. J., Johnson, B., and Tarasick, D.: Technical Note: Ozonesonde climatology between 1995 and 2011: description, evaluation and applications, Atmos. Chem. Phys., 12, 7475–7497, <a href="https://doi.org/10.5194/acp-12-7475-2012" target="_blank">https://doi.org/10.5194/acp-12-7475-2012</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>49</label><mixed-citation>
      
Tong, Y. L.: The multivariate normal distribution, Springer-Verlag, New
York, ISBN:&thinsp;978-1-4613-9655-0, 1990.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>50</label><mixed-citation>
      
Waterman, M. S. and Whiteman, D. E.: Estimation of probability densities by
empirical density functions, International Journal of Mathematical Education in Science and Technology, 9, 127–137, <a href="https://doi.org/10.1080/0020739780090201" target="_blank">https://doi.org/10.1080/0020739780090201</a>, 1978.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>51</label><mixed-citation>
      
Weatherhead, E. C., Reinsel, G. C., Tiao, G. C., Meng, X.-L., Choi, D.,
Cheang, W.-K., Keller, T., DeLuisi, J., Wuebbles, D. J., Kerr, J. B.,
Miller, A. J., Oltmans, S. J., and Frederick, J. E.: Factors affecting the
detection of trends: Statistical considerations and applications to
environmental data, J. Geophys. Res.-Atmos., 103, 17149–17161, <a href="https://doi.org/10.1029/98JD00995" target="_blank">https://doi.org/10.1029/98JD00995</a>, 1998.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>52</label><mixed-citation>
      
Weatherhead, E. C., Reinsel, G. C., Tiao, G. C., Jackman, C. H., Bishop, L.,
Frith, S. M. H., DeLuisi, J., Keller, T., Oltmans, S. J., Fleming, E. L.,
Wuebbles, D. J., Kerr, J. B., Miller, A. J., Herman, J., McPeters, R.,
Nagatani, R. M., and Frederick, J. E.: Detecting the recovery of total column
ozone, J. Geophys. Res.-Atmos., 105, 22201–22210, <a href="https://doi.org/10.1029/2000JD900063" target="_blank">https://doi.org/10.1029/2000JD900063</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>53</label><mixed-citation>
      
Wincek, M. A. and Reinsel, G. C.: An Exact Maximum Likelihood Estimation
Procedure for Regression-<i>ARMA</i> Time Series Models with Possibly
Nonconsecutive Data, J. Roy. Stat. Soc. B Met., 48, 303–313, <a href="https://doi.org/10.1111/j.2517-6161.1986.tb01414.x" target="_blank">https://doi.org/10.1111/j.2517-6161.1986.tb01414.x</a>, 1986.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>54</label><mixed-citation>
      
Zahumenský, I.: Guidelines on Quality Control Procedures for Data from
Automatic Weather Stations, World Meteorological Organization, 11 pp.,
<a href="https://www.researchgate.net/publication/228826920_Guidelines_on_Quality_Control_Procedures_for_Data_from_Automatic_Weather_Stations" target="_blank"/> (last access: 15 June 2023), 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>55</label><mixed-citation>
      
Zhang, Y., Bocquet, M., Mallet, V., Seigneur, C., and Baklanov, A.: Real-time
air quality forecasting, part II: State of the science, current research
needs, and future prospects, Atmos. Environ., 60, 656–676,
<a href="https://doi.org/10.1016/j.atmosenv.2012.02.041" target="_blank">https://doi.org/10.1016/j.atmosenv.2012.02.041</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>56</label><mixed-citation>
      
Zhou, Y., Chang, F.-J., Chang, L.-C., Kao, I.-F., and Wang, Y.-S.: Explore a
deep learning multi-output neural network for regional multi-step-ahead air
quality forecasts, J. Clean. Prod., 209, 134–145,
<a href="https://doi.org/10.1016/j.jclepro.2018.10.243" target="_blank">https://doi.org/10.1016/j.jclepro.2018.10.243</a>, 2019.

    </mixed-citation></ref-html>--></article>
