Articles | Volume 19, issue 15
https://doi.org/10.5194/amt-19-5071-2026
https://doi.org/10.5194/amt-19-5071-2026
Research article
 | 
06 Aug 2026
Research article |  | 06 Aug 2026

Improving multi-modal wind speed prediction of short and medium term with a bi-clustered machine learning method

Yan Zhang, Lei Li, Xiong Xiong, Xiang Yin, Xiaojun Zhang, Fuhai Cui, Rui Dang, Wei Liu, Liang Zhai, Pengzhao Wang, Peng Sun, Weixiao Lu, and Wenjie Zhang
Abstract

Accurate prediction of wind speed is of great importance for stable and reliable operation of wind farms. However, the single numerical model forecast cannot provide precise wind speed outputs due to the defect of its physical parameterization scheme, whose error will gradually grow with increasing prediction time. Therefore, we proposed a model named Bi-clustered Recursive Bayesian Forest (BCRBR) for wind speed prediction and correction. The approach incorporated Sea-land Breeze and weather stability effects, integrating an atmospheric circulation index as input features; wind farm data underwent modal classification via bi-clustering to mitigate wind speed magnitude interactions, followed by machine learning-based correction of wind speed. The method was proved to be effective for wind speed prediction correction. Compared to forecasts from the Weather Research and Forecasting model, wind speed error indicators were reduced by more than 60 %; and the forecast precision increased from 30.2 % to 78.4 %, of which the improvement is more than twice. Compared to other models, the proposed model presented favorable correction results in different types of wind field, indicating its greater versatility and stronger competitiveness than other models.

Share
1 Introduction

The demand for renewable energy is increasing globally due to the depletion of non-renewable fossil fuels and the deterioration of the ecological environment, leading to a gradual shift towards new energy as the primary power source. In 2024, the global newly installed capacity of wind power reached a record high of 117 GW, with the cumulative installed capacity reaching 1136 GW, an increase of 11 % compared to 2023 (Global Wind Energy Council, 2025). The increasing need for wind power is a positive sign for the energy transition in line with the goals of carbon neutrality. Offshore wind farms can generate more energy compared to onshore wind farms. As a crucial source of clean energy, its consistent and reliable operation enables better integration of a large volume of wind energy, thereby improving the stability of the power system and improving the efficiency of the generation (Enevoldsen and Valentine, 2016). Wind speed usually has a close and complex relationship with the output power of wind power generation, and wind speed prediction errors will be directly transferred to the wind power prediction (Liu et al., 2020). The power output of the wind farm shows significant variability and uncertainty due to the complex oceanic meteorological conditions, the atmospheric laminar flow, and other factors influencing the speed of the wind (Xiong et al., 2023; Zheng et al., 2016). Rapid fluctuation of wind power can disrupt the balance of supply, demand, and safe operation of the power system. In extreme weather conditions, it can even lead to widespread blackouts and the collapse of the power grid, resulting in significant economic losses to society (Khazaei et al., 2022; Yildiz et al., 2021; Xu et al., 2022; Demolli et al., 2019). Therefore, precise forecasting of wind speed is imperative for the achievement of renewable energy development objectives.

Based on the fundamental principles and operational mechanisms of the model, current wind speed prediction methods can be broadly categorized into two groups: physical approaches and statistical techniques (Wang et al., 2022; He et al., 2018). Numerical Weather Prediction (NWP), a physical approach, utilizes computer models to simulate the evolution of atmospheric systems for weather forecasting. It is capable of simulating large-scale meteorological processes, including the progression of wind speed, making it particularly suitable for long-term prediction of wind speed in the context of wind farms (Zhao et al., 2018; Son and Jung, 2021; Brotzge et al., 2023). Common NWP models include the high-resolution limited area model (HIRLAM) (Landberg, 1999), Mesoscale Model5 (MM5) (Salcedo-Sanz et al., 2011), European Centre for Medium-Range Weather Forecasts (ECMWF), and Weather Research and Forecasting model (WRF) (Prósper et al., 2019). Currently, the main method for predicting future wind speeds in the field of wind power prediction is the use of the WRF model (Zhou et al., 2023). However, due to factors such as the imperfection of physical parameterization schemes, low resolution, inaccurate terrain, and others, significant errors exist in numerical weather prediction. The process of atmospheric motion is the result of the joint effect of historical and current states. As the duration of the prediction increases, the NWP wind speed prediction error, which is the main input of the intelligent learning model, will accumulate over time, thereby leading to a notable decline in the accuracy of the wind power prediction model. Additionally, the considerable amount of computing time makes it unsuitable for short-term predictions (Zhao et al., 2019; Xu et al., 2021).

Statistical methods encompass both traditional statistical models and machine learning models, typically developed using a large volume of historical monitoring data and meteorological synchronous observation data. They serve as a crucial complement to numerical model approaches (Katinas et al., 2018; Ouarda and Charron, 2021). Traditional statistical models (e.g., time series, regression) do not simulate atmospheric physics; they are computationally efficient but lack nonlinear fitting ability, leading to lower prediction accuracy (Yousuf et al., 2022). In contrast, machine learning models learn from historical data without strict physical assumptions, effectively capturing complex nonlinear relationships and achieving higher accuracy. They also generalize well with large datasets. However, ML models are data-dependent and require extensive training and tuning (Ley et al., 2022). The typical range of RMSE for WRF offshore wind speed forecasts is currently 1.5–3.5 m s−1, and the error patterns are related to sea-land breezes and atmospheric stability – precisely the kind of nonlinear mapping that machine learning can capture (Gong et al., 2025; Kang et al., 2020). Since it is difficult to completely eliminate these errors by simply improving the WRF physical model or data assimilation, recent studies have begun to use machine learning methods to post-process and correct WRF outputs (Salvão et al., 2025; Tsai et al., 2021; Zhou and He, 2017).

With the development of research for a long time, the accuracy of any single prediction method is almost saturated. Therefore, in practical engineering applications, the single model should be supplemented with other methods to achieve high-precision wind power prediction. Parri and Teeparthi (Parri and Teeparthi, 2024) introduced a hybrid wind speed prediction model (SVMD-TF-QS) that integrates a novel query selection mechanism (QS), continuous variational mode decomposition (SVMD), and a Transformer (TF)-based model to accurately forecast wind speed while minimizing computational load. Zheng and Wang (2024) combined several algorithms based on recurrent neural networks and used the Levy crystal structure algorithm for weight optimization to create a short-term wind speed prediction model. These hybrid models effectively leveraged the strengths of individual models to improve forecast accuracy and reliability. Prediction models based on multi-algorithm fusion will become an important development direction in the field of wind speed prediction.

In offshore wind farms, wind speed is influenced by complex meteorological processes, among which sea-land breeze (SLB) plays a critical role due to diurnal temperature contrasts (Shen, 2021). SLB modulates turbulence intensity and wind speed gradients, potentially altering wind-power characteristics.However, to date, few studies have considered the impacts of weather systems, terrain, and day-night alternation on wind speed. In summary, the following problems remain to be solved in the field of wind speed prediction:

  1. The feature factors input to the corrected model for wind speed prediction have significant limitations. Traditional forecast revision models use historical wind speed data as input and have not considered the interactions between meteorological factors. It does not have any feature selection and lacks feature quantities that can indicate the trend of weather stability;

  2. The factors influencing the wind speed prediction correction method are not comprehensive enough. Previous studies have ignored the influence of meteorology and the principles of statistical models. The offshore wind field should consider the interaction between SLB, large-scale circulation, and other weather systems, as well as high and low wind speed values in the models. Due to the randomness of wind speed and the localized nature of meteorological features, wind speed correction requires a variety of models.

To solve the above two problems, this study proposed a model called Bi-clustered Recursive Bayesian Forest (BCRBR), which is based on a variety of machine learning methods and the idea of bi-clustering combined with the atmospheric circulation data to achieve the wind speed prediction correction of offshore wind farms.

To address the above-mentioned issues, the contributions of this paper are as follows:

  1. The input feature factors of the hidden feature-rich model were extracted from various meteorological elements predicted by NWP. Considering the background field of atmospheric circulation, the atmospheric circulation index (ACI) was added as the quantity of features indicating the trend of weather stability, and the input features were filtered by the recursive feature elimination (RFE) method.

  2. The idea of bi-clustering was adopted to classify the meteorological events from two perspectives, namely, SLB and weather stationarity. The machine learning algorithm was strengthened by the optimization algorithm combined with the multi-modal classification results to realize the multivariate and multi-scale wind speed prediction correction.

  3. The excellent performance of this method for wind speed prediction correction was verified by experiments. After testing and evaluating offshore wind farms in Jiangsu Province, the model was applied to mountainous wind farms in Jiangxi Province to verify its robustness and the validity of the patterns was examined through ablativity experiments.

The organization of this paper is as follows: Sect. 1 introduces the research status and shortcomings in the field of wind speed prediction correction. Section 2 introduces the data method and the modeling process applied in this study. Section 3 analyzed the experimental results of the BCRBR model in detail and Sect. 4 summarized the results of the model and discussed future directions for improvement.

2 Data and methods

2.1 Data

The target wind farm in this study is an offshore wind farm located in Jiangsu Province, China, which belongs to the northern subtropical monsoon climate zone with flat terrain. Wind speeds are higher in summer and relatively stable in winter. Sea–land breezes alternate markedly, and offshore wind directions vary frequently, predominantly from the east. This study utilized wind farm data from April to August 2023. Meteorological information, including wind speed and direction at hub height, was obtained from a meteorological mast installed within the farm. Power data at the individual turbine and station level were collected via the SCADA system mounted at the turbine nacelles. All data have a temporal resolution of 15 min. The installed capacity of the wind farm is 300 MW, comprising 96 Myse 3.0–135 turbines with a rotor diameter of 135 m and a hub height of 90 m. The annual mean wind speed at hub height is 7.3 m s−1. Since the forecast wind speed provided by NWP represents the regional wind speed over the wind farm area, and the target variable for prediction is the station-level wind speed, we used the average wind speed of the 96 turbines at the same time (i.e., the station-level wind speed) as the observed value for prediction. This approach also ensures the integrity and continuity of the wind speed data.

Since large-scale offshore wind farms are under complex meteorological conditions, the introduction of the ACI can reflect the meteorological phenomena related to wind speed changes, which can help to better adjust the bias and improve the prediction of future wind speed. Therefore, we added ACI as the quantity of characteristics in the wind speed correction. ACI includes the strength of the East Asian Major Trough (CQ) and the area, strength, ridge line, and west extension ridge point of Subtropical High (GM, GQ, GX, and GD). The CQ is standardized by the height field of 500 hPa of the ERA5 reanalysis data of 110–145° E and 25–45° N. GM, GQ, GX and GD were calculated using the average data of the height field of 500 hPa 0, 6, 12, 18 h of the ERA5 reanalysis data (Wang et al., 2021).

2.2 Methods

2.2.1 Bi-clustered Recursive Bayesian Forest Model

The BCRBR model proposed in this study contains four patterns which are the biclustered meteorological pattern classification model (BCMMC), RFE, Bayesian optimization (BO) and Random Forest (RF) algorithm. First, considering the influence of SLB generated by day and night changes and weather stability on wind speed, the BCMMC model was used for the modal classification of wind speed data and added ACI as input features. According to the division of the historical meteorological environment label and the prediction of the target data set label, the modal matching of the target data set and the historical data set was performed, to reduce the interaction between the extreme values of the wind speed of the input model. Furthermore, to reduce the complexity of the correction model and overfitting to enhance interpretability, the input features of the wind speed correction model were screened using the RFE method. Finally, the RF regression algorithm was chosen for the final correction of wind speed, and the parameters were optimized by the BO algorithm before each model training.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f01

Figure 1BCRBR model flow diagram to correct for predicted wind speed by WRF.

Table 1Names of input features and their abbreviations.

Download Print Version | Download XLSX

The flow chart of the hybrid model used in this study is shown in Fig. 1 and consists of the following three steps.

  • Step 1 is related to data preprocessing where we used three initial datasets, namely WRF forecast data, observed wind speed data and atmospheric circulation data. Twelve groups of data were selected from the WRF forecast, including 10, 90, 110, and 130 m wind speeds, 2 m temperature, 2 m relative humidity, surface air pressure, and precipitation. The forecast data were interpolated to the latitude and longitude of the target wind farm using a bilinear interpolation method from the WRFOUT grid point weather forecast data. The atmospheric circulation data included five ACIs. All data was fused on the basis of timestamp linkage to construct the input factors for the BCRBR model. Missing values and outliers were removed from the dataset and the data was standardized to ensure a more stable and efficient training process for the model. Table 1 shows the feature names and their abbreviations of the input model.

  • Step 2 of Fig. 1 Offshore Wind Farm Data Modal Split. In this study, data from April to July 2023 was used as the historical data used for model training, which is the training set. The data for August 2023 was used as the target dataset used to predict revisions, also known as the test set. A plot of the daily variation of the ACI is given in Fig. 2. The strength of the atmospheric circulation system increased markedly in August and the difficulty of correcting the target dataset increased, especially during the time framed by the black dotted line, when the ACI as a whole fluctuated considerably. We processed the divided historical dataset as well as the target dataset by the BCMMC model to obtain multiple patterns that have been matched between the training set and the test set to form a multi-modal dataset.

  • Step 3: Feature engineering and model hyperparameter optimization. For each pattern, the RF algorithm was used first to establish regression models and then RFE was used to screen input features. In this paper, 10 % of the training set was randomly divided as the validation set during model training for hyperparameter optimization, as well as evaluation of the model fit and generalization ability. The trained model was applied to the test set to obtain the corrected wind speed data, and the accuracy of the model was finally evaluated by the wind speed evaluation index.

To comprehensively verify and evaluate the performance, universality and effectiveness of each pattern of the model, the comparison experiment, the robustness experiment, and the ablativity experiment of the model were carried out.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f02

Figure 2Daily variation of selected ACI during the study period from April 2023 to August 2023.

Download

2.2.2 WRF model

In this study, the WRF 4.2 model developed by the National Center for Environmental Prediction (NCEP) is used, which has the characteristics of portability, extensibility, high efficiency, multiple nesting and rich parametric scheme design (Skamarock et al., 2019). Combined with the three-layer grid nesting configuration, the prediction region is shown in Fig. 3. The number of grids is 150 × 150, 90 × 90 and 150 × 180, and the horizontal grid resolutions were 9, 3, and 1 km, respectively. The center points of the grid were set at 34° N and 120° E. The system is updated every 12 h, once at 00:00 UTC and once at 12:00 UTC, with forecasts for the next 7 d. Considering that the time scale of the meteorological station data in the study area is 15 min, the time interval of the forecast data from the WRF model is also set to 15 min. The meteorological factors selected for the forecast include wind speed in the 10, 90, 110, and 130 m wind directions, temperature of 2 m, relative humidity of 2 m, surface air pressure, and precipitation. Use the bilinear interpolation method to interpolate the WRFOUT grid-based weather forecast data to the latitude and longitude of the target wind farm.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f03

Figure 3Schematic diagram of the simulation area of the WRF model.

2.2.3 Bi-clustered Meteorological Modal Classification Model

Taking into account the complex meteorological conditions at sea, we proposed using the BCMMC model (Lu, 2025; Sun et al., 2009). A biclustered modal classification of meteorological data was performed from two perspectives: the SLB of mesoscale meteorological systems and weather stability. The first clustering divided the dataset according to the time of day and night to improve the applicability of different periods. Since the synergy of different factors in the large-scale weather system can interfere with the accuracy of wind speed prediction, the second clustering was performed according to the interaction mechanism between different meteorological elements. The structure of the BCMMC model is shown in Fig. 4 in the form of a flowchart.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f04

Figure 4BCMMC model flow diagram.

Download

Bi-clustered modal classification

Sea breezes start from late morning to noon, whereas land breezes start at midnight and end at noon (Gille, 2005). According to domestic meteorological standards, 00:00 and 16:00 UTC were used as time points to divide the day and night periods. Temperature and wind speed have a direct impact on weather conditions and can be used as indicators to describe the stability of the weather (Ren et al., 2011).

This study used the K-means clustering method (Lloyd, 1982; MacQueen, 1967) to classify temperature and wind speed as classification features for weather stability. The Euclidean distance is used to measure the similarity between data objects, and the similarity is inversely proportional to the distance (Sinaga and Yang, 2020). By presetting the initial number of clustering centers, clustering can be performed based on the distance of data objects from the centers. During the clustering process, the center positions are continuously updated to minimize the intra-class variance (SSE). The clustering process ends when the SSE is stable or the objective function converges.

Label prediction

After the historical data set was divided into patterns, the label prediction of the target data set was needed to match the corresponding historical data, and the multi-modal model was formed by combining them. The Deep Forest algorithm (DF) (Zhou and Ji, 2019) is a new tree-based model that can be comparable to the deep neural network proposed by Professor Zhou Zhihua and Dr. Feng Ji on 28 February 2017. Its structure is shown in Fig. 5. DF is a powerful ensemble learning method, also known as gcForest with a multigranularity scan. These are the two core concepts in DF, cascade forest, and multigranularity scans. DF combines the advantages of RF and deep learning, and has advantages in handling high-dimensional data, automatic feature selection, and ensemble learning, making it a powerful classification method with good generalization ability and robustness. In addition, DF is naturally resistant to overfitting (with its cross-validation process). A better result can be obtained without any parameter adjustment.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f05

Figure 5The structure of DF.

Download

2.2.4 Elimination of recursive characteristics

RFE (Guyon et al., 2002) is a feature selection method used to select the most important features to improve model performance or reduce computational cost. It works by recursively training the model and removing features with minimal impact on performance in each round, continuing this process until a specified number of features or a performance threshold is reached. Therefore, we focus on the features that substantially contribute to the performance of the model and improve the generalizability and efficiency of the model (Lee et al., 2022).

2.2.5 Parameter optimization

Bayesian Optimization

The BO algorithm (Mockus, 1975) is commonly used to optimize machine learning models or other tasks that require parameter tuning (Shahriari et al., 2016). BO uses Bayesian statistical inference to construct a probabilistic model of the parameter space and updates the model based on historical observations at each iteration. This allows for smarter selection of the next combination of parameters to evaluate to maximize the outcome of the objective function. Here are the steps to implement the BO method:

  • Step 1: Define the objective function. First, define the objective function f(x) to be optimized, where x is a set of hyperparameters.

  • Step 2: Initialize the Gaussian process: Construct a Gaussian process GPm(x),kx,x, where m(x) is the mean function and kx,x is the kernel function. Usually, set m(x)=0 and choose an appropriate kernel function, such as the Gaussian kernel:

    (1) k x , x = exp - x - x 2 2 σ 2
  • Step 3: Initial sampling. Select an initial set of hyperparameters x1, x2, …, xn, compute the corresponding objective function values y1, y2, …, yn, and use this data to construct the initial Gaussian process model.

  • Step 4: Update the Gaussian process model: Based on the current Gaussian process model, calculate the next hyperparameter xn+1 to sample, such that the expected improvement (EI) is maximized:

    (2) x n + 1 = argmax x EI ( x ) = E max f ( x ) - f x + , 0

    where x+ is the current best known hyperparameter.

  • Step 5: Iterative optimization. For the newly sampled hyperparameter xn+1, compute the objective function value yn+1 and update the Gaussian process model. Repeat step 4 until the specified number of iterations or convergence criteria are met.

  • Step 6: Output the optimal hyperparameters. Throughout the optimization process, record all the hyperparameters sampled and their corresponding objective function values. Finally, output the set of hyperparameters that corresponds to the minimum objective function value as the optimal hyperparameters.

10-fold cross-validation method

At each step of using the BO algorithm, 10-fold cross-validation is used to evaluate the performance of the current hyperparameter configuration to avoid BO falling into a locally optimal solution. 10-fold cross-validation (Breiman et al., 1984) is a commonly used evaluation method for machine learning models to evaluate their generalizability on unseen data (Stone, 1974). In each iteration, the dataset is randomly divided into 10 equal-sized subsets (folds). One of the folds is selected as the validation set and the remaining 9 folds are used as the training set to train the model and evaluate the performance on the validation set. This process is repeated 10 times, each time using a different validation fold. The performance metrics of the 10 times are averaged as the final evaluation result of the model.

2.2.6 Random forest algorithm

The RF algorithm (Breiman, 2001) combines multiple decision trees to perform prediction and classification tasks. The RF modeling process is as follows.

First, define the wind speed prediction training set XiYi, where Yi is the real value in the RF prediction model, mapped to the real value of wind speed of the ith sample in the data; Xi is the feature vector established by the meteorological elements of the ith sample in the data, and is denoted by {Ii1,Ii2,,Iin}Xi denote the n influence factors of the ith sample.

Next, based on determining the training set, a single regression decision tree is built. Through the feature vector X and its corresponding true value Y in the training sample, the split variable and split value are searched, and the regression decision tree divides the whole vector space into m partitions {R1,R2,,Rn}. For any partition of which can be mapped to the model Cm, the vector space is divided into two parts by the value of a feature, and the expression is

(3)R1j,s=I|Ijs(4)R2j,s=I|Ijs

where, j represents an influencing factor and s represents the value when splitting. The objective function for performing the vector space split variable and split value search is

(5) z : min j s min c 1 x i ϵ R i j s y i - c i 2 + min c 2 x i ϵ R 2 j s y i - c 2 2

where, z is the minimum variance of the real value of wind speed; yi is the real wind speed value of the ith sample; xi is the corresponding value of the ith sample's impact factor vector; c1 is the real mean value of wind speed in the first part; c2 is the mean real value of wind speed in the second part.

Finally, build the entire RF based on a single decision tree. The RF prediction value is the average of the predicted values of all decision trees.

2.2.7 Evaluation metrics

In this study, the effect of the model was evaluated by statistical test. We use the following statistics: Root Mean Square Errors (RMSE), relative mean absolute error (rRMSE), and mean absolute error (MAE), where smaller values indicate smaller errors. The mean absolute percentage error (MAPE) has a value range of [0, +∞], the smaller the value, the better the prediction model has, a MAPE of 0 indicates a perfect model, and a MAPE greater than 100  % indicates a poor model. R, also known as the correlation coefficient, was used to verify the degree of “linear” relationship predicted by the model. The formulas for the four statistics are given below:

(6)RMSE=1ni-1ny^i-yi2(7)rRMSE=1ni=1nyi-y^i2ini=1nyi×100%(8)MAE=i=1nyi-y^in(9)MAPE=100ni=1nyi-y^iyi(10)R=i=1nxi-xyi-yi=1nxi-x2i=1nyi-y2

where y^i is the ith predicted value, yi is the ith observation, n is the total number of time samples, and y is the average of the observations. Forecast accuracy (FA) refers to the percentage of absolute deviations in wind speed forecasts that are not greater than 1 m s−1. The formula is as follows:

(11) FA = N r N f × 100 %

where Nr is the number of samples where the absolute deviation of the wind speed forecast is not more than 1 m s−1, and Nf is the number of samples forecasted.

3 Results and discussion

3.1 Results of the BCMMC model

The preprocessed training and test sets were loaded into the BCMMC model for modal classification. After the day and night division, the training set was divided into 4 clusters for each of the daytime and nighttime hours using K-means clustering. A total of 8 groups were waiting to be matched with the test set. Table 2 indicated that the data distribution within each group was concentrated over 2–3 months with a uniform time distribution. There was no concentration in any particular month, demonstrating a good clustering effect.

Table 2Distribution of bi-clustered data. The bold clusters are clusters that were matched to the test set to form multimodal data.

Download Print Version | Download XLSX

Each cluster represented the meteorological characteristics of different periods, which was conducive to subsequent analysis and modeling.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f06

Figure 6Results of the BCMMC model. (a) Bi-clustering classification, and (b) label prediction.

Download

The daytime and nighttime periods of the test set were processed by the DF algorithm, and two clusters were predicted for each, as shown in Fig. 6. These four clusters corresponded to the two scenarios of low and high wind speeds under hot weather conditions, reflecting the stability of the weather or not. The clusters of the test set were matched with the clusters of the corresponding training set to form multi-modal data. The BCMMC model was divided into four modes, and modes 1, 2, 3, and 4 corresponded to the four meteorological features of daytime smooth, daytime smooth, nighttime smooth, and nighttime smooth in August 2023, respectively.

3.2 Results of RFE and Parameter Optimization

The RF regression models were established for the four patterns and the parameters were optimized. The correlation analysis (Fig. 7) showed that the actual wind speeds of the patterns were strongly correlated with the wind speeds at different levels, meteorological factors, and some ACIs, indicating that the other factors and the large-scale circulation system have an important influence on wind speed prediction. The importance ranking of the features (Fig. 8) showed that the wind speed at different levels was the main characteristic, while the contributions of ACI and other meteorological factors varied according to the patterns, and the fluctuation of the RMSE with the number of features (Fig. 9) in the RFE screening of the input features of the modal models was due to the correlation between the features and the small amount of data.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f07

Figure 7Correlation between variables of the 4 patterns. Panels (a)(d) represent patterns 1, 2, 3, and 4 respectively. The red dashed boxes represent highly relevant meteorological elements; blue dashed boxes represent highly relevant ACI.

Download

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f08

Figure 8Importance of the 4 patterns. Panels (a)(d) represent patterns 1, 2, 3, and 4 respectively. The framed features represent the final input.

Download

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f09

Figure 9The result of RFE, variation of each pattern error (RMSE) with the number of features.

Download

The number of features with the lowest error was chosen as the final input features of the models: pattern 1 had 7 final inputs of 90, 130, 110, 10 m wind speed, 2 m relative humidity, and 2 m temperature with CQ; pattern 2 had all 17 features selected as final inputs; pattern 3 had 5 features of 90, 10 m wind speed, and 2 m relative humidity, and 130 m wind speed, and 2 m temperature; and pattern 4 had 5 features of 130, 90, 110, 10 m wind speed, and CQ. After feature selection, the best parameters of the models confirmed by the BO algorithm are shown in Table 3.

Table 3The best hyperparameters of the models.

Download Print Version | Download XLSX

The ACI reflected the complex relationship between meteorological variables and could improve the prediction performance of the model by correlated with other characteristics. The input features of all patterns except pattern 3 included the ACI, which indicated that atmospheric circulation played an important role in the correction of the wind speed prediction and had a significant effect on the model prediction. Pattern 2 was a non-smooth period during the daytime, which was affected by daylight, topography, and other factors, with frequent and complex changes of large-scale weather systems, and required highly correlated multi-featured factors as inputs to characterize the complex relationship with the fast-changing wind speeds. Pattern 3 was a stable period at night, with weak activity of large-scale weather systems and relatively static pressure, and no ACI was needed as a reference. Patterns 1 and 4 only screened CQ, a weather system that is an important indicator of changes in the offshore wind farm.

3.3 Results of wind speed correction

The corrected results for the four patterns were exported and aggregated to obtain the corrected results for the final August 2023 test set. The statistical errors are shown in Table 4, and the wind speed errors of each pattern after the correction of the BCRBR model were reduced to different degrees compared to the WRF forecast, the daytime dataset represented by patterns 1 and 2 showing better prediction performance compared to the nighttime dataset represented by patterns 3 and 4. On the contrary, the RMSE of the wind speed in the test set decreased from 2.692 to 0.965 m s−1, an overall reduction of 64.15 %, and the correlation coefficient, R, increased from 0.766 to 0.898. A comparison of the regression scatter density plots for each dataset (Fig. 10) showed that the wind speed forecasts of the WRF system were overall large. Patterns 1 and 3 correspond to a small range of wind speed values; in contrast to Patterns 2 and 4, which correspond to a large range of wind speed values, there were more outliers in the dataset, which would have a certain impact on the model's generalization performance and lead to a decrease in the prediction performance.

Table 4The four patterns and aggregated results. Values in bold are used to highlight the models/datasets that performed best in the experiment.

Download Print Version | Download XLSX

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f10

Figure 10Regression scatter density plots for WRF forecast and the forecast of the BCRBR model. Panels (a), (c), (e), (g) represent patterns 1, 2, 3, 4 respectively for WRF, and panels (b), (d), (f), (h) represent patterns 1, 2, 3, 4 respectively for BCRBR.

Download

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f11

Figure 11Box plot of daily changes in forecast and observed values for August 2023.

Download

The daily variation of wind speed in the test set was further analyzed, and the boxplot of the daily variation of actual wind speed between the WRF forecast and the BCRBR model forecast was shown in Fig. 11. The distribution of the actual data (grey) was relatively concentrated, with a median between 3 and 4 m s−1 and fewer outliers. The distribution of the WRF forecast data (blue) was more discrete, the median was more different from the observed data in different periods, and there were more outliers and obvious overestimation, which indicated that the WRF model had greater uncertainty in simulating the wind speeds. The median of the BCRBR model forecast data (in blue) was in better agreement with the actual observations for most of the periods, but there were some periods of slightly higher or lower values. In addition, as seen in the dashed box portion of Fig. 11b, the wind speed trend was smooth during the daytime hours, while the wind speed during the nighttime hours was larger and subject to larger fluctuations, which made it more difficult to predict, and that is why the wind speed prediction for the two patterns during the daytime hours was more accurate than the two patterns during the nighttime hours.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f12

Figure 12Time series plots of wind speed values, errors, and CQ during abrupt atmospheric circulation changes. (a) Process 1, (b) Process 2.

Download

From the above, it can be seen that wind speed was more difficult to predict in the unstable state of the atmosphere, and the change in atmospheric circulation was closely related to wind speed. Since CQ made a larger contribution compared to other ACIs in the BCRBR model for most patterns, here CQ was used as an indication of the change in weather stability, focusing on the comparison of the wind speed forecasts during the two abrupt ACI changes in the black box in Fig. 4. Since FA was defined as the percentage of wind speed forecasts with absolute deviation not greater than 1 m s−1, here the absolute deviation less than or equal to 1 was defined as a small error; otherwise, it was a large error. Comparing the time series plots of wind speed values and errors in the two processes (Fig. 12), it is easy to find that the WRF system produced a large and concentrated bias in the forecast when the weather was not stable and the wind speed error was significantly reduced after the revision of the BCRBR model. The maximum error of the two processes decreased from 4.21 to 3.11 and 11.18 to 5.25 m s−1, and the average error decreased from 1.43 to 0.54 and 3.26 to 0.73 m s−1, respectively. In particular, the corrected effect of process 2 was a good example of the validity of the BCRBR model.

Table 5The optimal hyper-parameters found for each model.

Download Print Version | Download XLSX

Table 6Comparing the performance of the models in the experiment in terms of error metrics. Values in bold are used to highlight the models/datasets that performed best in the experiment.

Download Print Version | Download XLSX

3.4 Comparison experiment

To comprehensively evaluate the wind speed prediction performance of the BCRBR model, we selected popular machine learning and deep learning methods in the wind power field, such as the RF, ERT, DF, XGBoost, lightGBM, and DNN models, and used the same dataset to predict wind speed for comparison experiments. The optimal hyper-parameters found for each model are summarized in Table 5. Table 6 shows the error metrics of the prediction effects of each model. After comparison, the BCRBR model performed the best in all error metrics and demonstrated good performance in the limited dataset, highlighting its strong competitiveness. The Taylor diagram in Fig. 15a visually displays the correction effects of each model, with the BCRBR model being closest to the reference point, indicating that its prediction results were closest to the measured values, had the smallest error, and exhibited the best prediction performance.

3.5 Ablativity experiment

To verify the validity of each pattern of the BCRBR model, an ablativity experiment was performed on the model. After removing each or some of the patterns in the model, respectively, the results after prediction on the same dataset were shown in the Taylor diagram in Fig. 13b, and the model that was closest to the observed value was still the BCRBR model.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f13

Figure 13Taylor diagrams of the predictions of August 2023 for each model. (a) Comparison experiment, (b) ablativity experiment.

Download

Statistically, it is not difficult to find that BCRBR performed the best in all the error indicators (Table 7). Interestingly, as long as the models containing the BCMMC pattern can get good prediction results, the performance was better than other models, indicating that the BCMMC model played a key role in improving the accuracy of wind speed prediction.

Table 7Performance of each model in ablativity experiment. Values in bold are used to highlight the models/datasets that performed best in the experiment.

Download Print Version | Download XLSX

3.6 Robustness experiment

The BCRBR model was applied to a mountainous wind farm in Jiangxi Province and a robustness experiment was conducted with data from 1 December 2023 to 19 March 2024 to test its generality. The data from December 2023 to February 2024 and March 2024 were used for training and testing, respectively. The time series of WRF predicted values, BCRBR predicted values, observed values, and forecast errors are given in Fig. 14, and the BCRBR model had a better correction effect, with the average error reduced from 1.84 to 1.23 m s−1, and the maximum error decreased from 7.04 to 5.35 m s−1, compared with the prediction of WRF. The blue dotted box in Fig. 14 exhibited the most effective correction, whereas the red dotted box showed a less effective correction, with a focus on high-wind-speed values. There was still potential for further improvement in the model.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f14

Figure 14Time series plot of wind speed values and errors for WRF forecast and the BCRBR model forecast.

Download

The regression scatter density plot of wind speed is shown in Fig. 15, comparing with the WRF forecast, the RMSE of wind speed decreased from 2.32 to 1.56 m s−1, and the correlation increased from 0.75 to 0.88. The linear distribution of the scatter had a more pronounced trend, and the regression line was closer to the congruent line y=x. In summary, the BCRBR model also had a good wind speed correction ability in different regions of the wind farm and a strong generality.

https://amt.copernicus.org/articles/19/5071/2026/amt-19-5071-2026-f15

Figure 15Regression scatter plot of wind speed. (a) WRF forecast, (b) BCRBR model forecast.

Download

In the robustness experiment, the overall accuracy of wind speed prediction for the wind farm in Jiangxi Province of China was slightly lower as compared to the offshore wind farm in Jiangsu Province, mainly because the wind farm in Jiangxi Province is a mountainous wind farm with a more complex topography than an offshore wind farm. In intricate terrain, numerous factors affect wind speed, including the drag effect caused by the ruggedness of mountain surfaces on the atmosphere, as well as localized circulations like valley winds and slope winds, stemming from the uneven heating and cooling of mountainous regions.

In the future, the following methods will be considered to enhance the accuracy of wind speed prediction in mountain areas. Incorporating terrain factors and underlying surface parameters, such as terrain height, slope, roughness, etc., as input features in the wind speed prediction model. Increasing the density and quality of observations, particularly in complex topographic regions, to provide a larger sample size for model training and validation.

4 Conclusions

This study presented a multimodal short- and medium-term wind speed prediction correction approach based on bi-clustering and machine learning, aiming to address the crucial scientific problems of incomplete consideration of influencing factors and limited feature extraction in traditional wind speed prediction. In this paper, the BCRBR model was innovatively constructed, with offshore wind farms as the research object, and the complex influence mechanisms of sea-land breeze and atmospheric circulation on wind speed were thoroughly considered from a meteorological perspective. By introducing bi-clustering strategies and innovative features such as atmospheric circulation indices, precise correction of wind speed predictions was achieved in this study. Compared to the traditional WRF model, the errors of the corrected wind speed predictions were significantly reduced, specifically the following. RMSE, rRMSE, MAE and MAPE all decreased by more than 60 %, the wind speed forecast accuracy rate increased from 30.2 % to 78.4 %, and the correlation coefficient R increased from 0.77 to 0.9.

This research not only offers new thoughts for wind speed prediction in methodology, but also provides data source guarantees to enhance the accuracy of wind power prediction. The research findings have significant theoretical value and practical significance in promoting the large-scale development of wind power and the sustainable development of power systems. The results of comparative experiments demonstrated that the BCRBR model outperforms the prevalent single-model prediction methods in the current literature. The results of the robustness experiments indicated that this method also performed better than the traditional WRF model in the prediction of wind speed in mountain wind farms in terms of various performance indicators, but there was a slight disparity compared to the results in offshore wind farms, which provided an important direction for subsequent improvement.

Future research will focus on the following directions to improve the accuracy of wind speed prediction in mountainous areas:

  1. Integrating terrain factors and surface parameters – such as elevation, slope, and roughness – as input features for wind speed prediction models.

  2. Enhancing the quality of observational data, especially in complex terrain regions, by implementing data quality control protocols to supply more reliable samples for model training and validation.

  3. Refining model parameters and extracting informative features to improve prediction performance under complex terrain conditions, thereby broadening the applicability of the method and supporting efforts to mitigate instability in renewable energy generation.

Appendix A: List of abbreviations
Abbreviations Full name
NWP Numerical Weather Prediction
WRF Weather Research and Forecasting
SLB Sea-land Breeze
SSE The Sum of Squares due to Error
DF Deep Forest
BO Bayesian Optimization
RFE Recursive Feature Elimination
RF Random Forest
RMSE Root Mean Square Errors
rRMSE Relative Root Mean Square Errors
MAE Mean Absolute Error
MAPE Mean Absolute Percentage Error
FA Forecast Accuracy
R Correlation Coefficient
ACI Atmospheric Circulation Index
ERT Extremely Randomized Trees
DNN Deep Neural Networks
BCMMC Bi-clustered Meteorological Modal Classification
RBR Recursive Bayesian Forests
BCRBR Bi-clustered Recursive Bayesian Forest
HIRLAM High-resolution Limited Area Model
MM5 Mesoscale Model 5
ECMWF European Centre for Medium-Range Weather Forecasts
Data availability

The data that has been used is confidential.

Author contributions

Yan Zhang: Funding acquisition, Resources, Supervision, Validation. Lei Li: Project administration, Investigation. Xiong Xiong: Resources, Supervision, Writing (review and editing). Xiang Yin: Visualization. Xiaojun Zhang: Formal analysis. Fuhai Cui: Software. Rui Dang: Data curation. Wei Liu: Formal analysis. Liang Zhai: Visualization. Pengzhao Wang: Validation. Peng Sun: Visualization. Weixiao Lu: Conceptualization, Investigation, Methodology, Formal analysis, Visualization, Writing (original draft preparation). Wenjie Zhang: Supervision.

Competing interests

The contact author has declared that none of the authors has any competing interests.

Disclaimer

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.

Financial support

This research was supported by the National Key R&D Program of China under grant no. 2023YFB4203301 and the National Natural Science Foundation of China under grant no. 42521006.

Review statement

This paper was edited by Simone Lolli and reviewed by two anonymous referees.

References

Breiman, L.: Random forests, Mach. Learn., 45, 5–32, https://doi.org/10.1023/A:1010933404324, 2001. 

Breiman, L., Friedman, J., Olshen, R. A., and Stone, C. J.: Classification and regression trees Chapman & Hall, New York, 22 pp., ISBN 9780412048418, 1984. 

Brotzge, J. A., Berchoff, D., Carlis, D. L., Carr, F. H., Carr, R. H., Gerth, J. J., Gross, B. D., Hamill, T. M., Haupt, S. E., Jacobs, N., McGovern, A., Stensrud, D. J., Szatkowski, G., Szunyogh, I., and Wang, X.: Challenges and Opportunities in Numerical Weather Prediction, B. Am. Meteorol. Soc., 104, E698–E705, https://doi.org/10.1175/BAMS-D-22-0172.1, 2023. 

Demolli, H., Dokuz, A. S., Ecemis, A., and Gokcek, M.: Wind power forecasting based on daily wind speed data using machine learning algorithms, Energ. Convers. Manage., 198, 111823, https://doi.org/10.1016/j.enconman.2019.111823, 2019. 

Enevoldsen, P. and Valentine, S. V.: Do onshore and offshore wind farm development patterns differ?, Energy Sustain. Dev., 35, 41–51, https://doi.org/10.1016/j.esd.2016.10.002, 2016. 

Gille, S. T.: Global observations of the land breeze, Geophys. Res. Lett., 32, L05605, https://doi.org/10.1029/2004GL022139, 2005. 

Global Wind Energy Council: GWEC Global Wind Report 2025, https://www.gwec.net (last access: 23 April 2025), 2025. 

Gong, H., Kan, Y., Yang, D., Huang, W., Fan, C., Hao, R., Sha, L., Deng, M., and Zhang, H.: Machine learning improves mesoscale offshore wind resource assessment over the Yellow Sea, J. Meteorol. Res., 39, 1510–1526, https://doi.org/10.1007/s13351-025-5051-z, 2025. 

Guyon, I., Weston, J., Barnhill, S., and Vapnik, V.: Gene selection for cancer classification using support vector machines, Mach. Learn., 46, 389–422, https://doi.org/10.1023/A:1012487302797, 2002. 

He, Q., Wang, J., and Lu, H.: A hybrid system for short-term wind speed forecasting, Appl. Energ., 226, 756–771, https://doi.org/10.1016/j.apenergy.2018.06.053, 2018. 

Kang, M., Ko, K., and Kim, M.: Verification of the Reliability of Offshore Wind Resource Prediction Using an Atmosphere–Ocean Coupled Model, Energies, 13, 254, https://doi.org/10.3390/en13010254, 2020. 

Katinas, V., Gecevicius, G., and Marciukaitis, M.: An investigation of wind power density distribution at location with low and high wind speeds using statistical model, Appl. Energ., 218, 442–451, https://doi.org/10.1016/j.apenergy.2018.02.163, 2018. 

Khazaei, S., Ehsan, M., Soleymani, S., and Mohammadnezhad-Shourkaei, H.: A high-accuracy hybrid method for short-term wind power forecasting, Energy, 238, 122020, https://doi.org/10.1016/j.energy.2021.122020, 2022. 

Landberg, L.: Short-term prediction of the power production from wind farms, J. Wind Eng. Ind. Aerod., 80, 207–220, https://doi.org/10.1016/S0167-6105(98)00192-5, 1999. 

Lee, M., Lee, J. H., and Kim, D. H.: Gender recognition using optimal gait feature based on recursive feature elimination in normal walking, Expert Syst. Appl., 189, 116040, https://doi.org/10.1016/j.eswa.2021.116040, 2022. 

Ley, C., Martin, R. K., Pareek, A., Groll, A., Seil, R., and Tischer, T.: Machine learning and conventional statistics: making sense of the differences, Knee Surg. Sport. Tr. A., 30, 753–757, https://doi.org/10.1007/s00167-022-06896-6, 2022. 

Liu, X., Zhang, H., Kong, X., and Lee, K. Y.: Wind speed forecasting using deep neural network with feature selection, Neurocomputing, 397, 393–403, https://doi.org/10.1016/j.neucom.2019.08.108, 2020. 

Lloyd, S. P.: Least squares quantization in PCM, IEEE T. Inform. Theory, 28, 129–137, https://doi.org/10.1109/TIT.1982.1056489, 1982. 

Lu, W.: Research on Power Prediction Algorithm for Offshore Wind Farms Based on Dual-Layer Deep Learning, Master's thesis, Nanjing University of Information Science and Technology, https://doi.org/10.27248/d.cnki.gnjqc.2025.000483, 2025. 

MacQueen, J.: Some methods for classification and analysis of multivariate observations, Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Vol. 1, 281–297, ISBN 9780520366701, 1967. 

Mockus, J.: On Bayesian Methods for Seeking the Extremum, Proceedings of the IFIP Technical Conference, 400–404, https://doi.org/10.1007/978-3-662-38527-2_55, 1975. 

Ouarda, T. B. M. J. and Charron, C.: Non-stationary statistical modelling of wind speed: A case study in eastern Canada, Energ. Convers. Manage., 236, 114028, https://doi.org/10.1016/j.enconman.2021.114028, 2021. 

Parri, S. and Teeparthi, K.: SVMD-TF-QS: An efficient and novel hybrid methodology for the wind speed prediction, Expert Syst. Appl., 249, 123516, https://doi.org/10.1016/j.eswa.2024.123516, 2024. 

Prósper, M. A., Otero-Casal, C., Fernández, F. C., and Miguez-Macho, G.: Wind power forecasting for a real onshore wind farm on complex terrain using WRF high resolution simulations, Renew. Energ., 135, 674–686, https://doi.org/10.1016/j.renene.2018.12.047, 2019. 

Ren, C., Ng, Y. Y., and Katzschner, L.: Urban climatic map studies: a review, Int. J. Climatol., 31, 2213–2233, https://doi.org/10.1002/joc.2237, 2011. 

Salcedo-Sanz, S., Ortiz-García, E. G., Pérez-Bellido, Á. M., Portilla-Figueras, A., and Prieto, L.: Short term wind speed prediction based on evolutionary support vector regression algorithms, Expert Syst. Appl., 38, 4052–4057, https://doi.org/10.1016/j.eswa.2010.09.067, 2011. 

Salvão, N., Monteiro, M., and Soares, C. G.: An assessment of two wind model uncertainties during storm events affecting Portugal, Ocean Eng., 333, 121395, https://doi.org/10.1016/j.oceaneng.2025.121395, 2025. 

Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and De Freitas, N.: Taking the human out of the loop: A review of Bayesian optimization, P. IEEE, 104, 148–175, https://doi.org/10.1109/JPROC.2015.2494218, 2016. 

Shen, C.: Climate-Driven Characteristics of Sea-Land Breezes Over the Globe, Geophys. Res. Lett., 48, e2020GL092308, https://doi.org/10.1029/2020GL092308, 2021. 

Sinaga, K. P. and Yang, M. S.: Unsupervised K-Means Clustering Algorithm, IEEE Access, 8, 100156–100172, https://doi.org/10.1109/ACCESS.2020.2988796, 2020. 

Skamarock, W. C., Klemp, J. B., Dudhia, J., Gill, D. O., Barker, D. M., Duda, M. G., Huang, X. Y., Wang, W., and Powers, J. G.: A description of the advanced research WRF version 4, NCAR Technical Note NCAR/TN-556+STR, https://doi.org/10.5065/1dfh-6p97, 2019. 

Son, N. and Jung, M.: Analysis of meteorological factor multivariate models for medium- and long-term photovoltaic solar power forecasting using long short-term memory, Appl. Sci., 11, 316, https://doi.org/10.3390/app11010316, 2021. 

Stone, M.: Cross-validatory choice and assessment of statistical predictions, J Roy. Stat. Soc. B Met., 36, 111–147, https://doi.org/10.1111/j.2517-6161.1974.tb00994.x, 1974. 

Sun, C., Tao, S., Luo, Y., Wang, S., and Song, L.: Land-sea breeze and the application of wind profile in the wind speed forecasting to wind farm along the coast, Chinese J. Geophys., 52, 630–636, 2009. 

Tsai, C.-C., Hong, J.-S., Chang, P.-L., Chen, Y.-R., Su, Y.-J., and Li, C.-H.: Application of bias correction to improve WRF ensemble wind speed forecast, Atmosphere, 12, 1688, https://doi.org/10.3390/atmos12121688, 2021. 

Wang, F., Tong, S., Sun, Y., Xie, Y., Zhen, Z., Li, G., Cao, C., Duić, N., and Liu, D.: Wind process pattern forecasting based ultra-short-term wind speed hybrid prediction, Energy, 255, 124509, https://doi.org/10.1016/j.energy.2022.124509, 2022. 

Wang, Q., Zhang, H., Zong, L., Su, H., Yang, Y., and Gao, Z.: Synoptic cause of a continuous conductor icing event on ultra-high-voltage transmission lines in northern Guangxi in 2015, J. Trop. Meteorol., 37, 579–589, https://doi.org/10.16032/j.issn.1004-4965.2021.055, 2021. 

Xiong, X., Zou, R., Sheng, T., Zeng, W., and Ye, X.: An ultra-short-term wind speed correction method based on the fluctuation characteristics of wind speed, Energy, 283, 129012, https://doi.org/10.1016/j.energy.2023.129012, 2023. 

Xu, W., Liu, P., Cheng, L., Zhou, Y., Xia, Q., Yu, G., and Liu, Y.: Multi-step wind speed prediction by combining a WRF simulation and an error correction strategy, Renew. Energ., 163, 772–782, https://doi.org/10.1016/j.renene.2020.09.032, 2021. 

Xu, Y., Jia, L., and Yang, W.: Correlation based neuro-fuzzy Wiener type wind power forecasting model by using special separate signals, Energ. Convers. Manage., 253, 115173, https://doi.org/10.1016/j.enconman.2021.115173, 2022. 

Yildiz, C., Acikgoz, H., Korkmaz, D., and Budak, U.: An improved residual-based convolutional neural network for very short-term wind power forecasting, Energ. Convers. Manage., 228, 113731, https://doi.org/10.1016/j.enconman.2020.113731, 2021. 

Yousuf, M. U., Al-Bahadly, I., and Avci, E.: Statistical wind speed forecasting models for small sample datasets: Problems, improvements, and prospects, Energ. Convers. Manage., 261, 115658, https://doi.org/10.1016/j.enconman.2022.115658, 2022. 

Zhao, J., Wang, J., Guo, Z., Guo, Y., and Lin, W.: Multi-step wind speed forecasting based on numerical simulations and an optimized stochastic ensemble method, Appl. Energ., 255, 113833, https://doi.org/10.1016/j.apenergy.2019.113833, 2019. 

Zhao, X., Liu, J., Yu, D., and Chang, J.: One-day-ahead probabilistic wind speed forecast based on optimized numerical weather prediction data, Energ. Convers. Manage., 164, 560–569, https://doi.org/10.1016/j.enconman.2018.03.036, 2018. 

Zheng, C., Pan, J., and Li, C.: Global oceanic wind speed trends, Ocean Coast. Manage., 129, 15–24, https://doi.org/10.1016/j.ocecoaman.2016.05.002, 2016. 

Zheng, J. and Wang, J.: Short-term wind speed forecasting based on recurrent neural networks and Levy crystal structure algorithm, Energy, 293, 130580, https://doi.org/10.1016/j.energy.2024.130580, 2024.  

Zhou, R. and He, X.: Effect of model horizontal resolution on the surface wind speed forecast in the Northeast China, Advances in Meteorological Science and Technology, 7, 95–100, 2017. 

Zhou, S., Gao, C. Y., Duan, Z., Xi, X., and Li, Y.: A robust error correction method for numerical weather prediction wind speed based on Bayesian optimization, variational mode decomposition, principal component analysis, and random forest: VMD-PCA-RF (version 1.0.0), Geosci. Model Dev., 16, 6247–6266, https://doi.org/10.5194/gmd-16-6247-2023, 2023. 

Zhou, Z. and Ji, F.: Deep forest, Natl. Sci. Rev., 6, 74–86, https://doi.org/10.1093/nsr/nwy108, 2019. 

Download
Short summary
We developed a Bi-clustered Recursive Bayesian Forest model that improves wind speed forecast accuracy by 48.2 % and reduces error metrics by over 60 %. The model incorporates sea-land breeze, weather stability, and atmospheric circulation indices as features, and uses bi-clustering modal classification to mitigate wind speed magnitude interactions. This machine learning-based correction technique outperforms traditional numerical models, providing more reliable wind speed forecasting.
Share