<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">ACP</journal-id><journal-title-group>
    <journal-title>Atmospheric Chemistry and Physics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">ACP</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Atmos. Chem. Phys.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1680-7324</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/acp-26-12243-2026</article-id><title-group><article-title>Measurement report: Quantifying the trade-off between station number and spatial layout in sparse GNSS networks for calibrating all-weather FY-4A precipitable water vapor</article-title><alt-title>Measurement report: Quantifying the trade-off between station number and spatial layout</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Ma</surname><given-names>Yongchao</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Chen</surname><given-names>Zhengsheng</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff2">
          <name><surname>Liu</surname><given-names>Tong</given-names></name>
          <email>tong2.liu@polyu.edu.hk</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Yu</surname><given-names>Zhibin</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff4">
          <name><surname>Wang</surname><given-names>Zhihao</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Rocket Force University of Engineering, Xi'an, China</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Department of Land Surveying and Geo-Informatics, the Hong Kong Polytechnic University, Hong Kong, China</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>School of Aerospace Science, Harbin Institute of Technology (Shenzhen), Shenzhen, China</institution>
        </aff>
        <aff id="aff4"><label>4</label><institution>Institute of Geospatial Information, Information Engineering University, Zhengzhou, China</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Tong Liu (tong2.liu@polyu.edu.hk)</corresp></author-notes><pub-date><day>31</day><month>August</month><year>2026</year></pub-date>
      
      <volume>26</volume>
      <issue>17</issue>
      <fpage>12243</fpage><lpage>12259</lpage>
      <history>
        <date date-type="received"><day>28</day><month>February</month><year>2026</year></date>
           <date date-type="rev-request"><day>8</day><month>May</month><year>2026</year></date>
           <date date-type="rev-recd"><day>14</day><month>July</month><year>2026</year></date>
           <date date-type="accepted"><day>4</day><month>August</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Yongchao Ma et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026.html">This article is available from https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026.html</self-uri><self-uri xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026.pdf">The full text article is available as a PDF file from https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e144">Integrating satellite-derived precipitable water vapor (PWV) provides data with high spatiotemporal resolution, which is crucial for monitoring and forecasting extreme weather. However, current fusion and calibration methods typically rely on dense GNSS networks, hindering its application in data-sparse regions. It remains unclear whether improving calibration under sparse conditions depends more on increasing station numbers or optimizing their spatial placement. To address this, we developed a machine learning-based calibration framework for  all-weather PWV and conducted controlled experiments across China. Our key finding is that for a fixed station budget, a spatially random layout consistently outperforms clustered or geographically biased distributions, reducing RMSE by up to 27 % under the same validation set. While increasing station density improves spatial generalization, with RMSE at independent stations dropping from 3.24–2.28 <inline-formula><mml:math id="M1" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> and bias converging near zero, performance gains saturate beyond approximately 120–160 stations. Spatially, errors under sparse, non-uniform networks concentrate in regions with strong humidity gradients or complex terrain; a uniform layout distributes errors more evenly. Temporally, all calibrated models capture seasonal cycles, with residual errors peaking in summer due to convective activity. This study demonstrates that in sparse network design, maximizing spatial coverage uniformity is more critical than simply adding stations. Thus, a transferable framework and a quantitative principle are provided for generating reliable satellite PWV products where GNSS observations are limited.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Natural Science Basic Research Program of Shaanxi Province</funding-source>
<award-id>2026JC-YBQN-0397</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e166">Atmospheric water vapor is a key component of the climate system, and its accurate monitoring via precipitable water vapor (PWV) is essential for weather and climate studies (Rocken et al., 1997; Trenberth et al., 2005). High-accuracy, spatiotemporally continuous PWV fields are therefore critical for understanding hydrological processes and improving meteorological forecasts (Lu et al., 2016; Chen and Liu 2016).</p>
      <p id="d2e169">At present, PWV observations are mainly derived from ground-based measurements and satellite remote sensing. Ground-based techniques, including radiosondes, microwave radiometers, and Global Navigation Satellite Systems (GNSS), generally provide high accuracy and high temporal resolution, making them suitable for capturing rapid variations in atmospheric water vapor (Ware et al., 2000; Durre et al., 2006; Namaoui et al., 2017). However, their spatial representativeness is strongly constrained by station density and distribution, limiting their ability to provide spatially continuous PWV fields over large regions. In contrast, satellite remote sensing provides regional to global coverage, complementing the spatial limitations of ground networks (Kaufman and Gao 1992; Zhang et al., 2019; Zhao et al., 2024). Nevertheless, near-infrared and thermal infrared PWV products are often affected by cloud contamination, viewing geometry, and retrieval assumptions, leading to insufficient accuracy, missing data, and poor spatiotemporal continuity. These issues are particularly pronounced over complex underlying surfaces or in regions with strong water vapor gradients (Jiang et al., 2024), thereby limiting the applicability of satellite-derived PWV in high-resolution moisture monitoring and process-oriented studies.</p>
      <p id="d2e172">To improve the quality of satellite PWV products, existing studies have generally followed two main technical pathways. One approach focuses on post-processing calibration of satellite PWV using ground-based observations to reduce systematic biases and enhance consistency across different observing systems. The other approach aims to improve retrieval algorithms and parameterization schemes at the source, thereby strengthening physical consistency and reducing retrieval uncertainty (Merrikhpour and Rahimzadegan 2017; He and Liu 2020).</p>
      <p id="d2e175">In post-calibration studies, GNSS-derived PWV has been widely adopted as a reference “truth” for satellite PWV calibration. Investigations based on relatively large GNSS networks consistently demonstrate that multi-station constraints can substantially improve the accuracy and stability of satellite PWV products. For example, Bai et al. (2021) constructed a linear MODIS PWV calibration model using 260 GNSS stations over China, achieving an overall accuracy improvement of approximately 20 %. Subsequently, Researchers developed machine-learning-based calibration models for MODIS and FY-3A PWV using 2794 and 214 GNSS stations over China, respectively, significantly enhancing the long-term performance of satellite PWV under all-sky conditions (Qin et al., 2023; Xu and Liu, 2023a). Xu and Liu (2023b) further validated the effectiveness of machine learning approaches in improving all-weather satellite PWV quality using 453 ground stations in Australia. To address the limited spatiotemporal continuity of PWV maps, Ma et al. (2023) proposed a distributed ensemble framework integrating multi-source heterogeneous data based on 207 GNSS stations in New South Wales. Ma et al. (2022a) reported that when only 14 GNSS stations in the Tibetan Plateau were used to construct a calibration model, the sparsity of stations prevented machine learning approaches from achieving high calibration accuracy. The consistent conclusions drawn from these studies across different regions and satellite products indicate that the number of GNSS stations is a key prerequisite influencing satellite PWV calibration performance. However, most existing studies treat the available station network as a given condition, and whether increasing the number of stations necessarily leads to linear or unlimited performance gains remains insufficiently understood.</p>
      <p id="d2e179">Meanwhile, dense GNSS PWV observations have also been used to develop multi-parameter satellite PWV mapping models. Researchers incorporated land-cover information using 173 GNSS stations in North America and achieved more than a 66 % improvement in MODIS near-infrared PWV retrieval accuracy through multi-channel coupling (Ma et al., 2022c). In addition, progress has been made in high-accuracy PWV retrieval from Fengyun satellites (Zhao et al., 2024), as well as in all-weather PWV retrieval through the fusion of infrared, near-infrared, and microwave observations (Sun et al., 2024; Du et al., 2025). Despite these advances, station number and spatial configuration are typically regarded as background conditions rather than explicitly treated as independent variables governing model performance.</p>
      <p id="d2e182">In summary, existing satellite PWV reconstruction and calibration studies primarily aim to establish mapping relationships between satellite observational information – including products, spectral features, spatiotemporal attributes, and environmental variables – and high-accuracy ground-based PWV, particularly GNSS-derived PWV. However, most studies have concentrated on regions with relatively dense GNSS station coverage, while quantitative understanding of model performance, stability, and generalization behavior under sparse station conditions is still lacking. How the number of training stations and their spatial configuration jointly constrain the spatial consistency and extrapolation capability of calibration models has not been systematically evaluated. To address this gap, we develop a machine learning-based calibration framework for the all-weather  PWV product over China. Through controlled experiments, we systematically isolate and quantify the impacts of training station density and spatial layout on model accuracy, spatial generalization, and temporal stability. The objective is to provide quantitative guidance on station configuration and methodological support for satellite PWV reconstruction under sparse observational conditions.</p>
      <p id="d2e187">The remainder of this paper is organized as follows. Section 2 describes the study area, datasets, and preprocessing procedures. Section 3 presents the all-weather PWV model construction and evaluation metrics. Section 4 reports the experimental results and discussion. Finally, Sect. 5 summarizes the main conclusions and outlines future research directions.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Study Area and Data</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Study Area</title>
      <p id="d2e205">This study focuses on mainland China as the study region (18–53° N, 73–135° E; Fig. 1). The region exhibits pronounced topographic heterogeneity characterized by a stepwise terrain pattern, including the Tibetan Plateau, inland basins, and eastern plains. Such complex terrain leads to strong regional contrasts in climatic background and water vapor transport and convergence mechanisms, resulting in pronounced spatial heterogeneity in PWV.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e210">Spatial Distribution of GNSS and IGRA Stations in China.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f01.png"/>

        </fig>

      <p id="d2e219">To systematically evaluate the influence of training sample size and spatial configuration on the generalization capability of all-weather PWV models, a total of 244 GNSS stations were selected as modeling samples. In addition, 16 GNSS stations were randomly excluded from the training process and reserved as an independent validation set. Furthermore, 80 radiosonde stations from the Integrated Global Radiosonde Archive (IGRA) were selected as an external independent validation dataset. The spatial distribution of all stations is shown in Fig. 1, where Fig. 1b highlights the randomly selected 16 independent GNSS validation stations.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Data and Processing</title>
<sec id="Ch1.S2.SS2.SSS1">
  <label>2.2.1</label><title>FY-4A</title>
      <p id="d2e237">The Fengyun-4A () satellite is China's second-generation geostationary meteorological satellite, operating in geostationary orbit and carrying the Advanced Geostationary Radiation Imager (AGRI). AGRI provides high-temporal-resolution observations; in addition to full-disk scanning, it performs regional observations over China at a 15 <inline-formula><mml:math id="M2" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">min</mml:mi></mml:mrow></mml:math></inline-formula> interval.  PWV products are freely available in near real time from the official data portal of National Satellite Meteorological Center (<uri>http://satellite.nsmc.org.cn/</uri>, last access: 22 August 2026). Because near-infrared channels cannot penetrate thick cloud cover, cloud detection is required prior to PWV retrieval. In this study, cloud-contaminated PWV observations were identified and excluded using the MERSI cloud mask product.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <label>2.2.2</label><title>GNSS</title>
      <p id="d2e263">GNSS-derived PWV data were obtained from the China Crustal Movement Observation Network (CMONOC), which provides continuous, high-precision, and high-temporal-resolution observations across mainland China. CMONOC has been widely used for monitoring crustal deformation, gravity field variations, tropospheric water vapor, and ionospheric electron content (Li et al., 2012). In this study, CMONOC observations for 2023 were processed using the GAMIT/GLOBK software to generate hourly zenith total delay (ZTD) time series. Zenith hydrostatic delay (ZHD) and zenith wet delay (ZWD) were separated, and PWV was derived using standard conversion factors. The basic relationship can be expressed as

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M3" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E1"><mml:mtd><mml:mtext>1</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>ZWD</mml:mtext><mml:mo>=</mml:mo><mml:mtext>ZTD</mml:mtext><mml:mo>-</mml:mo><mml:mtext>ZHD</mml:mtext></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>PWV</mml:mtext><mml:mo>=</mml:mo><mml:mtext>ZWD</mml:mtext><mml:mo>⋅</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">6</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi mathvariant="normal">w</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi mathvariant="normal">w</mml:mi></mml:msub><mml:mo>⋅</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">3</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:msub><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>T</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∫</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mi>e</mml:mi><mml:mi>T</mml:mi></mml:mfrac></mml:mstyle><mml:mtext>ds</mml:mtext></mml:mrow><mml:mrow><mml:mo>∫</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mi>e</mml:mi><mml:mrow><mml:msup><mml:mi>T</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mtext>ds</mml:mtext></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            where <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mn mathvariant="normal">3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">377</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">600.0</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M5" display="inline"><mml:mrow class="unit"><mml:msup><mml:mi mathvariant="normal">K</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">Pa</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:msub><mml:msup><mml:mi/><mml:mo>′</mml:mo></mml:msup><mml:mn mathvariant="normal">2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">16.52</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M7" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">hPa</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is weighted mean temperature (<inline-formula><mml:math id="M9" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula>), <inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">ρ</mml:mi><mml:mi mathvariant="normal">w</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the density of liquid water, <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi mathvariant="normal">w</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the specific gas constant for water vapor (461.5 <inline-formula><mml:math id="M12" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">J</mml:mi><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">kg</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:msup><mml:mi mathvariant="normal">K</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>), <inline-formula><mml:math id="M13" display="inline"><mml:mi>e</mml:mi></mml:math></inline-formula> is water vapor pressure (<inline-formula><mml:math id="M14" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">hPa</mml:mi></mml:mrow></mml:math></inline-formula>), <inline-formula><mml:math id="M15" display="inline"><mml:mi>T</mml:mi></mml:math></inline-formula> is temperature (<inline-formula><mml:math id="M16" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">K</mml:mi></mml:mrow></mml:math></inline-formula>), and ds denotes the vertical layer thickness. Surface pressure and weighted mean temperature required for GNSS PWV estimation were provided by ERA5.</p>
      <p id="d2e569">The ZHD was computed using the Saastamoinen model (Saastamoinen, 1972), which achieves sub-millimeter accuracy under standard atmospheric conditions. The derived GNSS PWV was then used as a reference truth dataset for all-weather PWV model training and validation.

              <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M17" display="block"><mml:mrow><mml:mtext>ZHD</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">0.0022768</mml:mn><mml:mo>⋅</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.00266</mml:mn><mml:mo>⋅</mml:mo><mml:mi>cos⁡</mml:mi><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>⋅</mml:mo><mml:mi mathvariant="italic">φ</mml:mi></mml:mrow></mml:mfenced><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.00028</mml:mn><mml:mo>⋅</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the pressure at the GNSS station; <inline-formula><mml:math id="M19" display="inline"><mml:mi mathvariant="italic">φ</mml:mi></mml:math></inline-formula> is the latitude of the station; <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the elevation of the station.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS3">
  <label>2.2.3</label><title>Radiosonde</title>
      <p id="d2e661">Radiosonde-derived PWV data were obtained from the Integrated Global Radiosonde Archive (IGRA), which is freely accessible at <uri>https://www.ncei.noaa.gov/products/weather-balloon/integrated-global-radiosonde-archive</uri> (last access: 22 August 2026). PWV derived from 2023 radiosonde observations was used as an external independent validation dataset to assess model consistency across different observation systems. To ensure temporal representativeness and continuity, stations with severe data gaps were excluded. Only stations with valid observations for no less than one third of the year (i.e., at least 120 <inline-formula><mml:math id="M21" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">d</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">yr</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>) were retained. After quality control, PWV data from 80 radiosonde stations in 2023 (Fig. 1) were selected for final model evaluation.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS4">
  <label>2.2.4</label><title>Topographic and Land Cover</title>
      <p id="d2e692">Digital elevation model (DEM) data from the Shuttle Radar Topography Mission (SRTM) were used as the elevation reference for  PWV. The SRTM DEM was jointly produced by NASA and the National Geospatial-Intelligence Agency (NGA) and is available at <uri>http://www.resdc.cn/</uri> (last access: 22 August 2026). In this study, SRTM DEM data with a spatial resolution of 90 <inline-formula><mml:math id="M22" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:math></inline-formula> were resampled to 4 <inline-formula><mml:math id="M23" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> to match the spatial resolution of the reconstructed  PWV. Land cover information was obtained from China's first annual land cover dataset derived from Landsat imagery. The dataset includes nine land cover types: cropland, forest, shrubland, grassland, water bodies, snow/ice, bare land, impervious surfaces, and wetlands. Considering the 4 <inline-formula><mml:math id="M24" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> spatial footprint of  PWV, a circular buffer with a radius of 2 <inline-formula><mml:math id="M25" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> was constructed around each GNSS station. The area fractions of different land cover types within each buffer were calculated and used as model input variables to enhance the representation of surface heterogeneity and regionally non-uniform errors.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS5">
  <label>2.2.5</label><title>NDVI</title>
      <p id="d2e746">Vegetation conditions and evapotranspiration processes play an important role in near-surface moisture sources and recycling. Therefore, the normalized difference vegetation index (NDVI) serves as an effective indicator for characterizing land-surface contributions to atmospheric water vapor and regional variability (Ma et al., 2022b). NDVI data were obtained from the MODIS vegetation index product MOD13A2, which is generated from atmospherically corrected daily bidirectional surface reflectance and provides NDVI at a spatial resolution of 1 <inline-formula><mml:math id="M26" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> and a temporal resolution of 16 <inline-formula><mml:math id="M27" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:math></inline-formula> (<uri>https://ladsweb.modaps.eosdis.nasa.gov/</uri>, last access: 22 August 2026). Similarly, a resampling method is employed to standardize the spatial resolution of NDVI and  PWV. Given its temporal characteristics, a fixed 16 <inline-formula><mml:math id="M28" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:math></inline-formula> cycle was adopted, and a constant NDVI value was assigned within each cycle to ensure data consistency and avoid additional uncertainty introduced by excessive temporal interpolation.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Methods</title>
      <p id="d2e788">To reduce the impact of cloud contamination and missing observations in  PWV products on statistical analysis and spatial comparison, an all-weather  PWV reconstruction was first performed to generate spatiotemporally continuous PWV fields. Subsequently, GNSS-derived PWV was introduced as an external constraint to calibrate reconstructed  PWV and reduce systematic biases. The stability and generalization of the all-weather PWV model were then evaluated under different training station numbers and spatial configurations. The overall workflow is illustrated in Fig. 2.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e799">Flowchart of the all-weather PWV map reconstruction and calibration.</p></caption>
        <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f02.png"/>

      </fig>

<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>All-weather FY-4A PWV Reconstruction</title>
      <p id="d2e815">The reconstruction aims to restore PWV continuity under cloudy or missing-data conditions by establishing a nonlinear mapping between  PWV and multiple auxiliary variables, including meteorological parameters, topography, time, and spatial information. To ensure spatiotemporal consistency across all input data, a rigorous pre-processing procedure was applied. Temporally, the ERA5 data were aligned to match the exact observation hours of the full-disk  PWV product. Spatially, the  PWV data and the DEM over the study domain were resampled onto the ERA5's <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.25</mml:mn><mml:mi mathvariant="italic">°</mml:mi><mml:mo>×</mml:mo><mml:mn mathvariant="normal">0.25</mml:mn><mml:mi mathvariant="italic">°</mml:mi></mml:mrow></mml:math></inline-formula> grid using a bilinear interpolation method, thereby unifying the spatial resolution and grid alignment for all datasets prior to model training.</p>
      <p id="d2e840">A random forest (RF) model was employed to describe the nonlinear relationship between PWV and auxiliary variables:

            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M30" display="block"><mml:mrow><mml:msub><mml:mtext>PWV</mml:mtext><mml:mtext>FY-4A</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mtext>RF</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mtext>Lat,Lon,Dem,Time,T2m,TCW,SP</mml:mtext><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          where T2m, SP, and PWV denote near-surface temperature, surface pressure, and water vapor-related parameters, respectively; Dem represents elevation; Lat and Lon denote spatial location; and TCW is total column water. After training, the model was applied to generate spatiotemporally continuous reconstructed  PWV fields.</p>
      <p id="d2e869">The model performance was evaluated using the test set, and the results show that the reconstructed model achieved an RMSE of 1.12 <inline-formula><mml:math id="M31" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> and a Bias of 0.51 <inline-formula><mml:math id="M32" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>. The reconstructed PWV has been verified to effectively retrieve the spatial distribution characteristics of water vapor over the entire region, and its distribution is consistent with that of ERA5 PWV. Detailed validation of the reconstruction has been reported in previous studies (Wang et al., 2026) and is not repeated here.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>GNSS-constrained FY-4A PWV Calibration</title>
      <p id="d2e897">Although reconstruction improves data completeness, regional systematic biases remain in  PWV. Therefore, GNSS PWV was used as a reference to calibrate reconstructed  PWV. To address the spatial mismatch between site-based GNSS observational data and gridded satellite PWV, this study employs a bilinear interpolation method to interpolate  gridded data to the GNSS site locations. Temporal mismatches were handled by averaging GNSS PWV at the two nearest epochs before and after satellite overpass times.</p>
      <p id="d2e906">An RF-based calibration model was constructed to map reconstructed  PWV, spatiotemporal information, NDVI, and land cover characteristics to GNSS PWV:

            <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M33" display="block"><mml:mrow><mml:msub><mml:mtext mathvariant="normal">PWV</mml:mtext><mml:mtext>GNSS</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mtext>RF</mml:mtext></mml:msub><mml:mo>(</mml:mo><mml:mtext>Lat,Lon,Elv,Time,NDVI,</mml:mtext><mml:msub><mml:mi>P</mml:mi><mml:mtext>LCTi</mml:mtext></mml:msub><mml:msub><mml:mtext>,PWV</mml:mtext><mml:mtext>FY-4A</mml:mtext></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mtext>LCTi</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> denotes the area fraction of the <inline-formula><mml:math id="M35" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th land cover type within a 2 <inline-formula><mml:math id="M36" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula> buffer around each GNSS station.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Station Configuration Experiments</title>
      <p id="d2e981">To quantitatively assess the effects of training station density and spatial structure on model generalization, two groups of comparative experiments were designed, as shown in Fig. 2b. <list list-type="custom"><list-item><label>1.</label>
      <p id="d2e986">Training station number experiment. Different numbers of GNSS stations were randomly selected from the available 228 stations to construct training datasets. Seven training sample sizes were considered: No. <inline-formula><mml:math id="M37" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 30, 40, 80, 120, 160, 200, and 244. Among them, Experiment No. 244 served as the reference. The remaining GNSS stations were used as independent validation samples, and IGRA radiosonde PWV data were further introduced for external validation.</p></list-item><list-item><label>2.</label>
      <p id="d2e997">Spatial structure experiment. Under the same training station number, three typical spatial configurations were constructed, including random, surrounding, and clustered distributions. These configurations were designed to evaluate how spatial coverage uniformity and representativeness influence spatial error patterns and extrapolation capability. Specific methods of construction are as follows. <list list-type="bullet"><list-item>
      <p id="d2e1002">Random distribution. A network composed of 200 stations is selected completely at random from all stations. This method simulates a scenario where station locations have no specific spatial pattern.</p></list-item><list-item>
      <p id="d2e1006">Surrounding distribution. To simulate the pattern of stations  distributed around a central area, the geographic center of all  stations is first identified. Then, 28 stations are randomly removed from the central area. The remaining 200 stations, primarily located in peripheral regions, constitute the training set.</p></list-item><list-item>
      <p id="d2e1010">Clustered distribution. To simulate the pattern of stations densely concentrated in a certain area, the opposite strategy is adopted: 28 stations are randomly removed from the peripheral areas of the station network, retaining 200 stations in the central and  relatively concentrated regions.</p></list-item></list></p></list-item></list></p>
      <p id="d2e1013">Through this experimental design, the relationships among sample size, spatial coverage, and error response were systematically examined, enabling the identification of major limiting factors governing model performance improvement and the corresponding saturation thresholds.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Model Parameter Optimization</title>
      <p id="d2e1024">Key RF parameters were optimized using out-of-bag (OOB) error as the evaluation criterion. Bayesian optimization was applied to automatically search the parameter space. This process was repeated for more than 5 cycles to minimize random variability in the optimal results. The range of tree numbers was 50–250 with an interval of 10. The final optimal parameter settings are summarized in Table 1.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e1030">Random forest hyperparameter values based on different station configurations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center" colsep="1"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col2" align="center" colsep="1">Training station number experiment </oasis:entry>
         <oasis:entry namest="col3" nameend="col4" align="center">Spatial structure experiment </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Number of Station/</oasis:entry>
         <oasis:entry colname="col2">Number of</oasis:entry>
         <oasis:entry colname="col3">Number of Station/</oasis:entry>
         <oasis:entry colname="col4">Number of</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Distribution Type</oasis:entry>
         <oasis:entry colname="col2">Decision Trees</oasis:entry>
         <oasis:entry colname="col3">Distribution Type</oasis:entry>
         <oasis:entry colname="col4">Decision Trees</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">30/Random</oasis:entry>
         <oasis:entry colname="col2">150</oasis:entry>
         <oasis:entry colname="col3">200/Surrounding</oasis:entry>
         <oasis:entry colname="col4">200</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">40/Random</oasis:entry>
         <oasis:entry colname="col2">150</oasis:entry>
         <oasis:entry colname="col3">200/Clustered</oasis:entry>
         <oasis:entry colname="col4">150</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">80/Random</oasis:entry>
         <oasis:entry colname="col2">150</oasis:entry>
         <oasis:entry colname="col3">200/ Random</oasis:entry>
         <oasis:entry colname="col4">200</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">120/Random</oasis:entry>
         <oasis:entry colname="col2">150</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">160/Random</oasis:entry>
         <oasis:entry colname="col2">200</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">200/Random</oasis:entry>
         <oasis:entry colname="col2">200</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">244/Random</oasis:entry>
         <oasis:entry colname="col2">200</oasis:entry>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>Evaluation Metrics and Validation Strategy</title>
      <p id="d2e1198">Model performance was evaluated at three levels, including training GNSS stations, independent GNSS validation stations, and IGRA radiosonde stations, to distinguish fitting capability from spatial extrapolation performance. The coefficient of determination, RMSE, and mean bias were adopted as the primary evaluation metrics:

                <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M38" display="block"><mml:mtable rowspacing="5.690551pt 5.690551pt" displaystyle="true"><mml:mlabeledtr id="Ch1.E7"><mml:mtd><mml:mtext>7</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>Bias</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>m</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:msubsup><mml:mfenced close=")" open="("><mml:mrow><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>model</mml:mtext></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>ref</mml:mtext></mml:msup></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E8"><mml:mtd><mml:mtext>8</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtext>RMSE</mml:mtext><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>m</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:msubsup><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>model</mml:mtext></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>ref</mml:mtext></mml:msup></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E9"><mml:mtd><mml:mtext>9</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>∑</mml:mo><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>ref</mml:mtext></mml:msup><mml:mo>-</mml:mo><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>model</mml:mtext></mml:msup></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>∑</mml:mo><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>ref</mml:mtext></mml:msup><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mrow><mml:msup><mml:mtext>Data</mml:mtext><mml:mtext>model</mml:mtext></mml:msup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

          where Data<sup>model</sup> denotes model-derived PWV and Data<sup>ref</sup> represents reference PWV observations.</p>
      <p id="d2e1380">In addition to overall statistical metrics, model robustness was further assessed in terms of spatial error patterns, PWV time series at representative stations, and monthly aggregated statistics. These analyses were conducted to evaluate model stability, seasonal dependence, and spatial consistency across different observational networks.</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results and Discussion</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Overall Performance</title>
      <p id="d2e1400">A central finding of our experiments is that in-sample fitting accuracy is remarkably insensitive to the number of training stations, whereas spatial generalization (validation at independent sites) critically depends on it. As shown in Fig. 3, for training stations, the accuracy of all-weather PWV model remains consistently high as the number of training stations increases from 30–244. The coefficient of determination <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> remains stable within the range of 0.988–0.990, while RMSE varies only slightly between 1.45 and 1.66 <inline-formula><mml:math id="M42" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>. Mean bias is close to zero throughout all configurations (<inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mo>|</mml:mo><mml:mtext>Bias</mml:mtext><mml:mo>|</mml:mo><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">0.02</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M44" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>). These results suggest that the random forest model is capable of sufficiently learning the regional-scale mapping relationship between GNSS PWV and  PWV, and that further increases in training samples yield only marginal improvements in fitting accuracy.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e1450">Validation results of models trained with varying numbers of GNSS stations. Panels <bold>(a)</bold>–<bold>(g)</bold> show the validation accuracy evaluated using stations included in the training set, whereas panels <bold>(a1)</bold>–<bold>(g1)</bold> depict the validation accuracy assessed using stations withheld from the training process.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f03.png"/>

        </fig>

      <p id="d2e1471">In contrast, the number of training stations has a pronounced effect on model performance at independent GNSS validation stations. As training station numbers increase, retrieval accuracy at independent stations improves steadily, with <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> increasing from approximately 0.96–0.97 and RMSE decreasing markedly from 3.24–2.28 <inline-formula><mml:math id="M46" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>. Meanwhile, Bias gradually converges from 0.42 <inline-formula><mml:math id="M47" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> toward near-zero values. These results demonstrate that enlarging the training dataset enhances the model's ability to represent regional PWV heterogeneity and effectively suppresses systematic errors during spatial extrapolation. Notably, when the number of training stations exceeds approximately 160, improvements in independent-station accuracy become marginal, indicating a saturation behavior in model performance and suggesting that the dominant statistical characteristics of PWV over the study region have been adequately captured.</p>
      <p id="d2e1502">This saturation identifies a cost-effective range for station numbers. Beyond this range, merely adding stations yields diminishing returns. To dissect the drivers of this saturation and explore how to optimize performance within a fixed station budget, we now turn to a spatial analysis of errors.</p>
      <p id="d2e1505">Figure 4 presents the inversion performance of models trained with different numbers of stations at 16 GNSS stations that were not involved in the training process. Uncalibrated  PWV exhibits substantial systematic underestimation (Bias: <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">2.17</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M49" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>) accompanied by large random errors (RMSE: 3.93 <inline-formula><mml:math id="M50" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>). With increasing training station numbers, calibration errors are significantly reduced and spatial performance becomes more stable, with particularly pronounced improvements observed in low-latitude regions. When the training station number reaches 200, the mean Bias decreases to 0.07 <inline-formula><mml:math id="M51" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>, while RMSE shows a monotonic decline from 3.93–1.90 <inline-formula><mml:math id="M52" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e1553">The consistency of these improvements across independent stations indicates that increasing training sample size effectively mitigates systematic bias and enhances model generalization. From a mechanistic perspective, when training data are limited, the model cannot adequately represent the spatial heterogeneity of regional PWV, leading to error accumulation and amplified random errors in regions with distinct climatic backgrounds, surface conditions, and moisture variability. As training station numbers increase, the diversity of PWV scenarios represented in the training data expands, enabling the model to learn more representative regional statistical features and thereby improving both accuracy and stability at independent locations.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e1558">Validation results of different models using GNSS stations not involved in model training.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f04.png"/>

        </fig>

      <p id="d2e1567">Independent validation using IGRA radiosonde PWV is summarized in Table 2. When the number of training stations increases from 30–120, model errors decrease substantially, with RMSE reduced from 2.68–2.36 <inline-formula><mml:math id="M53" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> and Bias converging to near-zero values (0.02 <inline-formula><mml:math id="M54" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>). This indicates that a moderate training sample size is enough to achieve a favorable balance between accuracy and robustness. Further increasing the number of training stations to 160–244 results in only minor fluctuations in RMSE (approximately 2.37–2.39 <inline-formula><mml:math id="M55" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>) and Bias, confirming that model performance gradually approaches saturation. Moreover, it was found that the model trained on 120 stations achieved higher accuracy than those trained on a larger number of stations. This phenomenon can be attributed to two main reasons. First, as the number of training stations increases, the model is compelled to learn more generalizable physical patterns, which may slightly reduce its fitting accuracy to the training data (i.e., a shift from “overfitting” toward “generalization”). Second, redundant stations fail to provide additional useful information; instead, they introduce a smoothing effect that dilutes the critical spatial features present in the original data.</p>

<table-wrap id="T2"><label>Table 2</label><caption><p id="d2e1598">Validation results of different models using IGRA stations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry colname="col2">Bias [<inline-formula><mml:math id="M56" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
         <oasis:entry colname="col3">RMSE [<inline-formula><mml:math id="M57" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">RF_30</oasis:entry>
         <oasis:entry colname="col2">0.29</oasis:entry>
         <oasis:entry colname="col3">2.68</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_40</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.86</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">2.72</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_80</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.47</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">2.46</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_120</oasis:entry>
         <oasis:entry colname="col2">0.02</oasis:entry>
         <oasis:entry colname="col3">2.36</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_160</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.14</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">2.38</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_200</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.39</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">2.39</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_244</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.14</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">2.37</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e1771">Overall, the consistent patterns revealed by Figs. 3 and 4 and Table 2 demonstrate that increasing training station density is essential for improving the spatial generalization capability of GNSS PWV calibration models. However, larger sample sizes do not necessarily translate into proportional performance gains. From the perspective of station number and random spatial distribution alone, a training station scale of approximately 120–200 provides an effective compromise between accuracy improvement and computational efficiency for regional-scale PWV calibration over China.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Spatial Dependence of Model Performance</title>
      <p id="d2e1782">While the analysis above confirms the role of station number and its saturation, a more pressing practical question emerges under resource constraints: Can the spatial layout of a fixed number of stations serve as a lever to enhance performance beyond what the number alone dictates? To answer this, we analyze the spatial patterns of model errors.</p>
      <p id="d2e1785">Figure 5 presents the spatial distributions of RMSE and Bias over China based on IGRA radiosonde validation under different training station numbers, together with the spatial distribution of training GNSS stations. Overall, both RMSE and Bias exhibit pronounced spatial heterogeneity, reflecting the complex spatial variability of atmospheric water vapor and its controlling factors.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e1791">Spatial distribution of validation accuracy for different models using IGRA stations.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f05.png"/>

        </fig>

      <p id="d2e1801">When the number of training stations is limited (e.g., No. <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">30</mml:mn></mml:mrow></mml:math></inline-formula> and No. <inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">40</mml:mn></mml:mrow></mml:math></inline-formula>), spatial discrepancies in model errors are particularly pronounced. RMSE is generally higher in regions with more active moisture variability, especially in eastern and southern China, where RMSE at some stations exceeds 3.0 <inline-formula><mml:math id="M65" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>. At the same time, Bias shows a clear pattern of alternating positive and negative values, indicating strong regional dependence during spatial extrapolation. These results suggest that insufficient training samples with limited spatial coverage hinder the model's ability to capture regional-scale PWV heterogeneity, leading to increased systematic errors in specific regions.</p>
      <p id="d2e1832">As the number of training stations increases to 80 and 120, the spatial error structure improves substantially. Overall RMSE levels decrease, the extent of high-error regions shrinks markedly, and RMSE at most stations concentrates within approximately 2.0–2.5 <inline-formula><mml:math id="M66" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>. Meanwhile, the spatial distribution of Bias becomes smoother, with extreme positive and negative deviations significantly reduced. This improvement indicates a substantially enhanced capacity of the model to adapt to regional differences in water vapor variability. When the number of training stations further increases to 160, the improvement in RMSE and Bias becomes relatively limited, suggesting that the dominant spatial variability patterns of PWV have largely been captured under the current data and regional setting, and that further increases in training samples yield diminishing returns in terms of spatial error reduction.</p>
      <p id="d2e1843">Overall, these spatial patterns confirm that small training datasets with insufficient spatial coverage are unable to represent regional PWV heterogeneity, resulting in localized systematic biases. In contrast, increasing training station numbers enhances the model's adaptability to diverse climatic backgrounds and moisture regimes, thereby effectively reducing spatially non-uniform errors. Once the number of training stations reaches approximately 120–160, model performance gradually approaches saturation, consistent with the statistical results discussed previously. This consistency further emphasizes the importance of jointly considering training sample size and spatial coverage in regional-scale PWV calibration.</p>
      <p id="d2e1846">To further examine the role of spatial density, Table 3 summarizes the relationship between calibration accuracy and station spatial resolution, quantified by the mean inter-station distance derived from a Delaunay triangulation. RMSE shows an overall positive correlation with mean station spacing, indicating that denser station networks provide stronger spatial constraints on the model. When the number of training stations exceeds approximately 120, improvements in accuracy become notably weaker, again exhibiting a diminishing marginal benefit. This behavior can be attributed to the effective reduction in inter-station distance with increasing station numbers, which enhances the spatial representativeness of training data and improves the model's ability to characterize PWV spatial variability. These results demonstrate that, beyond sample size, the spatial configuration of training stations is a critical factor limiting further improvements in calibration performance.</p>

<table-wrap id="T3" specific-use="star"><label>Table 3</label><caption><p id="d2e1852">Comparison of GNSS spatial resolution and model accuracy using IGRA station.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="center"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="center" colsep="1"/>
     <oasis:colspec colnum="5" colname="col5" align="center"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="center"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Model</oasis:entry>
         <oasis:entry namest="col2" nameend="col4" colsep="1">Outside the black box </oasis:entry>
         <oasis:entry namest="col5" nameend="col7">Inside the black box </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">Average Distance [<inline-formula><mml:math id="M67" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
         <oasis:entry colname="col3">Bias [<inline-formula><mml:math id="M68" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
         <oasis:entry colname="col4">RMSE [<inline-formula><mml:math id="M69" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
         <oasis:entry colname="col5">Average Distance [<inline-formula><mml:math id="M70" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">km</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
         <oasis:entry colname="col6">Bias [<inline-formula><mml:math id="M71" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
         <oasis:entry colname="col7">RMSE [<inline-formula><mml:math id="M72" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_N30</oasis:entry>
         <oasis:entry colname="col2">821.81</oasis:entry>
         <oasis:entry colname="col3">0.52</oasis:entry>
         <oasis:entry colname="col4">2.30</oasis:entry>
         <oasis:entry colname="col5">626.84</oasis:entry>
         <oasis:entry colname="col6">0.01</oasis:entry>
         <oasis:entry colname="col7">3.16</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_N40</oasis:entry>
         <oasis:entry colname="col2">646.35</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.52</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">2.12</oasis:entry>
         <oasis:entry colname="col5">784.79</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1.28</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">3.45</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_N80</oasis:entry>
         <oasis:entry colname="col2">481.59</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.32</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">1.95</oasis:entry>
         <oasis:entry colname="col5">473.30</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.66</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">3.08</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_N120</oasis:entry>
         <oasis:entry colname="col2">441.37</oasis:entry>
         <oasis:entry colname="col3">0.08</oasis:entry>
         <oasis:entry colname="col4">1.90</oasis:entry>
         <oasis:entry colname="col5">381.58</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">2.91</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">RF_N160</oasis:entry>
         <oasis:entry colname="col2">369.75</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.10</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">2.00</oasis:entry>
         <oasis:entry colname="col5">336.01</oasis:entry>
         <oasis:entry colname="col6"><inline-formula><mml:math id="M79" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.18</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col7">2.86</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e2148">Due to the combined influence of land–sea contrast, temperature fields, and large-scale circulation, southern China (black-box region) exhibits stronger spatiotemporal variability in water vapor and more pronounced short-term fluctuations than northern China. Under limited station coverage, the model struggles to capture fine-scale variability in this region, resulting in generally lower accuracy compared with northern China. For example, for the RF model, increasing the training configuration from RF_N30 to RF_N160 reduces mean station spacing by 46.40 % in southern China and leads to a 13.04 % reduction in RMSE, whereas in northern China, mean spacing is reduced by approximately 55 % and RMSE decreases by 17.10 %. These comparisons further demonstrate that reducing inter-station distance and improving spatial coverage density can significantly enhance calibration accuracy, although the magnitude of improvement is constrained in regions with more intense moisture variability.</p>
      <p id="d2e2151">Figure 6 further compares the spatial distribution of  PWV retrieved by different models under varying training station numbers. A comparison of Fig. 6a–c indicates that the all-weather model can effectively fill in regional missing values. At large spatial scales, the RF_N244 retrieval exhibits strong consistency with the GNSS-derived PWV field, characterized by a gradual increase from northwestern to southeastern China. To better illustrate the influence of training sample size, RF_N244 is used as a reference benchmark, and spatial differences between other models and RF_N244 are analyzed, as shown Fig. 6d–i. When training station numbers are small (e.g., RF_N30 and RF_N40), deviations relative to RF_N244 are large and spatially heterogeneous, with local differences exceeding 6 <inline-formula><mml:math id="M80" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula> in some regions. These discrepancies are particularly pronounced in areas with complex terrain or strong moisture gradients, indicating elevated uncertainty in representing local PWV structures under sparse training conditions.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e2166">Comparison of retrieval results at 12:00 UTC on the 232nd day of 2023 based on training models with different station configurations. Panels <bold>(a)–(c)</bold> respectively display  PWV, GNSS PWV, and RF_N244 model derived PWV. Panels <bold>(d)–(i)</bold> show the differences in retrieval results between each model and the RF_N244 model.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f06.png"/>

        </fig>

      <p id="d2e2183">As training station numbers increase, spatial differences relative to RF_N244 decrease markedly, with patterns transitioning from scattered to more continuous and smoother distributions. Extreme difference regions shrink substantially, and when training station numbers reach 160–200, only minor small-scale fluctuations remain. Notably, although station numbers increase beyond 120, the overall spatial coverage pattern does not improve substantially, reinforcing the conclusion that station layout and coverage uniformity are as important as sample size in constraining model performance.</p>
      <p id="d2e2186">The persistent spatial error patterns under sparse conditions suggest that the spatial representativeness and coverage uniformity of stations may be a more critical performance controller than sheer quantity. To test this hypothesis directly, we designed a controlled experiment isolating the effect of spatial configuration while holding the station number constant.</p>
      <p id="d2e2190">To explicitly isolate the effect of spatial configuration, calibration models were trained using GNSS stations arranged under three different spatial distribution structures, and their performance was independently evaluated using the 16 GNSS stations that were not involved in the model training process. The validation results are summarized in Table 4.</p>

<table-wrap id="T4" specific-use="star"><label>Table 4</label><caption><p id="d2e2196">Accuracy of models trained with different station distribution structures across the remaining GNSS stations.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="center"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row>

         <oasis:entry rowsep="1" colname="col1" morerows="1"/>

         <oasis:entry rowsep="1" namest="col2" nameend="col3" align="center" colsep="1">RF_Clustered </oasis:entry>

         <oasis:entry rowsep="1" namest="col4" nameend="col5" align="center" colsep="1">RF_Surrounding </oasis:entry>

         <oasis:entry rowsep="1" namest="col6" nameend="col7">RF_Random </oasis:entry>

       </oasis:row>
       <oasis:row rowsep="1">

         <oasis:entry colname="col2">RMSE [<inline-formula><mml:math id="M81" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>

         <oasis:entry colname="col3">Bias [<inline-formula><mml:math id="M82" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>

         <oasis:entry colname="col4">RMSE  [<inline-formula><mml:math id="M83" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>

         <oasis:entry colname="col5">Bias [<inline-formula><mml:math id="M84" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>

         <oasis:entry colname="col6">RMSE [<inline-formula><mml:math id="M85" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>

         <oasis:entry colname="col7">Bias [<inline-formula><mml:math id="M86" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>]</oasis:entry>

       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>

         <oasis:entry colname="col1">Area_Clustered</oasis:entry>

         <oasis:entry colname="col2">2.77</oasis:entry>

         <oasis:entry colname="col3"><inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.47</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>

         <oasis:entry colname="col4">–</oasis:entry>

         <oasis:entry colname="col5">–</oasis:entry>

         <oasis:entry colname="col6">2.02</oasis:entry>

         <oasis:entry colname="col7">0.16</oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col1">Area_Surrounding</oasis:entry>

         <oasis:entry colname="col2">–</oasis:entry>

         <oasis:entry colname="col3">–</oasis:entry>

         <oasis:entry colname="col4">3.61</oasis:entry>

         <oasis:entry colname="col5"><inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.51</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>

         <oasis:entry colname="col6">2.65</oasis:entry>

         <oasis:entry colname="col7"><inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.43</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>

       </oasis:row>
       <oasis:row>

         <oasis:entry colname="col1">Area_ Random</oasis:entry>

         <oasis:entry colname="col2">–</oasis:entry>

         <oasis:entry colname="col3">–</oasis:entry>

         <oasis:entry colname="col4">–</oasis:entry>

         <oasis:entry colname="col5">–</oasis:entry>

         <oasis:entry colname="col6">2.86</oasis:entry>

         <oasis:entry colname="col7">0.04</oasis:entry>

       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e2408">As shown in Table 4, the randomly distributed configuration consistently achieves the best calibration performance. Compared with the clustered and surrounding configurations, the random distribution reduces RMSE by 27.08 % and 26.59 %, respectively, and reduces Bias by 65.96 % and 15.69 %, respectively. These results indicate that a spatially uniform and representative training set is more effective in constraining regional PWV variability and suppressing systematic errors.</p>
      <p id="d2e2411">In contrast, the surrounding configuration exhibits the poorest performance among the three layouts. This degradation is likely attributable to the spatial mismatch between training and validation stations: most of the remaining validation stations are concentrated in the mid- and low-latitude regions, where water vapor variability is stronger and more dynamic. The lack of representative training information for these regions limits the model's ability to characterize intense moisture variability, leading to increased random errors and systematic bias during spatial extrapolation.</p>
      <p id="d2e2414">Overall, Table 4 show that, for regional-scale PWV calibration, spatial representativeness and coverage uniformity of training stations are more critical than merely increasing station numbers. Local clustering or surrounding layouts do not guarantee improved accuracy and may even degrade spatial extrapolation capability, whereas uniformly distributed random configurations offer the most robust and transferable solution for GNSS-constrained satellite PWV calibration.</p>
      <p id="d2e2418">Figure 7 compares model performance under different station layouts while keeping the training station number fixed (No. <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">200</mml:mn></mml:mrow></mml:math></inline-formula>). Figure 7d–f visualizes the spatial difference patterns of RF_N244 retrievals under distinct station layouts, while Fig. 7g–i depicts the actual geographical arrangements corresponding to clustered, surrounding, and randomly distributed stations, respectively. Even with identical sample sizes, significant performance differences emerge among spatial configurations: random distributions perform best (RMSE: 1.44 <inline-formula><mml:math id="M91" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>, Bias: <inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.34</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M93" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>), followed by surrounding configurations (RMSE: 1.75 <inline-formula><mml:math id="M94" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>, Bias: <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.68</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M96" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>), while clustered configurations perform worst (RMSE: 2.39 <inline-formula><mml:math id="M97" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>, Bias: <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.36</mml:mn></mml:mrow></mml:math></inline-formula> <inline-formula><mml:math id="M99" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>). This finding highlights the decisive role of spatial coverage uniformity and representativeness in determining model generalization capability.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e2513">Comparison of retrieval results at 12:00 UTC on the 232nd day of 2023 based on training models with different station distributions. Among them, panel <bold>(a)</bold> displays  PWV, panel <bold>(b)</bold> represents GNSS station-derived PWV, panel <bold>(c)</bold> presents PWV retrieved from the model trained on 244 stations, panels <bold>(d)–(f)</bold>    show PWV retrieved from models trained with different station    distributions, and panels <bold>(g)–(i)</bold> illustrate the    distribution of GNSS stations.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f07.png"/>

        </fig>

      <p id="d2e2540">Further inspection reveals that performance differences among configurations are relatively small over low-elevation plains, whereas the largest discrepancies occur in regions characterized by strong topographic and climatic gradients, such as plateaus, basins, and transition zones. A more balanced station layout provides more comprehensive spatial constraints, thereby improving the representation of water vapor variability in complex terrain. For instance, compared with a surrounding configuration, a more uniformly distributed configuration yields markedly improved performance over the Tibetan Plateau and the Yunnan–Guizhou Plateau, indicating reduced systematic bias at high elevations.</p>
      <p id="d2e2543">It is noteworthy that although clustered configurations increase station density locally, they result in the poorest overall performance, as shown Fig. 7d. In regions with strong spatial heterogeneity in water vapor and surface conditions, locally clustered observations can introduce spatial bias in the training dataset, making it difficult for a single global mapping function to simultaneously represent diverse environments such as plateaus, mountains, plains, lakes, and urban areas. Consequently, fitting errors accumulate and spatial extrapolation capability deteriorates. These results demonstrate that blindly increasing station density through localized clustering may be counterproductive, and that optimized spatial distribution is essential for robust regional calibration.</p>
      <p id="d2e2546">Taken together, the spatial analysis confirms that while training station number is important, optimizing spatial layout and coverage uniformity under a given sample size is even more critical for improving calibration accuracy and robustness in practical applications.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Temporal Stability of Model Performance</title>
      <p id="d2e2557">Finally, we evaluate temporal stability to assess whether optimizations in station number and layout yield robust performance under dynamic atmospheric conditions. Figure 8 presents PWV time series comparisons at three GNSS stations located at different latitudes. All three stations exhibit pronounced seasonal cycles, characterized by high PWV in summer and low PWV in winter. Across different training configurations, RF-based models generally capture the seasonal evolution of GNSS PWV well; however, differences remain in amplitude representation and short-term variability.</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e2562">Time-series comparison of PWV retrievals from models trained with different numbers of stations.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f08.png"/>

        </fig>

      <p id="d2e2571">When training station numbers are small (e.g., RF_No. <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mn mathvariant="normal">30</mml:mn></mml:mrow></mml:math></inline-formula>), the model tends to underestimate peak PWV during high-moisture periods, with evident smoothing effects in some intervals. As training station numbers increase to 80 and 200, agreement with GNSS observations improves substantially, particularly at mid- and low-latitude stations. This improvement indicates that enlarging the training dataset enables the model to better learn both the amplitude and temporal structure of PWV variability.</p>
      <p id="d2e2586">Analysis of error time series further reveals that, although errors generally fluctuate around zero, their magnitude and dispersion strongly depend on training station number. Sparse training configurations exhibit larger positive and negative oscillations and higher error dispersion, whereas increasing training station numbers leads to more concentrated error distributions and a marked reduction in extreme errors. These results demonstrate that increasing training station numbers enhances not only mean accuracy but also temporal stability and robustness. Nevertheless, occasional synchronized biases across different training configurations persist during some high-PWV periods, suggesting that retrieval errors are not solely controlled by training sample size, but may also be influenced by rapid moisture transport, strong convective activity, and observation or matching uncertainties.</p>
      <p id="d2e2589">Monthly-scale variations in Bias and RMSE under different training station numbers are shown in Fig. 9, with the original  PWV product included for comparison. Overall, calibrated models exhibit substantially lower Bias and RMSE than the original  product throughout the year, while also displaying a clear seasonal dependence, with larger errors in summer and smaller errors in winter.</p>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e2598">Comparison of the accuracy of  PWV and their calibrated versions using different models based on GNSS validation.</p></caption>
          <graphic xlink:href="https://acp.copernicus.org/articles/26/12243/2026/acp-26-12243-2026-f09.png"/>

        </fig>

      <p id="d2e2609">In terms of Bias, monthly mean values fluctuate around zero for all calibrated models, but their amplitudes depend on training station number. Models trained with fewer stations show more pronounced positive and negative deviations in certain months, particularly during seasonal transitions or under high-PWV conditions. As training station numbers increase to 80–120 and beyond, monthly Bias variations converge markedly toward zero, indicating effective suppression of systematic errors and improved monthly stability.</p>
      <p id="d2e2612">RMSE exhibits a consistent seasonal pattern across all models, with maximum values occurring during June–August and minimum values in winter months. Under sparse training conditions, RMSE during high-PWV months can reach 4–5 <inline-formula><mml:math id="M101" display="inline"><mml:mrow class="unit"><mml:mi mathvariant="normal">mm</mml:mi></mml:mrow></mml:math></inline-formula>. As training station numbers increase, overall RMSE decreases significantly and monthly variability becomes smoother. When training station numbers reach approximately 120–200, RMSE differences among months are substantially reduced, demonstrating enhanced robustness under varying seasonal moisture regimes.</p>
      <p id="d2e2625">It should be noted that even with increased training station numbers, RMSE remains relatively high during summer months characterized by intense moisture variability. This suggests that monthly-scale errors are not solely determined by training sample size, but are also influenced by enhanced convective activity, rapid moisture transport, and increased observational uncertainties. Nevertheless, the convergence of Bias and RMSE with increasing training station numbers provides a solid basis for further investigations into seasonal dependence and extreme moisture processes.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d2e2637">This study developed a GNSS-constrained calibration framework for  PWV and systematically investigated how training station number and spatial configuration influence all-weather PWV model's accuracy, spatial generalization, and temporal stability at the regional scale over China. By first reconstructing all-weather  PWV fields and subsequently applying a random forest–based calibration using GNSS PWV as an external constraint, the proposed approach effectively reduced systematic bias and random errors in satellite-derived PWV. The results demonstrate that, while high in-sample fitting accuracy can be achieved with relatively limited training data, the spatial generalization capability of the calibration model is strongly controlled by the density and spatial representativeness of GNSS training stations.</p>
      <p id="d2e2644">Comprehensive analyses across spatial, temporal, and seasonal dimensions reveal a clear performance saturation behavior. Increasing the number of training stations leads to substantial improvements in independent validation accuracy and spatial error homogeneity, particularly when training sample sizes increase from sparse to moderate levels. However, once the dominant spatial variability of PWV is sufficiently represented, further increases in training station numbers yield diminishing returns. For regional-scale applications over China, a training station number of approximately 120–200 provides an effective balance between accuracy, robustness, and computational efficiency, assuming a quasi-uniform spatial distribution. Moreover, experiments isolating spatial configuration effects demonstrate that station layout and coverage uniformity are at least as important as sample size, and that locally clustered station distributions may even degrade model performance in regions with strong moisture heterogeneity.</p>
      <p id="d2e2647">Despite the overall improvements achieved, residual uncertainties remain, particularly during summer months characterized by intense water vapor variability and rapid moisture transport. These errors indicate that calibration performance is not solely governed by training sample characteristics, but is also influenced by atmospheric dynamics, surface–atmosphere interactions, and observation or matching uncertainties. Future work will focus on incorporating additional dynamic predictors related to moisture transport and convection, refining season-dependent calibration strategies, and extending the proposed framework to other satellite platforms and regions with sparse or unevenly distributed GNSS networks. Furthermore, further efforts will be directed toward achieving optimal calibration of satellite-derived water vapor using existing ground-based measurements in station-sparse regions, such as the Tibetan Plateau and oceanic islands.</p>
      <p id="d2e2650">Beyond the specific case of  PWV over China, this study elucidates a generalizable principle for calibrating satellite geophysical products under sparse ground constraints: when observational resources are limited, prioritizing spatially representative and uniformly distributed stations is more effective than pursuing a higher station count alone. This “layout-over-density” principle stems from a fundamental requirement for spatial generalization – the model must capture the spatial heterogeneity of the target variable. A uniformly distributed network maximizes the spatial representativeness of the training data with a minimal number of stations.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e2659">The generated  PWV calibration model are available from Zenodo, at DOI: <ext-link xlink:href="https://doi.org/10.5281/zenodo.18751647" ext-link-type="DOI">10.5281/zenodo.18751647</ext-link> (Ma, 2025). The  water vapor data are openly and freely available at <uri>http://satellite.nsmc.org.cn/</uri> (last access: 22 August 2026). The radiosonde data can be accessed through <uri>https://www.ncei.noaa.gov/products/weather-balloon/integrated-global-radiosonde-archive</uri> (last access: 22 August 2026). The NDVI data can be download at <uri>https://ladsweb.modaps.eosdis.nasa.gov/</uri> (last access: 22 August 2026). The SRTM DEM is available at <uri>http://www.resdc.cn/</uri> (last access: 22 August 2026). Any additional information or data used in this study are available from the corresponding author upon reasonable request. The raw GNSS data are proprietary and cannot be shared publicly due to privacy and confidentiality restrictions imposed by the data provider. Access to the restricted data can be requested by contacting Zhihao Wang at zhihaowang1997@outlook.com. Requests are subject to approval by the data owner and may require a data-use agreement; the data will be provided for non-commercial academic research only.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e2685">CRediT: Yongchao Ma: Data curation, Formal analysis, Investigation, Methodology, Visualization, Writing – original draft; Zhengsheng Chen: Data curation, Project administration; Tong Liu: Methodology, Validation, Writing – review and editing; Zhibin Yu: Resources, Supervision; Zhihao Wang: Methodology, Resources, Validation, Writing – review and editing.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e2691">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e2697">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e2703">The reviewers' and editors' comments are highly appreciated. Acknowledgement is made to the Fengyun Satellite Data Center, the Integrated Global Radiosonde Archive (IGRA, provided by NOAA), the European Centre for Medium-Range Weather Forecasts (ECMWF), and the National Aeronautics and Space Administration (NASA) for providing the data.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e2708">This study was financially supported by the Science  and Technology Program of Guangdong Province (grant no. 2025B1212050001); in part by the Natural Science Basic Research Program of Shaanxi Province (grant no. 2026JC-YBQN-0397); in part by Shenzhen Science and Technology Program (grant  no. JCYJ20240813105116022); and in part by Rocket Force University of Engineering Foundation for Young Scientists (grant no. 2024QN-B037).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e2714">This paper was edited by Rolf Müller and reviewed by four anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><mixed-citation>Bai, J., Lou, Y., Zhang, W., Zhou, Y., Zhang, Z., and Shi, C.: Assessment and calibration of MODIS precipitable water vapor products based on GPS network over China, Atmos. Res., 254, 105504, <ext-link xlink:href="https://doi.org/10.1016/j.atmosres.2021.105504" ext-link-type="DOI">10.1016/j.atmosres.2021.105504</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><mixed-citation>Chen, B. and Liu, Z.: Assessing the performance of troposphere tomographic modeling using multi-source water vapor data during Hong Kong's rainy season from May to October 2013, Atmos. Meas. Tech., 9, 5249–5263, <ext-link xlink:href="https://doi.org/10.5194/amt-9-5249-2016" ext-link-type="DOI">10.5194/amt-9-5249-2016</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><mixed-citation>Du, Z., Zhang, B., Yao, Y., Zhao, Q., and Zhang, L.: Integrating near-infrared, thermal infrared, and microwave satellite observations to retrieve high-resolution precipitable water vapor, Remote Sens. Environ., 318, 114611, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2025.114611" ext-link-type="DOI">10.1016/j.rse.2025.114611</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><mixed-citation>Durre, I., Vose, R. S., and Wuertz, D. B.: Overview of the Integrated Global Radiosonde Archive, J. Climate, 19, 53–68, <ext-link xlink:href="https://doi.org/10.1175/JCLI3594.1" ext-link-type="DOI">10.1175/JCLI3594.1</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><mixed-citation>He, J. and Liu, Z.: Refining MODIS NIR atmospheric water vapor retrieval algorithm using GPS-derived water vapor data, IEEE T. Geosci. Remote, 59, 3682–3694, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2020.3016655" ext-link-type="DOI">10.1109/TGRS.2020.3016655</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><mixed-citation>Jiang, N., Wu, Y., Li, S., Xu, Y., Wang, Y., and Xu, T.: First PWV Retrieval Using MERSI-LL Onboard FY-3E and Cross Validation With Co-Platform Occultation and Ground GNSS, Geophys. Res. Lett., 51, e2024GL108681, <ext-link xlink:href="https://doi.org/10.1029/2024GL108681" ext-link-type="DOI">10.1029/2024GL108681</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><mixed-citation>Kaufman, Y. J. and Gao, B.-C.: Remote sensing of water vapor in the near IR from EOS/MODIS, IEEE T. Geosci. Remote, 30, 871–884, <ext-link xlink:href="https://doi.org/10.1109/36.175321" ext-link-type="DOI">10.1109/36.175321</ext-link>, 1992.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><mixed-citation>Li, Q., You, X., Yang, S., Du, R., Qiao, X., Zou, R., and Wang, Q.: A precise velocity field of tectonic deformation in China as inferred from intensive GPS observations, Sci. China Earth Sci., 55, 695–698, <ext-link xlink:href="https://doi.org/10.1007/s11430-012-4412-5" ext-link-type="DOI">10.1007/s11430-012-4412-5</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><mixed-citation>Lu, C., Li, X., Li, Z., Heinkelmann, R., Nilsson, T., Dick, G., Ge, M., and Schuh, H.: GNSS tropospheric gradients with high temporal resolution and their effect on precise positioning, J. Geophys. Res.-Atmos., 121, 912–930, <ext-link xlink:href="https://doi.org/10.1002/2015JD024255" ext-link-type="DOI">10.1002/2015JD024255</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><mixed-citation>Ma, X., Yao, Y., Zhang, B., and Du, Z.: FY-3A/MERSI precipitable water vapor reconstruction and calibration using multi-source observation data based on a generalized regression neural network, Atmos. Res., 265, 105893, <ext-link xlink:href="https://doi.org/10.1016/j.atmosres.2021.105893" ext-link-type="DOI">10.1016/j.atmosres.2021.105893</ext-link>, 2022a.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><mixed-citation>Ma, X., Yao, Y., Zhang, B., and He, C.: Retrieval of high spatial resolution precipitable water vapor maps using heterogeneous earth observation data, Remote Sens. Environ., 278, 113100, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2022.113100" ext-link-type="DOI">10.1016/j.rse.2022.113100</ext-link>, 2022b.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><mixed-citation>Ma, X., Yao, Y., Zhang, B., Qin, Y., Zhang, Q., and Zhu, H.: An Improved MODIS NIR PWV Retrieval Algorithm Based on an Artificial Neural Network Considering the Land-Cover Types, IEEE T. Geosci. Remote., 60, 1–12, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2022.3170078" ext-link-type="DOI">10.1109/TGRS.2022.3170078</ext-link>, 2022c.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><mixed-citation>Ma, Y.: Calibration model for FY-4A PWV based on different GNSS station network, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.18751647" ext-link-type="DOI">10.5281/zenodo.18751647</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><mixed-citation>Ma, Y., Liu, T., Yu, Z., Jiang, C., Xu, G., and Lu, Z.: All-weather precipitable water vapor map reconstruction using data fusion and machine learning-based spatial downscaling, Atmos. Res., 296, 107068, <ext-link xlink:href="https://doi.org/10.1016/j.atmosres.2023.107068" ext-link-type="DOI">10.1016/j.atmosres.2023.107068</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><mixed-citation>Merrikhpour, M. H. and Rahimzadegan, M.: Improving the Algorithm of Extracting Regional Total Precipitable Water Vapor Over Land From MODIS Images, IEEE T. Geosci. Remote, 55, 5889–5898, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2017.2716414" ext-link-type="DOI">10.1109/TGRS.2017.2716414</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><mixed-citation>Namaoui, H., Kahlouche, S., Belbachir, A. H., Van Malderen, R., Brenot, H., and Pottiaux, E.: GPS water vapor and its comparison with radiosonde and ERA-Interim data in Algeria, Adv. Atmos. Sci., 34, 623–634, <ext-link xlink:href="https://doi.org/10.1007/s00376-016-6111-1" ext-link-type="DOI">10.1007/s00376-016-6111-1</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><mixed-citation>Qin, Y., Wang, Y., Zhang, B., Fang, X., Yao, Y., and Ma, X.: A Novel Model Integrating the Spherical Cap Harmonic Analysis with the XGBoost Algorithm to Improve the MODIS NIR PWV, IEEE T. Geosci. Remote, 61, 1–12, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2023.3326659" ext-link-type="DOI">10.1109/TGRS.2023.3326659</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><mixed-citation>Rocken, C., Anthes, R., Exner, M., Hunt, D., Sokolovskiy, S., Ware, R., Gorbunov, M., Schreiner, W., Feng, D., Herman, B., Kuo, Y.-H., and Zou, X.: Analysis and validation of GPS/MET data in the neutral atmosphere, J. Geophys. Res., 102, 29849–29866, <ext-link xlink:href="https://doi.org/10.1029/97JD02400" ext-link-type="DOI">10.1029/97JD02400</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><mixed-citation>Saastamoinen, J.: Atmospheric correction for the troposphere and stratosphere in radio ranging satellites, in: The use of artificial satellites for geodesy [Internet], American Geophysical Union (AGU), 247–251, <ext-link xlink:href="https://doi.org/10.1029/GM015p0247" ext-link-type="DOI">10.1029/GM015p0247</ext-link>, 1972.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><mixed-citation>Sun, Q., Ji, D., Letu, H., Ni, X., Zhang, H., Wang, Y., Li, B., and Shi, J.: A method for estimating high spatial resolution total precipitable water in all-weather condition by fusing satellite near-infrared and microwave observations, Remote Sens. Environ., 302, 113952, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2023.113952" ext-link-type="DOI">10.1016/j.rse.2023.113952</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><mixed-citation>Trenberth, K. E., Fasullo, J., and Smith, L.: Trends and variability in column-integrated atmospheric water vapor, Clim. Dynam., 24, 741–758, <ext-link xlink:href="https://doi.org/10.1007/s00382-005-0017-4" ext-link-type="DOI">10.1007/s00382-005-0017-4</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><mixed-citation>Wang, Z., Chai, H., Zhu, C., Ma, H., Zheng, N., and Chen, P.: Reconstruction of high-resolution precipitable water vapor of  based on GNSS and remote sensing data, Meas. Sci. Technol., 37, 025803, <ext-link xlink:href="https://doi.org/10.1088/1361-6501/ae2b8d" ext-link-type="DOI">10.1088/1361-6501/ae2b8d</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><mixed-citation>Ware, R. H., Fulker, D. W., Stein, S. A., Anderson, D. N., Avery, S. K., Clark, R. D., Droegemeier, K. K., Kuettner, J. P., Minster, J. B., and Sorooshian, S.: SuomiNet: A Real–Time National GPS Network for Atmospheric Research and Education, B. Am. Meteorol. Soc., 81, 677–694, <ext-link xlink:href="https://doi.org/10.1175/1520-0477(2000)081&lt;0677:SARNGN&gt;2.3.CO;2" ext-link-type="DOI">10.1175/1520-0477(2000)081&lt;0677:SARNGN&gt;2.3.CO;2</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><mixed-citation>Xu, J. and Liu, Z.: Long-Term Calibration of Satellite-Based All-Weather Precipitable Water Vapor Product From FengYun-3A MERSI Near-Infrared Bands From 2010 to 2017 in China, IEEE T. Geosci. Remote, 61, 1–14, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2023.3300880" ext-link-type="DOI">10.1109/TGRS.2023.3300880</ext-link>, 2023a.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><mixed-citation>Xu, J. and Liu, Z.: Improving the Accuracy of MODIS Near-Infrared Water Vapor Product Under all Weather Conditions Based on Machine Learning Considering Multiple Dependence Parameters, IEEE T. Geosci. Remote, 61, 1–15, <ext-link xlink:href="https://doi.org/10.1109/TGRS.2023.3252024" ext-link-type="DOI">10.1109/TGRS.2023.3252024</ext-link>, 2023b. </mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><mixed-citation>Zhang, B., Yao, Y., Xin, L., and Xu, X.: Precipitable water vapor fusion: an approach based on spherical cap harmonic analysis and Helmert variance component estimation, J. Geod., 93, 2605–2620, <ext-link xlink:href="https://doi.org/10.1007/s00190-019-01322-1" ext-link-type="DOI">10.1007/s00190-019-01322-1</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><mixed-citation>Zhao, Q., Ma, Z., Yin, J., Yao, Y., Yao, W., Du, Z., and Wang, W.: General method of precipitable water vapor retrieval from remote sensing satellite near-infrared data, Remote Sens. Environ., 308, 114180, <ext-link xlink:href="https://doi.org/10.1016/j.rse.2024.114180" ext-link-type="DOI">10.1016/j.rse.2024.114180</ext-link>, 2024.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Measurement report: Quantifying the trade-off between station number and spatial layout in sparse GNSS networks for calibrating all-weather FY-4A precipitable water vapor</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
       Bai, J., Lou, Y., Zhang, W., Zhou, Y., Zhang, Z., and Shi, C.: Assessment and calibration of MODIS precipitable water vapor products based on GPS network over China, Atmos. Res., 254, 105504, <a href="https://doi.org/10.1016/j.atmosres.2021.105504" target="_blank">https://doi.org/10.1016/j.atmosres.2021.105504</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
       Chen, B. and Liu, Z.: Assessing the performance of troposphere tomographic modeling using multi-source water vapor data during Hong Kong's rainy season from May to October 2013, Atmos. Meas. Tech., 9, 5249–5263, <a href="https://doi.org/10.5194/amt-9-5249-2016" target="_blank">https://doi.org/10.5194/amt-9-5249-2016</a>, 2016. 
    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
       Du, Z., Zhang, B., Yao, Y., Zhao, Q., and Zhang, L.: Integrating near-infrared, thermal infrared, and microwave satellite observations to retrieve high-resolution precipitable water vapor, Remote Sens. Environ., 318, 114611, <a href="https://doi.org/10.1016/j.rse.2025.114611" target="_blank">https://doi.org/10.1016/j.rse.2025.114611</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
       Durre, I., Vose, R. S., and Wuertz, D. B.: Overview of the Integrated Global Radiosonde Archive, J. Climate, 19, 53–68, <a href="https://doi.org/10.1175/JCLI3594.1" target="_blank">https://doi.org/10.1175/JCLI3594.1</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
       He, J. and Liu, Z.: Refining MODIS NIR atmospheric water vapor retrieval algorithm using GPS-derived water vapor data, IEEE T. Geosci. Remote, 59, 3682–3694, <a href="https://doi.org/10.1109/TGRS.2020.3016655" target="_blank">https://doi.org/10.1109/TGRS.2020.3016655</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
       Jiang, N., Wu, Y., Li, S., Xu, Y., Wang, Y., and Xu, T.: First PWV Retrieval Using MERSI-LL Onboard FY-3E and Cross Validation With Co-Platform Occultation and Ground GNSS, Geophys. Res. Lett., 51, e2024GL108681, <a href="https://doi.org/10.1029/2024GL108681" target="_blank">https://doi.org/10.1029/2024GL108681</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
       Kaufman, Y. J. and Gao, B.-C.: Remote sensing of water vapor in the near IR from EOS/MODIS, IEEE T. Geosci. Remote, 30, 871–884, <a href="https://doi.org/10.1109/36.175321" target="_blank">https://doi.org/10.1109/36.175321</a>, 1992.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
       Li, Q., You, X., Yang, S., Du, R., Qiao, X., Zou, R., and Wang, Q.: A precise velocity field of tectonic deformation in China as inferred from intensive GPS observations, Sci. China Earth Sci., 55, 695–698, <a href="https://doi.org/10.1007/s11430-012-4412-5" target="_blank">https://doi.org/10.1007/s11430-012-4412-5</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
       Lu, C., Li, X., Li, Z., Heinkelmann, R., Nilsson, T., Dick, G., Ge, M., and Schuh, H.: GNSS tropospheric gradients with high temporal resolution and their effect on precise positioning, J. Geophys. Res.-Atmos., 121, 912–930, <a href="https://doi.org/10.1002/2015JD024255" target="_blank">https://doi.org/10.1002/2015JD024255</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
       Ma, X., Yao, Y., Zhang, B., and Du, Z.: FY-3A/MERSI precipitable water vapor reconstruction and calibration using multi-source observation data based on a generalized regression neural network, Atmos. Res., 265, 105893, <a href="https://doi.org/10.1016/j.atmosres.2021.105893" target="_blank">https://doi.org/10.1016/j.atmosres.2021.105893</a>, 2022a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
       Ma, X., Yao, Y., Zhang, B., and He, C.: Retrieval of high spatial resolution precipitable water vapor maps using heterogeneous earth observation data, Remote Sens. Environ., 278, 113100, <a href="https://doi.org/10.1016/j.rse.2022.113100" target="_blank">https://doi.org/10.1016/j.rse.2022.113100</a>, 2022b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
       Ma, X., Yao, Y., Zhang, B., Qin, Y., Zhang, Q., and Zhu, H.: An Improved MODIS NIR PWV Retrieval Algorithm Based on an Artificial Neural Network Considering the Land-Cover Types, IEEE T. Geosci. Remote., 60, 1–12, <a href="https://doi.org/10.1109/TGRS.2022.3170078" target="_blank">https://doi.org/10.1109/TGRS.2022.3170078</a>, 2022c.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
       Ma, Y.: Calibration model for FY-4A PWV based on different GNSS station network, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.18751647" target="_blank">https://doi.org/10.5281/zenodo.18751647</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
       Ma, Y., Liu, T., Yu, Z., Jiang, C., Xu, G., and Lu, Z.: All-weather precipitable water vapor map reconstruction using data fusion and machine learning-based spatial downscaling, Atmos. Res., 296, 107068, <a href="https://doi.org/10.1016/j.atmosres.2023.107068" target="_blank">https://doi.org/10.1016/j.atmosres.2023.107068</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
       Merrikhpour, M. H. and Rahimzadegan, M.: Improving the Algorithm of Extracting Regional Total Precipitable Water Vapor Over Land From MODIS Images, IEEE T. Geosci. Remote, 55, 5889–5898, <a href="https://doi.org/10.1109/TGRS.2017.2716414" target="_blank">https://doi.org/10.1109/TGRS.2017.2716414</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
       Namaoui, H., Kahlouche, S., Belbachir, A. H., Van Malderen, R., Brenot, H., and Pottiaux, E.: GPS water vapor and its comparison with radiosonde and ERA-Interim data in Algeria, Adv. Atmos. Sci., 34, 623–634, <a href="https://doi.org/10.1007/s00376-016-6111-1" target="_blank">https://doi.org/10.1007/s00376-016-6111-1</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
      
Qin, Y., Wang, Y., Zhang, B., Fang, X., Yao, Y., and Ma, X.: A Novel Model Integrating the Spherical Cap Harmonic Analysis with the XGBoost Algorithm to Improve the MODIS NIR PWV, IEEE T. Geosci. Remote, 61, 1–12, <a href="https://doi.org/10.1109/TGRS.2023.3326659" target="_blank">https://doi.org/10.1109/TGRS.2023.3326659</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
      
Rocken, C., Anthes, R., Exner, M., Hunt, D., Sokolovskiy, S., Ware, R., Gorbunov, M., Schreiner, W., Feng, D., Herman, B., Kuo, Y.-H., and Zou, X.: Analysis and validation of GPS/MET data in the neutral atmosphere, J. Geophys. Res., 102, 29849–29866, <a href="https://doi.org/10.1029/97JD02400" target="_blank">https://doi.org/10.1029/97JD02400</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
      
Saastamoinen, J.: Atmospheric correction for the troposphere and stratosphere in radio ranging satellites, in: The use of artificial satellites for geodesy [Internet], American Geophysical Union (AGU), 247–251, <a href="https://doi.org/10.1029/GM015p0247" target="_blank">https://doi.org/10.1029/GM015p0247</a>, 1972.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
       Sun, Q., Ji, D., Letu, H., Ni, X., Zhang, H., Wang, Y., Li, B., and Shi, J.: A method for estimating high spatial resolution total precipitable water in all-weather condition by fusing satellite near-infrared and microwave observations, Remote Sens. Environ., 302, 113952, <a href="https://doi.org/10.1016/j.rse.2023.113952" target="_blank">https://doi.org/10.1016/j.rse.2023.113952</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
       Trenberth, K. E., Fasullo, J., and Smith, L.: Trends and variability in column-integrated atmospheric water vapor, Clim. Dynam., 24, 741–758, <a href="https://doi.org/10.1007/s00382-005-0017-4" target="_blank">https://doi.org/10.1007/s00382-005-0017-4</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
       Wang, Z., Chai, H., Zhu, C., Ma, H., Zheng, N., and Chen, P.: Reconstruction of high-resolution precipitable water vapor of  based on GNSS and remote sensing data, Meas. Sci. Technol., 37, 025803, <a href="https://doi.org/10.1088/1361-6501/ae2b8d" target="_blank">https://doi.org/10.1088/1361-6501/ae2b8d</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
       Ware, R. H., Fulker, D. W., Stein, S. A., Anderson, D. N., Avery, S. K., Clark, R. D., Droegemeier, K. K., Kuettner, J. P., Minster, J. B., and Sorooshian, S.: SuomiNet: A Real–Time National GPS Network for Atmospheric Research and Education, B. Am. Meteorol. Soc., 81, 677–694, <a href="https://doi.org/10.1175/1520-0477(2000)081&lt;0677:SARNGN&gt;2.3.CO;2" target="_blank">https://doi.org/10.1175/1520-0477(2000)081&lt;0677:SARNGN&gt;2.3.CO;2</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
       Xu, J. and Liu, Z.: Long-Term Calibration of Satellite-Based All-Weather Precipitable Water Vapor Product From FengYun-3A MERSI Near-Infrared Bands From 2010 to 2017 in China, IEEE T. Geosci. Remote, 61, 1–14, <a href="https://doi.org/10.1109/TGRS.2023.3300880" target="_blank">https://doi.org/10.1109/TGRS.2023.3300880</a>, 2023a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
       Xu, J. and Liu, Z.: Improving the Accuracy of MODIS Near-Infrared Water Vapor Product Under all Weather Conditions Based on Machine Learning Considering Multiple Dependence Parameters, IEEE T. Geosci. Remote, 61, 1–15, <a href="https://doi.org/10.1109/TGRS.2023.3252024" target="_blank">https://doi.org/10.1109/TGRS.2023.3252024</a>, 2023b.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
       Zhang, B., Yao, Y., Xin, L., and Xu, X.: Precipitable water vapor fusion: an approach based on spherical cap harmonic analysis and Helmert variance component estimation, J. Geod., 93, 2605–2620, <a href="https://doi.org/10.1007/s00190-019-01322-1" target="_blank">https://doi.org/10.1007/s00190-019-01322-1</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
       Zhao, Q., Ma, Z., Yin, J., Yao, Y., Yao, W., Du, Z., and Wang, W.: General method of precipitable water vapor retrieval from remote sensing satellite near-infrared data, Remote Sens. Environ., 308, 114180, <a href="https://doi.org/10.1016/j.rse.2024.114180" target="_blank">https://doi.org/10.1016/j.rse.2024.114180</a>, 2024.

    </mixed-citation></ref-html>--></article>
