Section: Workflow architecture | AGROECOLVRE - Soil and subsoil carbon sequestration metadata analysis as a function of land uses and agroecological practices | LifeWatch ERIC Training Platform

Main course page

Workflow architecture

  • Workflow architecture

    The process is organized across four sequential phases: it starts with Step 1, dedicated to the import and preliminary cleaning of the datasets (LUCAS-Core 2009 and LUCAS-Topsoil 2018), before moving on to the exploration, filtering, and selection phase of the sampling points and variables of interest (Step 2a and Step 2b), with the simultaneous generation of summary tables. Finally, the data are subjected to two different advanced statistical analysis approaches, Principal Component Analysis (Step 3, PCA) and Analysis of Variance (Step 4, ANOVA), both designed to produce analytical outputs in both tabular (.csv) and graphical (.png) formats (Fig. 1).

    Fig. 1. Illustrative workflow diagram

    Fig. 1. Illustrative workflow diagram

    STEP 1. Reading in LUCAS data

    This initial service reads the core and topsoil datasets for 2009 and 2018 performing also a data cleaning and standardisations. If personal raw LUCAS-Topsoil datasets are used, they can’t be distributed with any 3rd party, and under any circumstance, outside the VRE due to data governance constraints.


    STEP 2a. LUCAS data selection and exploration

    This step allow users to filter and explore the synchronized datasets, selecting LUCAS-Soil 2018 points that share the exact same Land Use (LU) as in 2009. It aggregates points by region, counting the total sampled rows for both Topsoil and Core datasets. It counts unchanged points by region and Land Cover (LC) class (Broadleaf, Coniferous, Shrubland, and mixed Forest+Shrubland).

    Output Generated:

    A comprehensive document containing a summary table with 8 specific columns tracking the point counts per region, designed to easily fit a standard page with a caption (summary_table1.docx).


    STEP 2b. LUCAS data selection (Area of Study)

    In this step, the user defines the geographical boundaries and relevant variables for the analysis for both PCA and ANOVA.

    Parameters: The user can provide a specific region study argument (as a vector supporting NUTS0, NUTS1, NUTS2, or NUTS3). If left blank, the system applies ‘ES11’ (Galicia, Spain) as the default baseline region. 

    Filtering: The workflow restricts rows exclusively to forests, shrublands, and grasslands, isolating the specific land covers relevant to the agroecological study. 


    STEP 3. Principal Component Analysis (PCA)

    Executes a PCA on the eight primary LUCAS soil chemical variables to reduce dimensionality and extract patterns.

    Missing Data Management: To handle missing values (NAs), rows containing any incomplete data are excluded from the computation for simplicity and transparency.

    Outputs Generated:

    1. A PCA chart plotting PC1 against PC2 (pca_plot.png).
    2. A visual chart plotting PC1 against PC2 considering Land Cover class (pca_plot_byLC.png).
    3. An Excel spreadsheet summarizing the statistical results and eigenvalues of the PCA (pca_summary.xlsx)
    4. An Excel spreadsheet summarizing the statistical results and eigenvalues of the PCA considering Land Cover class (pca_summary_byLC.xlsx)


    STEP 4. ANOVA (Analysis of Variance)

    To evaluate whether the 8 soil chemical characteristics differ significantly among distinct ecosystem categories, a one-way ANOVA is conducted across the Land Cover groups. Post-hoc Tukey tests are implemented to identify specific group variances. 

    Outputs Generated:

    1. Exploratory boxplots for each land cover grouped by soil parameter(bxplt_variable_LC.png).
    2. Frequency distribution histograms for each category (histograms_variable_LC.png).
    3. Graphical representation of ANOVA results combined with Tukey honest significant differences (anova_boxplots_tukey.png).
    4. A structured spreadsheet containing the complete ANOVA statistical summary tables (ANOVA_summary.xlsx).