Skip to content
Global Ocean Accounts Partnership Technical Guidance

Integrating Research Data into Official Statistics

Circular ID TG-4.5
Version 6.0
Badge Applied
Status Draft
Last Updated May 2026

1. Outcome

1This Circular provides guidance on bridging the gap between scientific research data and official statistics for the compilation of ocean accounts. Ocean accounting requires diverse data sources that extend beyond traditional statistical surveys and administrative records to include oceanographic observations, biodiversity monitoring, ecosystem assessments, and other forms of research data generated by universities, research institutions, and international scientific programmes. This Circular addresses the systematic integration of research data into statistical production, maintaining the quality standards expected of official statistics whilst drawing on the distinct capabilities of scientific research organisations.

2By implementing the guidance in this Circular, practitioners will be able to identify research data sources relevant to ocean accounting applications, assess their fitness for statistical purposes using quality frameworks adapted from the United Nations National Quality Assurance Framework (UN NQAF), establish institutional partnerships with research organisations, and compile ocean accounts that combine traditional statistical sources with research data in a transparent and methodologically sound manner. The specific applications addressed include the use of oceanographic research data for ecosystem condition accounts (see TG-3.5 Ecosystem Condition), stock assessment data for fisheries accounts (see TG-6.7 Fisheries Stock Assessment), and biodiversity monitoring data for extent accounts (see TG-2.9 Ecosystem Extent).

3The quality assessment dimensions presented here build on the overarching quality framework established in TG-0.7 Quality Assurance, whilst the data harmonisation techniques needed to reconcile research data with other sources are detailed in TG-4.6 Data Harmonisation. Key terms used in this Circular are defined in TG-0.6 Glossary.

2. Requirements

1This Circular requires familiarity with:

  • 2

    TG-0.1 General Introduction to Ocean Accounts — provides foundational understanding of ocean accounts components and the relationship between environmental and economic accounting frameworks, including the conceptual basis for integrating diverse data sources into a coherent accounting system.

  • 3

    TG-0.7 Quality Assurance — establishes the overarching quality framework applicable to ocean accounting data, including the quality dimensions and assessment procedures that this Circular applies specifically to research data sources.

3. Guidance Material

1The compilation of ocean accounts frequently requires data from scientific research programmes that were not originally designed for statistical purposes. Research data can fill gaps in ocean observation and ecosystem monitoring that traditional statistical sources cannot address. Integrating such data into official statistics requires careful attention to quality assessment, metadata documentation, and institutional coordination. This Circular sets out a systematic approach to that integration.

2The quality considerations discussed here should be understood within the broader quality assurance framework described in TG-0.7 Quality Assurance. For guidance on reconciling data from multiple sources with different classifications and spatial boundaries, see TG-4.6 Data Harmonisation.

3.1 Decision Use Cases for Research Data

1Research data integration supports specific decision-making applications in ocean accounting. This section identifies the primary use cases where research data sources provide essential inputs that traditional statistical sources cannot supply.

3.1.1 Integration modes and governance accountability

1Before identifying specific use cases, practitioners should classify each prospective integration of research data into one of three modes, as each mode triggers different governance requirements within the NSO. This classification is material to the NSO’s legal and audit accountability, since substituting an official source with research data may require formal endorsement from the national statistical authority rather than a technical decision by the compiling unit.

2Mode A — Gap-filling. No official statistical source exists for the required variable. Research data are the only feasible input. Examples include subsurface dissolved oxygen profiles in pelagic waters, eDNA-based species occurrence in remote habitats, and acoustic biomass estimates for non-commercial species. Governance: technical judgement by the compiling unit, documented in account metadata under the standard quality-assessment workflow set out in Section 3.5.

3Mode B — Supplementary use. An official source exists but research data are used in parallel as a secondary validation source, a cross-check, or a higher-resolution overlay. Examples include using satellite chlorophyll products alongside ship-based water-quality monitoring, or using research stock assessments to validate administrative catch records. Governance: technical judgement by the compiling unit, with a coherence test required under Section 3.5.3 Step 3.3.

4Mode C — Primary substitution. An official source exists but is replaced by research data judged of higher quality. Examples include replacing a discontinued shipboard nutrient monitoring programme with an autonomous Argo-derived product, or substituting an obsolete remote-sensing land-cover product with a research-grade national mapping. Governance: requires chief statistician (or equivalent national statistical authority) endorsement, plus a documented UN NQAF Level A institutional-framework compliance check, before the substituted source can be published in an official account. The substitution decision and its endorsement must be recorded in the account compilation metadata.

5The mode applicable to a given source should be determined by the compiling NSO when applying the Section 3.5 procedure.

6Figure 4.5.1 summarises the screening path: three sequential admissibility gates (Q1—Q3) check variable coverage, metadata completeness, and quality-threshold compliance; mode routing gate (Q4) then classifies the admissible source as Mode A (gap-filling), Mode B (supplementary), or Mode C (primary substitution — chief statistician endorsement required before publication in official accounts).

Q1: Variable covered?Q2: NQAF metadata?Q3: Quality thresholds?Q4: Official source?Mode A: Gap-fillingMode B: SupplementaryNSO endorsementMode C: SubstitutionNot admissible

Figure 4.5.1 Research-data admissibility gates Q1–Q4 and integration modes A–C.

3.1.2 Ecosystem condition accounts

1Ecosystem condition accounts require measurements of biophysical and chemical variables that characterise the state of marine ecosystems1. For ocean environments, many of these measurements are only available through oceanographic research programmes:

2Oceanographic surveys provide temperature, salinity, dissolved oxygen, nutrient concentrations, and pH measurements throughout the water column. The Global Ocean Observing System (GOOS) coordinates international efforts to standardise observation methods and improve data sharing2. For condition accounts in pelagic waters (see TG-6.5 Pelagic and Open Ocean Accounting), research vessel surveys and autonomous profiling floats (such as the Argo network) are the primary sources for subsurface condition data.

3Acoustic surveys estimate biomass of schooling pelagic fish and benthic invertebrates using scientific echosounders. Regional fisheries management organisations (RFMOs) often conduct acoustic surveys to assess stock condition, providing data that can be integrated into ecosystem condition accounts where fish biomass serves as a condition indicator3.

4Water quality monitoring by environmental research agencies tracks pollution levels, turbidity, and harmful algal blooms. For coastal condition accounts, research monitoring of nitrogen and phosphorus loading provides essential data on eutrophication pressure.

5The SEEA Ecosystem Accounting framework identifies specific condition characteristics that typically require research data inputs, including biotic characteristics (species diversity, biomass, community composition), abiotic characteristics (water temperature, pH, salinity), and functional characteristics (primary productivity, nutrient cycling)4.

3.1.3 Fisheries stock assessment

1Stock assessment for commercial and subsistence fisheries depends heavily on scientific research data5:

2Catch-at-age data from research vessel surveys provide independent estimates of population abundance and age structure, complementing catch data reported by commercial vessels. For highly migratory species such as tuna, international research programmes coordinated through RFMOs are often the only source of systematic catch-at-age information across the species’ range.

3Tagging programmes using electronic tags, acoustic telemetry, and genetic markers reveal migration patterns, population structure, and survival rates. Close-kin mark-recapture methods using genetic analysis allow estimation of absolute abundance for high-value species such as southern bluefin tuna, where traditional survey methods are impractical6.

4Life history parameters including growth rates, natural mortality, and fecundity are typically derived from research laboratory studies and field sampling programmes rather than routine statistical collection.

5National Statistical Offices (NSOs) compiling fisheries accounts rely on stock assessments produced by fisheries research institutes to estimate the physical and monetary value of aquatic resources, as described in TG-3.1 Asset Accounts.

6Data-rich versus data-poor stock assessments. NSOs must distinguish between data-rich and data-poor stock assessment methods. Data-rich methods (e.g., age-structured Virtual Population Analysis, integrated statistical catch-at-age models) combine catch-at-age, abundance indices, and life-history parameters to produce biomass estimates with characterisable confidence intervals. Data-poor methods (e.g., Length-Based Spawning Potential Ratio (LBSPR), Catch-Maximum Sustainable Yield (CMSY), Depletion-Based Stock Reduction Analysis (DB-SRA)) rely on far fewer inputs and produce estimates with much wider, often order-of-magnitude, uncertainty ranges7. The quality tier reported alongside fisheries asset accounts (Tier 1, 2, or 3 in SEEA EA terminology) must reflect the underlying assessment method: data-rich age-structured assessments typically support Tier 3 reporting, whilst data-poor length-based or catch-only methods should generally be reported as Tier 1, with the methodological caveats made explicit in the account release notes. Publishing a fisheries asset account derived from a data-poor assessment without communicating that limitation can mislead policy audiences about the precision of the stock estimate.

3.1.4 Biodiversity and extent mapping

1Research programmes provide essential data for ecosystem extent accounts and biodiversity indicators:

2Benthic habitat mapping using multibeam echosounder surveys, remotely operated vehicles (ROVs), and drop cameras classifies seabed substrates and identifies ecosystem types such as cold-water coral reefs, sponge beds, and seagrass meadows. National hydrographic offices and marine research institutes are typically the custodians of these data.

3Species occurrence records from biodiversity surveys, museum collections, and citizen science programmes are aggregated in global repositories such as the Ocean Biodiversity Information System (OBIS) and the Global Biodiversity Information Facility (GBIF)8. For ecosystem extent accounts based on the IUCN Global Ecosystem Typology, these occurrence data support delineation of ecosystem functional groups.

4Coral reef monitoring through the Global Coral Reef Monitoring Network and regional programmes such as the Coral Triangle Initiative provides systematic assessments of coral cover, bleaching events, and reef condition that underpin extent and condition accounts for coral reef ecosystems (see TG-6.1 Coral Reef Ecosystem Accounting).

5The Framework for the Development of Environment Statistics (FDES) notes that “scientific research data can be used to address data gaps” in environmental statistics, particularly for parameters that require specialised measurement techniques9.

3.2 Types of Research Data

1Research data relevant to ocean accounting come from a range of sources, each with distinct characteristics, collection methodologies, and quality considerations. These distinctions determine which data sources are appropriate and how their fitness for accounting purposes is assessed.

3.2.1 Oceanographic observation data

1Oceanographic observation systems generate continuous or periodic measurements of physical, chemical, and biological ocean parameters10. These include temperature, salinity, currents, dissolved oxygen, nutrient concentrations, and chlorophyll levels. Major international programmes such as the Global Ocean Observing System (GOOS), the Argo float network, and regional ocean observing systems (e.g., the Integrated Marine Observing System in Australia) provide standardised data streams that can support ecosystem condition accounts11. The Intergovernmental Oceanographic Commission (IOC) of UNESCO coordinates international efforts to improve ocean observation capacity and data sharing12.

2Such data are typically collected using instrumented platforms including research vessels, moored buoys, autonomous underwater vehicles, satellite remote sensing, and profiling floats. The primary advantages include (1) systematic temporal coverage enabling trend analysis, (2) standardised measurement protocols developed through international scientific consensus, and (3) growing open access through global data repositories. However, spatial coverage can be uneven, with data density varying widely between coastal and open ocean areas, and between developed and developing country waters13.

3Among the most directly relevant oceanographic variables for ecosystem condition accounts are the Essential Ocean Variables (EOVs) defined by GOOS. The EOV framework organises ocean observations into physics, biogeochemistry, and biology/ecosystems domains14. For ecosystem condition accounts, priority EOVs include sea surface temperature, dissolved oxygen, inorganic carbon, ocean colour (as a proxy for phytoplankton biomass), and marine habitat properties. For ecosystem extent accounts, relevant EOVs include hard coral cover, seagrass cover, mangrove cover, and macroalgal canopy cover. Practitioners should consult the current GOOS EOV specification sheets, which document readiness levels, observation requirements, and data product availability for each variable, to determine which EOVs are feasible data sources for their national accounting context.

3.2.2 Biodiversity survey data

1Biodiversity surveys document species occurrence, abundance, and distribution through structured sampling programmes15. For marine environments, these include fish stock assessments, invertebrate surveys, marine mammal and seabird censuses, coral reef monitoring, and seagrass mapping exercises. Such data support ecosystem extent accounts (mapping ecosystem types), ecosystem condition accounts (assessing biodiversity indicators), and ecosystem services accounts (quantifying provisioning services such as fisheries).

2Biodiversity data are collected through diverse methodologies including visual census, acoustic surveys, environmental DNA (eDNA) sampling, and citizen science programmes16. The Global Biodiversity Information Facility (GBIF) aggregates species occurrence records from research institutions worldwide and provides standardised access through its data portal17. Regional initiatives such as the Ocean Biodiversity Information System (OBIS) focus specifically on marine biodiversity data18.

3Biodiversity surveys often follow sampling designs optimised for scientific research questions rather than complete spatial coverage. Such designs can bias sampling towards accessible locations or areas of particular scientific interest, a bias that must be addressed when the data are used for area-based ecosystem accounts19.

3.2.3 Ecosystem monitoring programmes

1Long-term ecosystem monitoring programmes track changes in ecosystem structure, function, and condition over time. Examples relevant to ocean accounting include coral reef monitoring networks (e.g., the Global Coral Reef Monitoring Network), mangrove forest assessments, seagrass habitat mapping, and kelp forest monitoring programmes20. These programmes typically combine remote sensing data with ground-truthing surveys to map ecosystem extent and assess condition indicators.

2The SEEA Ecosystem Accounting framework recommends using a tiered approach for ecosystem monitoring, with Tier 1 using globally available default data, Tier 2 using regionally appropriate data, and Tier 3 using nationally collected data with full spatial and temporal coverage21. Research monitoring programmes often provide the foundation for Tier 2 and Tier 3 approaches, particularly for marine ecosystems where routine statistical collection is limited.

3.2.4 Remote sensing and Earth observation data

1Satellite-based Earth observation provides systematic, repeated coverage of ocean and coastal areas at scales relevant to national accounting22. Key parameters measurable from space include sea surface temperature, ocean colour (indicative of chlorophyll and primary productivity), sea level, surface currents, coastal land cover change, and wetland extent. The European Union’s Copernicus Marine Service and NASA’s Ocean Biology Processing Group provide validated ocean data products derived from multiple satellite sensors23.

2Remote sensing data are valuable for their temporal frequency (enabling detection of change) and their spatial reach (enabling complete national coverage). Limitations include cloud cover interference, the inability to observe below the sea surface, and the requirement for ground-truthing to validate derived products. The SEEA Technical Guidance on Biophysical Modelling notes that remote sensing offers “enormous opportunities to disseminate data with very short time-lags and high-frequency”24.

3The major satellite platforms relevant to ocean accounting (Sentinel-2, Landsat, MODIS, and Sentinel-1 SAR) span spatial resolutions from 10 metres to 1 kilometre, with revisit times from daily (MODIS) to 16 days (Landsat). For full guidance on sensor characteristics and selection, see TG-4.1 Remote Sensing and Geospatial Data.25

4Atmospheric correction validation in coastal and optically complex waters. Standard atmospheric correction algorithms used for open-ocean (“Case 1”) waters, including DARK_PIXEL, SeaDAS default, and the Case-2 Regional CoastColour processor (C2RCC), are known to perform poorly in turbid, shallow, sediment-laden, or chlorophyll-rich coastal (“Case 2”) waters. The systematic biases this introduces propagate directly into derived products such as chlorophyll-a, suspended sediment, and turbidity, which are the principal remote sensing inputs to coastal condition accounts. For NSOs in tropical and developing country contexts, where Case 2 waters dominate the coastal zone, in situ validation of satellite-derived products is mandatory before use in official accounts. Practitioners should (i) confirm that the product is derived from a Case 2-capable processor and consult the Copernicus Marine Service product quality factsheet for the specific product version, (ii) verify validation status against EUMETSAT/ACRI atmospheric correction validation protocols, and (iii) where feasible, perform a matchup analysis against in situ radiometric measurements from the AERONET Ocean Color network (AERONET-OC) or a comparable national field campaign. Uncorrected or unvalidated products should not be published as the basis of official condition statistics for coastal waters.

3.2.5 Scientific research publications and datasets

1Peer-reviewed scientific publications are a source of coefficients, conversion factors, and methodological parameters for ecosystem service modelling26. For example, estimates of carbon sequestration rates in mangroves, nutrient retention by seagrass meadows, or coastal protection values from coral reefs are frequently derived from published research rather than direct national measurement. Whilst individual studies may be site-specific, synthesis studies and meta-analyses can provide generalised values applicable across similar ecosystem types.

2Many funders and publishers now require research datasets underlying publications to be deposited in open repositories as a condition of publication or funding. The Framework for the Development of Environment Statistics (FDES) notes that scientific research data “are usually available at no or low cost” and “can be used to address data gaps”27. However, the FDES also cautions that such data “often use terms and definitions that differ from those used in statistics”, may have “limited scope”, and are “often available on a one-time basis only”28.

3Coefficient selection protocol where multiple published estimates exist. It is common, especially for developing country ocean accounts compiling blue carbon, nutrient retention, or coastal protection coefficients, to encounter multiple published estimates for the same parameter that span an order of magnitude or more. Mangrove soil carbon sequestration estimates, for example, span roughly 0.5 to 15 Mg C ha⁻¹ yr⁻¹ in the published literature, with central tendency varying by biogeography, stand age, and methodology. Where multiple estimates exist, NSOs should follow this four-step coefficient selection protocol:

  1. 4Prefer systematic reviews and meta-analyses over individual studies. Where a published systematic review or meta-analysis exists for the parameter of interest, use its central estimate as the default value rather than averaging across individual studies. Meta-analyses adjust for study quality and weight estimates by sample size.
  2. 5Where multiple meta-analyses exist, document the range and select the most geographically and ecologically appropriate value. Geographic and ecological matching (e.g., tropical vs. temperate mangroves, or high-rainfall vs. arid coastal settings) should be explicit and justified in writing. A simple arithmetic mean of meta-analyses with different geographic scopes is not recommended.
  3. 6Apply a sensitivity analysis across the plausible range. Re-compute the affected account aggregate using at least the lower and upper bounds of the range of credible estimates, and report the resulting account values as a sensitivity envelope alongside the central estimate. This propagates coefficient uncertainty into the account.
  4. 7Document the selection rationale in account metadata. The selected coefficient, its source, the alternatives considered and rejected, the geographic-matching justification, and the sensitivity envelope must be recorded in the account compilation metadata to support audit and replication.

8Where SEEA EA biome-specific default values are available, these should be used as the baseline starting point for the protocol and any departure justified.

3.3 Quality Assessment Frameworks

1Research data must be assessed for fitness for statistical purposes before incorporation into ocean accounts. The quality dimensions applied to research data differ in emphasis from those applied to survey data, reflecting the distinct characteristics of scientific data production. For the overarching quality framework applicable to all ocean accounting data, see TG-0.7 Quality Assurance.

3.3.1 Dimensions of data quality

1The UN NQAF quality dimensions (relevance, accuracy, reliability, timeliness, accessibility, coherence, comparability) and their general definitions are in TG-0.7 Quality Assurance Principles29. Applied to research data, relevance turns on spatial/temporal/variable alignment with the account structure. Accuracy turns on sampling design, methodology validation, and error propagation in derived products, with SEEA EA Technical Guidance on Biophysical Modelling for modelled outputs30. Coherence turns on mapping research classifications to statistical standards (see TG-4.6 Data Harmonisation). Comparability turns on methodology breaks and documented bridging factors31. Accessibility turns on licensing, format compatibility, documented access protocols, and FAIR principles (Section 3.4.1).

2Minimum acceptability thresholds. Identifying the relevant dimensions is necessary but not sufficient: NSOs require a decision rule for when a research dataset can be admitted to an official account. Table 3.3.1 specifies, for each UN NQAF dimension, the conditions under which a data source is (a) acceptable without qualification, (b) acceptable with documented caveats, or (c) not acceptable for use in official accounts. The thresholds are anchored to international standards where these exist (e.g., the IOC/GOOS ±0.2 mg/L accuracy precedent for dissolved oxygen) and otherwise expressed as documentation requirements that allow defensible compilation. The ±0.2 mg/L figure is used here as an exemplar of how a published international standard can be carried directly into the integration decision. Equivalent dimension-specific thresholds should be substituted where they exist for other variables.

3Table 3.3.1: Minimum acceptability thresholds by UN NQAF quality dimension

DimensionAcceptable without qualificationAcceptable with documented caveatsNot acceptable
RelevanceSpatial coverage, temporal resolution, and measured variables directly match account requirementPartial match; transformation or scaling required and documentedVariable is conceptually misaligned with the account; no defensible transformation
AccuracyValidated against an international standard (e.g., IOC/GOOS ±0.2 mg/L for DO); error budget publishedAccuracy estimate available but not against a recognised standard; uncertainty disclosed in account releaseNo accuracy assessment available and no alternative validation feasible
ReliabilityDocumented sampling design, repeated comparable surveys, published precisionSingle-event survey with documented methods; precision inferred from methodMethods undocumented; no precision basis
CoherenceClassification crosswalk to SEEA/IUCN GET exists; co-located comparison availableMapping requires interpretive judgement; documented and reviewedCategories not mappable to statistical classifications
ComparabilityStable methodology across full time series; no breaksDocumented methodological change with bridging factors appliedUndocumented method change creating an unresolvable break
AccessibilityOpen licence, persistent identifier, machine-readable formatRestricted licence with documented terms; access via MoUNo defined access path; licence prohibits statistical use
TimelinessUpdate schedule compatible with the accounting cycleLag exceeds accounting cycle by no more than one period; documentedLag exceeds two accounting periods with no commitment to update

4Note on Mode C (primary substitution) entries: Where the integration mode is Mode C under Section 3.1.1 (i.e., a research dataset replaces an existing official source), all threshold assessments in Table 3.3.1 remain applicable, but the overall admission decision additionally requires chief-statistician (or equivalent national statistical authority) endorsement before the source can be published in an official account. Mode C sources that pass all Table 3.3.1 thresholds but lack this endorsement may not be published.

5A data source rated “Not acceptable” on any compilation-blocking dimension (accuracy, coherence, or accessibility) cannot be admitted to an official account without remediation. Sources rated “Acceptable with caveats” may be admitted provided the caveats are recorded in account metadata and surfaced in the account release.

3.3.2 Reproducibility, replicability, and scientific rigour

1Reproducibility, obtaining consistent results using the same data and methods32, is essential for updatable ocean accounts. Replicability, obtaining consistent results when new data are collected using the same methods33, provides confidence that coefficients derived in one context can be applied in another. The SEEA Technical Guidance on Biophysical Modelling emphasises that “transparency of approaches is essential”.34

2Three-level reproducibility standard for integration. To convert reproducibility from an aspirational requirement into an auditable gate, NSOs should apply the following three-level standard when assessing research data:

  • 3Level 1 — Minimum acceptable for integration of empirical (non-model) data. Methods are fully described in a peer-reviewed publication or equivalent technical report. Raw inputs and processing code need not be publicly archived, but the methods must be sufficient to allow a competent practitioner to reproduce the analysis given the inputs.
  • 4Level 2 — Preferred standard, and the minimum required for any model-based output (e.g., biophysical model results, modelled ecosystem service values, statistically interpolated condition indices). Processing code and processed data are archived in a public repository (e.g., Zenodo, GitHub) with documented inputs and a usable readme. Cited datasets carry persistent identifiers.
  • 5Level 3 — Highest standard. All of Level 2, plus archived raw inputs and a containerised or fully scripted processing environment (e.g., Docker, Singularity, or a versioned conda environment) that enables byte-exact reproduction of outputs.

6Legacy empirical datasets that pre-date open-data mandates may be integrated under Level 1, provided their methods are documented in peer-reviewed publications. Outputs of biophysical or statistical models must meet at least Level 2 before being admitted to an official account. This is the level at which audit exposure is highest, as the result depends on assumptions and code rather than direct measurement.

3.3.3 Uncertainty quantification

1Research data are characterised by various sources of uncertainty that must be documented and, where possible, quantified. Uncertainty can arise from sampling variability, measurement error, model parameter uncertainty, and structural model uncertainty. The SEEA Technical Guidance on Biophysical Modelling notes that “uncertainty matrices, which outline possible sources of uncertainty for each model” provide a basic approach to uncertainty documentation35.

2For ecosystem accounts derived from biophysical models, “outputs should be seen as best estimates, rather than absolute values” unless detailed parameterisation and validation has been conducted36. Statistical agencies should communicate uncertainty alongside point estimates, enabling users to assess fitness for their specific purposes. The tiered approach recommended in SEEA EA supports this, with lower tiers acknowledged to have greater uncertainty but serving valuable purposes for awareness-raising and broad trend analysis37.

3The UN NQAF recommends that “statistical agencies should publish information on the quality of the statistics they compile and disseminate” and that “quality information should include measures of accuracy”38. For research data integrated into ocean accounts, this translates to publishing confidence intervals, standard errors, or qualitative uncertainty assessments alongside the data values used in the accounts.

4Combining uncertainty components — methods note. Where an account value carries multiple uncertainty components (typically measurement uncertainty, spatial interpolation uncertainty, and model parameter uncertainty), the components must be combined into a single published uncertainty using a method that is both transparent and methodologically defensible. NSOs should apply the following:

  • 5Quadrature combination (root sum of squares) for uncertainty components that are statistically independent and uncorrelated. If $u_1, u_2, \ldots, u_n$ are the independent component uncertainties (expressed as standard uncertainties), the combined uncertainty is $u_c = \sqrt{u_1^2 + u_2^2 + \cdots + u_n^2}$. This is the method specified by the Joint Committee for Guides in Metrology (JCGM 100:2008 — Guide to the Expression of Uncertainty in Measurement, “GUM”), the international standard for measurement uncertainty propagation, and should be the default for ocean accounts.
  • 6Arithmetic addition ($u_c = u_1 + u_2 + \cdots + u_n$) as a conservative upper bound where the correlation structure between components is unknown or where components are believed to be positively correlated. This approach overstates uncertainty for independent components and should be clearly identified as a conservative bound when used.

7Structural model uncertainty (also called scenario uncertainty or model-form uncertainty) arises when the mathematical structure of a model may be wrong, independently of whether its parameter values are correct. As structural uncertainty has no standard probabilistic form, it cannot be combined in quadrature with measurement or parameter uncertainties. It must instead be reported as a separate qualitative statement or as a scenario-based range (e.g., low/central/high scenario outputs from alternative model structures). This separate reporting requirement applies whenever biophysical model outputs are used as primary inputs to ocean accounts. IPCC AR6 Working Group I uncertainty guidance and JCGM 100:2008 Annex notes on model-form uncertainty provide relevant precedent. The chosen combination method, the components combined, and the assumed correlation structure must be documented alongside the account release. Where uncertainty components have been combined arithmetically because correlation structure is unknown, future work to characterise the correlation structure and move to a GUM-aligned quadrature combination should be flagged in the account quality report.

3.4 Metadata Standards

1Complete metadata documentation is essential for integrating research data into statistical systems. Metadata enable data discovery, support quality assessment, and provide the documentation required for reproducible compilation of accounts.

3.4.1 FAIR principles for research data

1The FAIR Guiding Principles (Findable, Accessible, Interoperable, Reusable) provide the overarching data management framework that underpins research-data integration. Full definitions and practical requirements for each principle are provided in TG-4.6 Data Harmonisation and Interoperability Section 3.3. NSOs should prioritise research data sources that adhere to FAIR principles and advocate for FAIR practices when negotiating data sharing arrangements with research organisations, since non-FAIR research datasets impose additional metadata remediation work at Phase 2 of the compilation procedure (Section 3.5.2).39

3.4.2 Domain-specific metadata standards

1Several domain-specific metadata standards are relevant for ocean accounting, including ISO 19115 (geospatial metadata), Darwin Core (biodiversity occurrence), CF Conventions (gridded oceanographic data in NetCDF), SDMX (statistical data and metadata exchange), and the IHO S-100 Universal Hydrographic Data Model40. Practitioners drawing on biodiversity survey data, spatially-referenced ocean observations, or hydrographic products should adopt the relevant standard at the point of ingest so that metadata records remain interoperable with national and international statistical systems. For full treatment of these standards, the SDMX adoption pathway, and the harmonisation workflow that connects them, see TG-4.6 Data harmonisation.

3.4.3 Data provenance documentation

1Provenance documentation tracks the history of a dataset through its processing chain, so that users can understand how data have been transformed from raw observations to derived products41. For research data used in ocean accounting, provenance should document:

  • 2Original data sources and collection methods
  • 3Processing steps and algorithms applied
  • 4Software and version numbers used
  • 5Personnel responsible for processing
  • 6Dates of processing steps
  • 7Quality control procedures applied

8The SEEA Technical Guidance on Biophysical Modelling recommends maintaining “a data provenance system” that “improves users’ ability to understand the fitness for purpose of data sets”42.

3.5 Compilation Procedure

1This section outlines a systematic procedure for assessing, acquiring, and integrating research data into ocean accounting programmes. The procedure consists of four phases: research data assessment, metadata alignment, quality assurance, and account integration.

3.5.1 Phase 1: Research data assessment

1The first phase involves identifying candidate research data sources and conducting a preliminary fitness assessment:

2Step 1.1: Identify data requirements — Determine which components of the ocean accounts require research data inputs. This assessment should be guided by the account structure and the availability of alternative data sources. For example, if ecosystem condition accounts for pelagic waters are planned (see TG-6.5 Pelagic and Open Ocean Accounting), identify which condition variables (dissolved oxygen, chlorophyll-a, sea surface temperature) are available from research programmes versus traditional statistical sources.

3Step 1.2: Survey available research data — Conduct a systematic survey of research data sources within the accounting domain. This survey should cover:

  • 4National research institutions (marine laboratories, oceanographic institutes, fisheries research centres)
  • 5International research programmes (GOOS, Argo, regional ocean observing systems)
  • 6Global data repositories (OBIS, GBIF, Copernicus Marine Service)
  • 7Published scientific literature and associated datasets

8Step 1.3: Apply integration checklist — For each candidate data source, complete the integration checklist presented in Table 3.5.1. This checklist draws together the quality, metadata, and institutional considerations discussed in Sections 3.3 and 3.4. Each criterion is designated as either mandatory (must be met or formally waived before progression to Phase 2) or conditional (may remain open during Phase 2 or 3, subject to a documented resolution path).

9Table 3.5.1: Research Data Integration Checklist

Integration CriterionStatusAssessment QuestionsDocumentation Required
Spatial coverageMandatoryDoes it cover the accounting area?Geographic metadata (bounding box, coordinate system)
Temporal alignmentMandatoryDoes it match accounting periods?Date/time stamps, temporal resolution
Methodological consistencyMandatoryAre methods comparable to official statistics?Methods documentation, peer-reviewed publications
Institutional accessMandatoryCan the NSO access/use the data?Data sharing agreement, licensing terms
Quality assuranceConditionalWhat QA procedures were applied?Quality reports, validation studies
Classification alignmentConditionalAre categories mappable to SEEA/ISIC?Classification concordance or crosswalk
Metadata completenessConditionalAre ISO 19115/FAIR metadata available?Metadata catalogue entry
ReproducibilityConditionalCan results be reproduced from documented inputs?Code repository, processing documentation

10Phase 1 gate. A data source may proceed to Phase 2 only if all mandatory criteria are met or formally waived. A waiver may be issued only by a named authority within the NSO (typically the head of the responsible statistical unit, or, for Mode C primary substitutions under Section 3.1.1, the chief statistician). The waiver, its scope, its expiry, and the remediation pathway must be recorded in the account compilation file. Conditional criteria that remain open at the gate must be assigned a target resolution date within Phase 2 or Phase 3. Failure to resolve a conditional criterion within Phase 3 escalates it to a mandatory issue at the Phase 4 release decision.

3.5.2 Phase 2: Metadata alignment

1Once candidate data sources have been identified, the second phase addresses metadata harmonisation:

2Step 2.1: Extract existing metadata — Retrieve available metadata from the research data source. This may exist in ISO 19115 format (for geospatial data), Darwin Core format (for biodiversity observations), or NetCDF-CF format (for oceanographic model outputs).

3Step 2.2: Assess metadata completeness against a two-tier standard — Compare existing metadata against the requirements for statistical use. NSOs should apply a two-tier metadata standard.

4Statistical minimum (required to proceed to Phase 3 quality assurance). The dataset must have documented:

  • 5Temporal extent (reference period and, where relevant, temporal resolution and update frequency)
  • 6Spatial coverage (geographic extent and coordinate reference system)
  • 7Measurement method (instrument, protocol, or modelling approach)
  • 8Uncertainty estimate (quantitative where available, or qualitative otherwise)
  • 9Access terms (licence, citation requirements, restrictions)

10Where formal metadata records (e.g., ISO 19115, Darwin Core) are absent, as is common for high-quality legacy datasets and many older oceanographic programmes, the statistical minimum may be satisfied using metadata extracted from a peer-reviewed publication that documents the dataset. The compiling NSO must record the source of each minimum element, including the publication citation and section reference where applicable.

11Full standard (long-term target for repeated and high-profile sources). The dataset is described by a fully-populated ISO 19115 record (for geospatial data) or Darwin Core record (for biodiversity data), including all of the elements identified by the UN NQAF Level D essential metadata list: identification (title, abstract, keywords), temporal extent (reference period, temporal resolution, update frequency), spatial extent (geographic coverage, coordinate reference system), data quality (accuracy, completeness, consistency), lineage (data sources, processing steps), distribution (access constraints, usage licences), and contact information (data custodian, responsible party)43.

12A dataset meeting the statistical minimum may progress to Phase 3 whilst work to reach the full standard is undertaken in parallel.

13Step 2.3: Fill metadata gaps — Where research data lack statistical metadata elements, work with the data provider to document missing information. Priority gaps include: correspondence to statistical classifications (e.g., mapping research ecosystem types to IUCN GET categories), uncertainty quantification (confidence intervals, accuracy assessments), and update schedules (will data be available on a recurring basis to support time series accounts?).

14Step 2.4: Document provenance — Create or enhance provenance documentation following the framework in Section 3.4.3. For research data that have undergone multiple processing steps (e.g., satellite imagery processed to ocean colour products, acoustic survey data processed to biomass estimates), the provenance chain must be fully documented to support reproducibility.

3.5.3 Phase 3: Quality assurance

1The third phase applies the quality assessment framework from Section 3.3:

2Step 3.1: Assess relevance (as defined in Section 3.3.1) — Verify that the research data address the specific accounting requirements identified in Phase 1, considering both conceptual fit (do the measured variables correspond to the accounting concepts?) and practical fit (are the data sufficiently timely, granular, and complete?).

3Step 3.2: Evaluate accuracy (as defined in Section 3.3.1) — Assess accuracy using available validation studies, inter-comparison exercises, or ground-truthing campaigns. For satellite-derived products, consult published accuracy assessments. For modelled outputs, assess model skill metrics against independent observations.

4Step 3.3: Test coherence (as defined in Section 3.3.1) — Verify that research data can be combined with other account data sources. Coherence testing should identify discrepancies in spatial boundaries, temporal reference periods, or measurement units that require harmonisation (see TG-4.6 Data Harmonisation).

5Where co-located independent monitoring is available (e.g., research station data co-located with an environmental protection agency monitoring station), coherence is tested by direct comparison of measurements at common locations and times. Where co-located data are unavailable, which is the typical situation for pelagic, open-ocean, and deep-shelf accounts, the following fallback coherence testing hierarchy applies. At least one fallback test must be completed and documented:

  1. 6Temporal climatology test. Compare research data values to climatological norms from global ocean databases (e.g., NOAA World Ocean Atlas, Copernicus Marine Service reanalyses) for the matching season, depth, and biogeographic region. Values falling outside the 5th-95th percentile range of the climatology should be reviewed for outlier status and either corrected, flagged, or accepted with documented justification.
  2. 7Cross-variable coherence test. Verify that correlated variables observed in the same dataset exhibit expected relationships (e.g., dissolved oxygen and temperature follow expected solubility relationships, and chlorophyll-a and nitrate exhibit expected biogeochemical coupling). Departures from expected relationships indicate either real ecosystem signal or measurement issue and must be examined.
  3. 8Literature-based plausibility check. Compare summary statistics (mean, range, seasonal cycle) to published ranges for equivalent ecosystem types in the peer-reviewed literature. Values outside the published range should be reviewed for plausibility.

9These fallback tests are not substitutes for direct co-location where it is feasible, but they ensure that the coherence check does not silently fail in the open-ocean and deep-water contexts where research data are most essential.

10Step 3.4: Check comparability (as defined in Section 3.3.1) — Document any methodological changes that create breaks in time-series comparability, and verify consistent methods across geographic domains.

11Step 3.5: Quantify uncertainty — Where feasible, quantify the uncertainty associated with research data values. This may take the form of standard errors (for survey-based estimates), confidence intervals (for modelled values), or qualitative uncertainty categories (high/medium/low confidence). Uncertainty estimates should be documented in metadata and, where appropriate, published alongside account values. Where multiple uncertainty components must be combined into a single published figure, follow the methods note in Section 3.3.3.

3.5.4 Phase 4: Account integration

1The final phase integrates quality-assured research data into ocean accounts:

2Step 4.1: Apply classification concordances — Where research data use different classifications from statistical standards, apply the concordances or crosswalks developed in Phase 2. For example, if research biodiversity data use scientific taxonomic names but the account structure requires aggregation to functional groups, apply the taxonomic-to-functional-group mapping.

3Step 4.2: Reconcile spatial and temporal boundaries — Align research data to the spatial and temporal structure of the accounts. This may require spatial aggregation (from fine-resolution survey points to accounting spatial units), temporal aggregation (from monthly observations to annual accounting periods), or gap-filling (interpolating missing values).

4Spatial interpolation method selection. The choice of spatial interpolation method must be justified by the density and spatial distribution of the source data, not selected by default. Table 3.5.4 specifies the recommended method as a function of station density and spatial pattern. The choice and its justification must be recorded in the account compilation file.

5Table 3.5.4: Recommended spatial interpolation method by station density and distribution

Station density / distributionRecommended methodNotes
Dense (>30 stations over the accounting area) with approximately random or stratified spatial distributionOrdinary kriging (or universal kriging where trend is present)Variogram should be fitted and inspected; cross-validation mandatory
Moderate density (10-30 stations) with clustered or coastal-biased distributionInverse distance weighting (IDW) with documented search radiusSensitive to clustering; report edge-effect zones as higher-uncertainty
Sparse coverage (<10 stations) or strongly heterogeneous environmentThiessen (Voronoi) polygon assignment, or assignment by ecosystem stratum meanResulting account values must be flagged as Tier 1 (broad indicative) under SEEA EA tiering
Sparse and gridded reference data availableRegression-based methods (e.g., regression-kriging) using auxiliary covariates (bathymetry, SST)Requires documented covariate justification

6Cross-validation is mandatory whichever interpolation method is selected. Leave-one-out cross-validation should produce a root-mean-square error no greater than the tolerance threshold, defined as the published measurement uncertainty of the source data multiplied by 1.5 (i.e., tolerance threshold = baseline RMSE × 1.5, a multiplicative allowance of 50% above the measurement uncertainty). Values exceeding this threshold indicate that interpolation uncertainty dominates and that the interpolated product must either be downgraded in tier, supplemented with additional stations, or replaced with a coarser spatial aggregation.

7Step 4.3: Document data sources — Record the use of research data in the account compilation metadata. Documentation should identify: the research data source (with citation and persistent identifier), the account components that use the research data, the processing steps applied, and the quality assessment results. This documentation supports transparency and reproducibility.

8Step 4.4: Establish update procedures — Where research data will be used on a recurring basis for time series accounts, establish procedures for data updates. Coordinate with research data providers to understand their publication schedule and arrange for regular data transfers. Monitor for methodological changes that may affect comparability across accounting periods.

9Back-revision policy for retroactively corrected research data. Research datasets are frequently revised: stock assessments are updated as new cohorts enter the record, satellite data products undergo reprocessing, and bathymetric and hydrographic surveys are revised as better data become available. A retrospective revision to historical research data can propagate through multiple accounting periods of a time series account, creating apparent trends that are artefacts of data revision rather than real ecosystem change. NSOs must adopt a back-revision policy with the following elements:

  • 10Monitoring. NSOs should subscribe to dataset versioning notices from research data providers and incorporate a versioning check into the standing data transfer protocol (Section 3.7.4).
  • 11Materiality threshold. A retrospective revision triggers formal revision of previously published accounts where it produces a change of more than 5% in any key account indicator for any covered accounting period, or where it materially changes the direction of a published trend. Below the materiality threshold, the revision is recorded in the next account release without re-publication of historical accounts.
  • 12Version identification. Every account release must identify the specific version (DOI or equivalent persistent identifier) of each research dataset used. This allows users to trace account values to specific dataset versions and supports reconstruction of historical account values.
  • 13Metadata revision history. The account compilation metadata must include a revision history listing the date, dataset, version change, materiality assessment, and the action taken (re-published, recorded in next release, or no action).

14This policy aligns with SEEA EA guidance on revision policy in statistical frameworks and with UN NQAF Level B revision policy requirements.

3.6 Worked Example: Integrating Oceanographic Survey Data into Condition Accounts

1This worked example demonstrates the application of the compilation procedure to a realistic scenario: a National Statistical Office seeking to compile ecosystem condition accounts for coastal shelf waters using dissolved oxygen data from a national oceanographic research programme.

2Setting: A coastal nation with 150,000 km² of exclusive economic zone (EEZ) shelf waters (depths <200m) seeks to compile annual condition accounts for the Marine Shelf (M1) ecosystem type following the IUCN Global Ecosystem Typology. One of the selected condition variables is dissolved oxygen concentration, which serves as an indicator of ecosystem health and hypoxia risk. The NSO has identified the National Oceanographic Research Institute (NORI) as a potential data provider.

Note on illustrative scope. This worked example uses a single condition variable (dissolved oxygen) and a linear-rescaling condition index purely for pedagogical clarity. SEEA EA Table 5.3 envisages multi-variable condition indices that combine abiotic chemical, abiotic physical, biotic compositional, structural, and functional characteristics, and a production-grade Marine Shelf (M1) condition account would aggregate dissolved oxygen with at least temperature, chlorophyll-a, and a biotic indicator. The dissolved oxygen reference (8.0 mg/L) and minimum-threshold (2.0 mg/L hypoxia) values used below are drawn from the OSPAR Ecological Quality Objectives for dissolved oxygen. National standards (e.g., EU Water Framework Directive good ecological status thresholds, regional sea convention values) may differ and should be used where authoritative locally.

3Phase 1: Research data assessment

4Step 1.1: Identify data requirements — The condition account requires dissolved oxygen measurements representative of the shelf ecosystem. Following SEEA EA guidance, the reference condition is defined as the dissolved oxygen level corresponding to a healthy, well-mixed shelf ecosystem (typically 6-8 mg/L). The account structure requires annual average values aggregated to ecosystem asset spatial units. Applying the Section 3.1.1 integration-mode typology, the NSO classifies this as Mode A (gap-filling): no official statistical source for subsurface dissolved oxygen exists in this jurisdiction.

5Step 1.2: Survey available data — NORI conducts quarterly oceanographic surveys at 45 fixed stations distributed across the shelf. Each station is sampled at 5 depth intervals (surface, 25m, 50m, 75m, 100m). Dissolved oxygen is measured using calibrated Winkler titration (precision ±0.1 mg/L). The programme has operated continuously since 2010 with consistent methodology. Data are archived in NORI’s institutional repository.

6Step 1.3: Apply integration checklist — Applying Table 3.5.1 with mandatory/conditional designations:

CriterionStatusAssessment ResultDocumentation
Spatial coverageMandatory45 stations cover 150,000 km² shelf area; spatial interpolation requiredStation coordinates in WGS84
Temporal alignmentMandatoryQuarterly surveys provide seasonal coverage; annual averaging feasibleSurvey dates documented per cruise
Methodological consistencyMandatoryWinkler titration is standard oceanographic methodNORI Standard Operating Procedures manual
Institutional accessMandatoryNORI willing to share data; MoU requiredDraft MoU provided by NORI legal office
Quality assuranceConditionalInter-laboratory bias <0.15 mg/L vs. IOC/GOOS ±0.2 mg/L standardQA reports 2012-2020
Classification alignmentConditionalDissolved oxygen is a standard SEEA EA condition variableSEEA EA Table 5.3
Metadata completenessConditionalStation metadata exist; cruise-level metadata incompleteISO 19115 records for stations
ReproducibilityConditionalRaw titration data archived; processing code not version-controlledNORI agrees to deposit in GitHub

7Phase 1 gate finding: All four mandatory criteria are met (subject to MoU finalisation, which the NSO head of environmental statistics waives for a 90-day remediation window). Conditional criteria are open but on a documented resolution path. The source proceeds to Phase 2.

8Phase 2: Metadata alignment

9Step 2.1: Extract existing metadata — NORI provides station-level metadata in CSV format including: station ID, latitude, longitude, depth, seafloor substrate type, and sampling history. Dissolved oxygen data are provided in a separate CSV with fields: station ID, cruise ID, date, depth, dissolved oxygen (mg/L), temperature (°C), salinity (PSU).

10Step 2.2: Assess metadata completeness against the two-tier standard (Section 3.5.2) — Existing metadata are evaluated against the statistical minimum: temporal extent (partially documented, with gaps in cruise-level dates), spatial coverage and CRS (CRS implicit, must be made explicit), measurement method (documented), uncertainty estimate (available from QA reports), access terms (under MoU). The dataset meets the statistical minimum once CRS is explicitly documented and cruise dates are filled in. The full ISO 19115 standard is set as a 12-month target.

11Step 2.3: Fill metadata gaps — NSO and NORI jointly develop enhanced metadata including:

  • 12Temporal coverage: Date range for each quarterly cruise added to cruise metadata table
  • 13Spatial reference: Coordinate reference system explicitly documented as WGS84 (EPSG:4326)
  • 14Quality flags: NORI applies automated QC checks (range test, climatology test, spike test) following GOOS recommendations and adds QC flags to data file (1=good, 2=probably good, 3=probably bad, 4=bad)
  • 15Provenance: Processing workflow documented: raw titration volume → dissolved oxygen calculation using modified Winkler equation → temperature and salinity correction → final value in mg/L

16Step 2.4: Document provenance — Provenance record created:

Dissolved oxygen values are measured using Winkler titration following the GOOS BioEco Panel recommendations. Seawater samples are collected using Niskin bottles mounted on a CTD rosette. Titration is performed shipboard within 6 hours of collection. Raw titration volumes are converted to dissolved oxygen concentration using the modified Winkler equation with temperature and salinity corrections applied. Processing code (Python) is version-controlled at https://github.com/NORI/oceanography/DO-processing [fictional URL] (Reproducibility Level 2 per Section 3.3.2). Quality control follows GOOS Real-Time Quality Control procedures (GOOS, 2021).

17Phase 3: Quality assurance

18Step 3.1: Assess relevance — The dissolved oxygen data directly address the condition account requirement for chemical state characteristics. Quarterly temporal resolution provides adequate seasonal coverage for annual aggregation. Spatial coverage (45 stations across 150,000 km²) is sparser than ideal but sufficient for broad-scale condition assessment given the relatively homogeneous shelf environment.

19Step 3.2: Evaluate accuracy — NORI’s QA reports indicate inter-laboratory comparison results within ±0.15 mg/L of reference standards, satisfying the IOC/GOOS ±0.2 mg/L accuracy threshold referenced in Table 3.3.1. Sampling precision (replicate measurements at same station) averages ±0.08 mg/L. The accuracy dimension is rated “Acceptable without qualification”.

20Step 3.3: Test coherence — NSO compares dissolved oxygen data against coastal water quality monitoring data from the environmental protection agency at 12 co-located stations. Average difference is 0.12 mg/L (within measurement uncertainty), confirming coherence. For shelf-edge stations where no co-located independent monitoring exists, the NSO additionally applies the temporal climatology test against NOAA World Ocean Atlas seasonal climatology: all 45 stations’ annual means fall within the 5th-95th percentile range of climatological values for the matching region and depth. This confirms the absence of systematic bias.

21Step 3.4: Check comparability — NORI methodology has remained unchanged since programme inception (2010). All data are directly comparable over time. Spatial comparability verified by consistent station locations and sampling protocols.

22Step 3.5: Quantify uncertainty — Based on QA assessment, dissolved oxygen values carry measurement uncertainty of ±0.15 mg/L (combining measurement precision and inter-laboratory comparison). This is combined in quadrature with spatial interpolation uncertainty (Step 4.2 cross-validation) per the JCGM 100:2008 (GUM) procedure described in Section 3.3.3. Combined uncertainty $u_c = \sqrt{(0.15)^2 + u_{spatial}^2}$ varies by asset unit depending on distance to the nearest station.

23Phase 4: Account integration

24Step 4.1: Apply classification concordances — No classification mapping is required, as dissolved oxygen is used directly.

25Step 4.2: Reconcile spatial and temporal boundaries — The accounting area is divided into 500 ecosystem asset spatial units (300 km² each) based on seabed substrate type. Applying the Table 3.5.4 method-selection rubric: with 45 stations over 150,000 km² in an approximately stratified spatial distribution, the rubric points to ordinary kriging. However, because variogram fitting at this station density is unstable, the NSO uses inverse distance weighting with a 30 km search radius as a pragmatic alternative and documents the reasoning. Mandatory leave-one-out cross-validation yields RMSE = 0.22 mg/L, within the 50% tolerance over the measurement uncertainty (0.15 × 1.5 = 0.225 mg/L), so the interpolated product is accepted. Quarterly values are averaged to produce annual mean dissolved oxygen per asset unit.

26Step 4.3: Document data sources — Account compilation metadata records:

Dissolved oxygen condition data sourced from National Oceanographic Research Institute Quarterly Shelf Survey (NORI-QSS), 2015-2020, dataset version 2021.1 (DOI: 10.xxxx/nori-qss-2021.1) [fictional]. Data access via MoU between NSO and NORI dated 2021-03-15. Data processing conducted by NSO Environmental Accounts Unit using scripts deposited at https://github.com/NSO/ocean-accounts/condition-processing [fictional URL]. Integration mode: A (gap-filling). Spatial interpolation: inverse distance weighting, 30 km search radius (justified per Table 3.5.4, cross-validation RMSE 0.22 mg/L). Temporal aggregation: arithmetic mean of quarterly values. Uncertainty combined in quadrature per JCGM 100:2008: ±0.15 mg/L measurement plus variable spatial interpolation uncertainty.

27Step 4.4: Establish update procedures — NSO and NORI agree that NORI will provide annual data extracts by 31 March each year (covering the previous calendar year). NSO will re-run spatial interpolation and update condition accounts by 30 June. NORI will notify NSO of any methodological changes at least 6 months prior to implementation. Under the back-revision policy applied here, NORI dataset versioning notices are subscribed to. Any retrospective revision producing >5% change in the asset-unit condition index for any previously published year triggers re-publication of affected accounts. Each account release identifies the NORI dataset DOI.

28Resulting condition account entry (example for one asset unit):

Accounting yearDissolved oxygen (mg/L)Indicator valueUncertainty
20156.80.80 (good condition)±0.35 mg/L
20166.50.75 (good condition)±0.33 mg/L
20175.90.65 (moderate condition)±0.38 mg/L
20186.20.70 (moderate condition)±0.36 mg/L
20195.70.62 (moderate condition)±0.40 mg/L
20205.40.57 (moderate condition)±0.42 mg/L

29Note: Condition is classified as “good” where the indicator value is 0.75 or above, and “moderate” where it falls below 0.75. Uncertainty values vary across years, as some monitoring stations had intermittent data gaps in later years. The gaps increased spatial interpolation distance and, with it, the combined-in-quadrature uncertainty for this asset unit.

30Outcome: The NSO has integrated research data from NORI into the ecosystem condition account. The dissolved oxygen time series indicates a declining trend from good to moderate condition over 2015-2020. This decline has prompted policy attention to potential eutrophication drivers. The documented uncertainty estimates and provenance information support transparent communication of account results and reproducible updates in future accounting periods.

3.7 Institutional Arrangements

1Effective integration of research data into official statistics requires institutional arrangements that bridge the different cultures, incentives, and practices of statistical offices and research organisations.

3.7.1 Roles of National Statistical Offices

1The SEEA Technical Recommendations identify several roles that NSOs can play in ecosystem accounting that are relevant to research data integration. Table 3.7.1 below summarises these roles.44

RoleDescription
Data organisationNSOs have expertise in collecting and organising data from diverse sources, building coherent pictures from varied inputs.
Standards stewardshipNSOs establish and maintain definitions, concepts, and classifications, addressing the multiple definitions common in research contexts.
Data integrationNSOs integrate data from various sources within national and international statistical frameworks.
Quality frameworksNSOs apply data quality frameworks enabling consistent assessment and accreditation of information sources.
National coverageNSOs create national pictures, applying techniques for scaling information to national level.
AuthorityNSOs present an authoritative voice through application of standard measurement approaches and quality frameworks.

2Whilst NSOs may not have deep expertise in marine science, they bring essential capabilities for transforming research data into official statistics45.

3.7.2 Roles of research institutions

1Research institutions contribute domain expertise, data collection infrastructure, and methodological innovation. The SEEA Technical Recommendations note that “agencies that lead work on geographic and spatial data — particularly the mapping of environmental data and the use of remote sensing information — including for spatial and temporal modelling of ecosystem services” play important roles46. For ocean accounting, relevant research institutions include:

  • 2Universities with marine science programmes
  • 3National oceanographic and hydrographic agencies
  • 4Fisheries research institutes
  • 5Environmental monitoring agencies
  • 6International research programmes (e.g., IOC, ICES)

7These institutions are often the primary custodians of ocean observation data, biodiversity records, and ecosystem assessments needed for ocean accounts.

3.7.3 Establishing partnerships

1The Global Statistical Geospatial Framework (GSGF) provides general guidance on collaboration between statistical offices and geospatial agencies47, but agreements with research institutions need to address dimensions that NSO-geospatial agency templates (such as the UN-GGIM MoU template) do not cover. Research institutions operating under university or grant-funded regimes have data governance requirements (publication embargo periods, intellectual property ownership, attribution under Creative Commons or funder mandates) that differ materially from those of national mapping agencies. NSOs negotiating data sharing arrangements with research partners should therefore use a purpose-built research data sharing agreement rather than carrying over the UN-GGIM template unchanged. The OECD Principles and Guidelines for Access to Research Data from Public Funding (2007) provide the governing international framework.

2Research data sharing agreement checklist. A research data sharing agreement should address the following at minimum:

  • 3Intellectual property ownership. Identify the rights holders of the data (institution, individual researcher, funder, or joint), confirm that the NSO has the rights necessary to ingest, transform, and republish the data in statistical products, and record any conditions imposed by funders (e.g., European Research Council, NIH, Wellcome, NSF).
  • 4Publication embargo periods. Where the research team needs to publish primary findings before the data are released, specify the embargo duration. The maximum acceptable embargo for ocean-accounting purposes is 6 months from the end of the data collection period, consistent with the annual account publication cycle: data collected by 31 December must be available to the NSO by 30 June of the following year to meet standard account publication deadlines (see Step 4.4). Shorter embargoes should be sought wherever possible. Embargoes beyond 6 months are not compatible with the annual publication cycle of official accounts and should not be accepted. Where a research partner requires a longer embargo as a condition of data sharing, this should be escalated to the chief statistician as an exceptional governance case. Note that many major international funder mandates (e.g., Horizon Europe, NSF) require data availability within 12 months of collection. This funder deadline does not override the NSO’s 6-month operational requirement.
  • 5Attribution requirements. Specify the citation form required by the data provider (data DOI plus accompanying publication) and confirm that the NSO will discharge these in account release documentation.
  • 6Data lifecycle. Address versioning (how new versions will be announced and made available), deprecation (how withdrawn versions will be handled), and the back-revision interaction (Section 3.5.4 Step 4.4).
  • 7Confidentiality of pre-publication data. Where the NSO is provided access to pre-publication data, specify the handling regime (named individuals, secured storage, no onward sharing) and the point at which the data may be released into the NSO’s normal compilation environment.

8Communities of practice. Regular engagement through working groups or committees maintains relationships and addresses emerging issues. The SEEA Technical Recommendations emphasise that “appropriate institutional arrangements and resourcing to support ongoing engagement and communication are also required”48.

9Capacity building. Joint training and skill-sharing activities build mutual understanding between statistical and research communities. Research personnel may need orientation on statistical concepts and quality frameworks, whilst statistical personnel may need training on oceanographic data and methods.

10Several countries provide practical models for NSO-research institution partnerships in ocean accounting. In Australia, the ABS compiles environmental-economic accounts in partnership with DCCEEW, drawing on scientific data inputs from CSIRO and IMOS.49 In the Netherlands, Statistics Netherlands (CBS) has worked with Wageningen University and NIOZ to compile experimental natural capital accounts for the North Sea using research-derived biophysical models and real-time sensor networks.49 Successful partnerships require sustained engagement over multiple accounting cycles, clear allocation of responsibilities, and mutual recognition of complementary capabilities.

3.7.4 Data transfer protocols

1The SEEA Technical Note on Air Emission Accounts recommends establishing “data transfer protocols” given that “data may be acquired from a number of institutions or agencies”50. Such protocols should address:

  • 2Data formats and transmission methods
  • 3Timing and frequency of data provision
  • 4Procedures for handling system changes and upgrades
  • 5Metadata to be provided with each data transfer
  • 6Dataset versioning notices (to support the Section 3.5.4 back-revision policy)
  • 7Feedback mechanisms for data quality issues

8Reliable protocols prevent disruption to statistical production when research systems are upgraded or personnel change.

3.7.5 Guidance for low-capacity and SIDS contexts

1The institutional arrangements set out in Sections 3.7.1-3.7.4 assume that one or more national research institutions exist with which the NSO can establish a formal partnership. For many small island developing states (SIDS), least developed countries, and other contexts where national research capacity is still developing, ocean accounts must be compiled using data sourced entirely from international repositories and global research programmes. This is the dominant scenario for a substantial part of the GOAP target audience. The absence of a domestic research institution partner should not be treated as a barrier to compiling ocean accounts.

2The following adapted procedure applies where no national research institution partner exists:

  • 3Use pre-validated global data products as research data inputs without a national MoU. Acceptable sources include the Copernicus Marine Service (validated ocean colour, sea surface temperature, sea level, and biogeochemistry products), GOOS Argo (subsurface temperature, salinity, dissolved oxygen), and OBIS / GBIF (biodiversity occurrence). These products are issued under open licences that permit statistical use without bespoke agreements. The product DOI or persistent identifier serves the role that the MoU would otherwise play in establishing access provenance.
  • 4Base quality assurance on the product’s published quality reports. In place of a bespoke Section 3.5.3 quality assessment, the NSO may rely on the product quality factsheet, validation report, or technical note published by the data provider, provided this is referenced in the account compilation metadata together with the product version. The Section 3.5.3 Step 3.3 fallback coherence tests still apply.
  • 5Engage regional bodies as institutional partners in lieu of national research institutes. Examples include the Secretariat of the Pacific Regional Environment Programme (SPREP) and the Pacific Community (SPC) for the Pacific Islands, the Indian Ocean Global Ocean Observing System (IOGOOS), the Caribbean Community Climate Change Centre (5Cs), and the Intergovernmental Oceanographic Commission Sub-Commission for Africa and the Adjacent Island States (IOCAFRICA). Regional bodies can provide brokered access to data, technical support, and a community of practice that substitutes for a national institutional partner.
  • 6Simplified metadata and provenance documentation. Pre-validated global products carry mature metadata, and the NSO’s documentation obligation is reduced to citing the product, version, access date, and the elements of the product’s quality report relied upon in the account.

7NSOs operating in low-capacity contexts should not be deterred by the formal arrangements set out earlier in Section 3.7: those arrangements describe the upper end of the institutional spectrum, whilst this sub-section describes the lower end at which ocean accounting is equally feasible.

4. Acknowledgements

1This Circular has been approved for public circulation and comment by the GOAP Technical Experts Group in accordance with the Circular Publication Procedure.

2Authors: [To be confirmed]

3Reviewers: [To be confirmed]

5. References

Footnotes

  1. 1

    SEEA EA, para. 5.14-5.30. Ecosystem condition accounts record “the quality of an ecosystem” through biophysical and chemical characteristics.

  2. 2

    Intergovernmental Oceanographic Commission. (2019). The Global Ocean Observing System 2030 Strategy. Paris: UNESCO-IOC. Available from: https://www.goosocean.org/

  3. 3

    Simmonds, J., & MacLennan, D.N. (2005). Fisheries Acoustics: Theory and Practice, 2nd ed. Oxford: Blackwell Science. Acoustic surveys provide fishery-independent biomass estimates used in stock assessment.

  4. 4

    SEEA EA, Table 5.3. The ecosystem condition typology identifies abiotic (physical state, chemical state), biotic (compositional, structural, functional), and landscape characteristics.

  5. 5

    Hilborn, R., & Walters, C.J. (1992). Quantitative Fisheries Stock Assessment: Choice, Dynamics and Uncertainty. New York: Chapman and Hall. Stock assessment integrates fishery-dependent catch data with fishery-independent survey data.

  6. 6

    Bravington, M.V., Skaug, H.J., and Anderson, E.C. (2016). “Close-kin mark-recapture.” Statistical Science, 31(2), 259-274. https://doi.org/10.1214/16-STS552. Applied to southern bluefin tuna by Hillary, R.M. et al. (2018), Scientific Reports, 8, 13767.

  7. 7

    Data-poor stock assessment methods referenced in Section 3.1.2: Hordyk, A., Ono, K., Valencia, S., Loneragan, N., & Prince, J. (2015). A novel length-based empirical estimation method of spawning potential ratio (SPR), and tests of its performance, for small-scale, data-poor fisheries. ICES Journal of Marine Science, 72(1), 217—228. doi:10.1093/icesjms/fsu004; Froese, R., Demirel, N., Coro, G., Kleisner, K. M., & Winker, H. (2017). Estimating fisheries reference points from catch and resilience. Fish and Fisheries, 18(3), 506—526. doi:10.1111/faf.12190; Dick, E. J., & MacCall, A. D. (2011). Depletion-Based Stock Reduction Analysis: a catch-based method for determining sustainable yields for data-poor fish stocks. Fisheries Research, 110(2), 331—341. doi:10.1016/j.fishres.2011.05.007. For a contemporary FAO synthesis of data-poor methods, see: Punt, A. E., Butterworth, D. S., de Moor, C. L., De Oliveira, J. A. A., & Haddon, M. (2014). Stock Assessment Methods Used by National and Regional Fisheries Management Organizations (FAO Fisheries and Aquaculture Technical Paper No. 569). FAO. fao.org/3/i3953e. SEEA EA para. 12.15 supports tiered reporting of asset accounts.

  8. 8

    OBIS. (2025). Ocean Biodiversity Information System. Available from: https://obis.org/ — GBIF. (2025). Global Biodiversity Information Facility. Available from: https://www.gbif.org/

  9. 9

    FDES 2013, para. 1.32-1.33. Scientific research and special projects “can be used to address data gaps” but “often use terms and definitions that differ from those used in statistics.”

  10. 10

    Intergovernmental Oceanographic Commission. (2019). The Global Ocean Observing System 2030 Strategy. Paris: UNESCO-IOC.

  11. 11

    Roemmich, D., et al. (2019). “On the future of Argo: A global, full-depth, multi-disciplinary array.” Frontiers in Marine Science, 6, 439. https://doi.org/10.3389/fmars.2019.00439

  12. 12

    SDG Framework. SDG Target 14.a: “Increase scientific knowledge, develop research capacity and transfer marine technology, taking into account the Intergovernmental Oceanographic Commission Criteria and Guidelines on the Transfer of Marine Technology.”

  13. 13

    GOOS. (2023). GOOS Essential Ocean Variables. Available from: https://www.goosocean.org/eov — Observation coverage is denser in developed country waters and coastal zones.

  14. 14

    GOOS. (2023). GOOS Essential Ocean Variables. Available from: https://www.goosocean.org/eov

  15. 15

    IUCN. (2020). IUCN Global Ecosystem Typology 2.0: Descriptive profiles for biomes and ecosystem functional groups. Gland: IUCN. https://doi.org/10.2305/IUCN.CH.2020.13.en

  16. 16

    Thomsen, P.F., & Willerslev, E. (2015). “Environmental DNA - An emerging tool in conservation for monitoring past and present biodiversity.” Biological Conservation, 183, 4-18.

  17. 17

    GBIF. (2025). Global Biodiversity Information Facility. Available from: https://www.gbif.org/

  18. 18

    OBIS. (2025). Ocean Biodiversity Information System. Available from: https://obis.org/

  19. 19

    SEEA Technical Recommendations, para. 1.34. Fully spatial approaches “will generally be more resource intensive and implementation will require more ecological and geo-spatial expertise.”

  20. 20

    GCRMN. (2020). Status of Coral Reefs of the World: 2020. Available from: https://gcrmn.net/

  21. 21

    SEEA EA, para. 12.15 on tiered approaches to measurement.

  22. 22

    SEEA Biophysical Modelling, para. 370. “Remote sensing data and modelling approaches provides enormous opportunities to disseminate data with very short time-lags and high-frequency.”

  23. 23

    Copernicus Marine Service. (2025). Available from: https://marine.copernicus.eu/

  24. 24

    SEEA Biophysical Modelling, para. 370.

  25. 25

    SEEA Biophysical Modelling, para. 370. Remote sensing provides “very short time-lags and high-frequency” data but requires validation against in situ measurements.

  26. 26

    SEEA Biophysical Modelling, para. 363. “The accuracy of modelled data can be assessed, although different approaches may be needed depending on the type of model used.”

  27. 27

    FDES 2013, para. 1.32. Scientific research data “are usually available at no or low cost.”

  28. 28

    FDES 2013, para. 1.33. Research data “often use terms and definitions that differ from those used in statistics”, have “limited scope”, and are “often available on a one-time basis only.”

  29. 29

    United Nations. (2019). United Nations National Quality Assurance Frameworks Manual for Official Statistics. New York: United Nations Statistics Division. Available from: https://unstats.un.org/unsd/methodology/dataquality/un-nqaf-manual/

  30. 30

    SEEA Biophysical Modelling, paras. 363-369 on model validation approaches.

  31. 31

    SEEA Biophysical Modelling, para. 373. “Modelling approaches have been rapidly improving…This creates challenges in including data produced from biophysical models into accounts.”

  32. 32

    National Academies of Sciences, Engineering, and Medicine. (2019). Reproducibility and Replicability in Science. Washington, DC: The National Academies Press. https://doi.org/10.17226/25303

  33. 33

    National Academies of Sciences, Engineering, and Medicine. (2019). Reproducibility and Replicability in Science. Washington, DC: The National Academies Press. Ch. 3 distinguishes reproducibility (same data, same methods) from replicability (new data, same or similar methods).

  34. 34

    SEEA Biophysical Modelling, para. 378. Transparency of workflow and uncertainty quantification are essential for reproducibility.

  35. 35

    SEEA Biophysical Modelling, para. 360. Uncertainty matrices outline possible sources of uncertainty for each model.

  36. 36

    SEEA Biophysical Modelling, para. 369. Model outputs should be seen as best estimates unless detailed validation has been conducted.

  37. 37

    SEEA Biophysical Modelling, para. 379. “Tier 1 and Tier 2 approaches may be best for awareness raising or analysis of broad spatiotemporal trends.”

  38. 38

    UN NQAF Manual, recommendation on quality reporting. Statistical agencies should publish quality information including accuracy measures.

  39. 39

    Wilkinson et al. (2016). “The FAIR Guiding Principles.”

  40. 40

    Relevant domain-specific metadata standards include ISO 19115-1:2014 Geographic information — Metadata (geospatial); Darwin Core (biodiversity occurrence, used by GBIF and OBIS); the CF Conventions for climate and forecast data in NetCDF; SDMX 3.0 (Statistical Data and Metadata Exchange); and the IHO S-100 Universal Hydrographic Data Model (covering bathymetry S-102, surface currents S-111, and water level S-104). See TG-4.6 Data harmonisation for the full treatment, classification crosswalks, and SDMX adoption pathway.

  41. 41

    SEEA Biophysical Modelling, para. 374 on data provenance systems.

  42. 42

    SEEA Biophysical Modelling, para. 374. Data provenance systems “improve users’ ability to understand the fitness for purpose of data sets.”

  43. 43

    UN NQAF Manual, Level D (Managing statistical outputs). Essential metadata elements for statistical dissemination.

  44. 44

    SEEA Technical Recommendations, Box 1.2. Potential roles of National Statistical Offices in ecosystem accounting.

  45. 45

    SEEA Technical Recommendations, paras. 1.60-1.61 on roles of NSOs and non-NSO agencies.

  46. 46

    SEEA Technical Recommendations, para. 1.61. Agencies with geospatial and remote sensing expertise play important roles.

  47. 47

    GSGF v2, Section on Principle 1. “Establishing strong communication and institutional collaboration mechanisms between NSOs and NGIAs is essential. This can be facilitated by, for example, country-level laws and policies, Memorandum of Understandings (MoUs), data sharing agreements, and other communities of practice.” See also OECD (2007), Principles and Guidelines for Access to Research Data from Public Funding, Paris: OECD, as the governing framework for research-data sharing agreements.

  48. 48

    SEEA Technical Recommendations, para. 1.57. “Given the need for involving many areas of expertise, an important aspect of implementation is the allocation of resources to co-ordination, data sharing and communication.”

  49. 49

    SEEA Technical Recommendations, paras. 1.55-1.62 on institutional arrangements for ecosystem accounting, including examples of multi-agency collaboration models. 2

  50. 50

    SEEA Technical Note: Air Emission Accounts, para. 81. “Given that data may be acquired from a number of institutions or agencies, it is important to establish data transfer protocols.”

Comment on this passage

Your comment may be published in the consultation record. Your name, email and organisation will not be.

Passage

CircularTG-4.5

Section

Paragraph

Quoted text

About you
Your comment

0 / 500 words