WA State Agricultural Biotechnology Centre

Real-time digital twin platform for predicting grain yield using genotype-by-environment data

Technology
Conceptual
Company

A cloud-based digital twin platform that combines genotype, phenotype, and environmental data to deliver near real-time grain yield predictions for barley and other cereals. By integrating IoT sensor inputs, relational databases, and machine learning pipelines, the platform aims to reduce reliance on large-scale field trials and enable faster decision-making for growers and breeders.

Overview

This solution addresses a critical limitation in agricultural forecasting: grain yield predictions based on genotype-by-environment (GxE) interactions are typically constrained by the statistical significance of genome-wide association studies, which cannot capture real-time environmental variation. The proposed digital twin platform integrates genotypic, phenotypic, and environmental datasets into a continuously updated knowledge graph, enabling near real-time grain yield forecasts without the need for repeated large-scale field trials.

The platform is designed to serve cereal breeders, agronomists, and precision agriculture operators who require faster, data-driven yield forecasts. By combining drone-collected field data, glasshouse phenotypic measurements, IoT sensor streams, and external weather APIs, the system aims to minimize prediction drift caused by year-to-year GxE variability and accelerate breeding and crop management decisions.

Technical specifications

Core technology stack:

  • Cloud-based relational databases with scalable virtual machine hosting
  • Near real-time API and JSON data injection using dynamic SQL and Python pipelines
  • Azure Digital Twins console with SQL joins and queries for knowledge graph generation
  • Automation layer including VPN, CRON scheduling, Python scripts, MS Power Automate, SSMS, and Power BI for visualization
  • Integration of NASA and Bureau of Meteorology weather APIs for environmental data streams

Data inputs:

  • Phenotypic data: canopy temperature, chlorophyll content, NDVI/RGB imagery, Zadoks maturity stage, plant height, dry-matter weight, stomatal conductance, thermal images, biomass and grain yield, root biomass, seed color, 100-grain weight, leaf and head type, harvest height, straw strength, and root traits including length, number, and root hairs or mycorrhizae
  • Genotypic data: whole genome sequences for approximately 120 doubled-haploid barley lines with seed color and trait annotation
  • Environmental data: LTE IoT sensor streams and complementary public weather feeds

Analytical approach:

  • Machine learning pipelines validated against state-of-the-art GxE modelling algorithms
  • Knowledge graph construction linking genotype, phenotype, and environment variables
  • Bias correction framework to account for residual GxE drift across growing seasons
Technology readiness level

The platform is at a mid-to-late validation stage. The team has already collected and modelled phenotype and genotype data for a large doubled-haploid barley population, established the framework for digital twin construction, and verified that existing machine learning approaches to GxE variation are sufficiently robust to underpin twin-based predictions. Complementary in-house cereal sequence datasets and relational database automation expertise are in place.

Remaining work focuses on quantifying how much environmental data is required to minimize forecast variation, determining whether a physical limit exists to GxE variation for the digital twin, and deploying the full automation stack within the Murdoch University agritech and AI infrastructure. Once validated, the platform will be ready for pilot engagement with breeding programs and precision agriculture partners seeking real-time yield decision support.

Halo home
Partner smarter. Move faster.
Get new partnering requests
delivered to your inbox.