Simulation models, meta-analyses, and machine learning all require standardized data, yet in agricultural research much valuable in published literature is found only in unstructured, non-machine-readable formats. Converting this information into analysis-ready datasets can support data reuse though it is labour-intensive, creating a persistent barrier to advancing research.
FAIRagro Use Case 2 introduces a novel automated workflow to address data limitations by systematically extracting and harmonizing data from peer-reviewed journals. Based on Chat Generative Pre-Trained Transformer (GPT)-4.1-mini model (OpenAI), the model on manually curated datasets covering genotypes, context metadata, and management practices (e.g., irrigation, fertilization) is fine tuned to extract structured and consistent metadata from scientific publications.
This work serves as an exemplary approach and resource for accessing, extracting, and utilizing scientific information for agroecosystem models and other analytic tools that require explicit, quantitative data.
Please note:
As of now the FAIRagro Talk series will take place at a new location:

