MediaIngredientMech: Autonomous Knowledge Factory for Media Ingredients

Overview

MediaIngredientMech is an autonomous knowledge factory for culture-media ingredient identity and ontology mappings, with LLM-assisted curation and human oversight. It maintains ingredient records, synonyms, mapping quality, environmental context, and an audit trail for curation decisions. Repository overview.

The published browser contains 2,951 ingredients: 2,616 MAPPED, 261 UNMAPPED, and 74 REJECTED, giving 89% mapped coverage after rounding. These figures were checked on September 20, 2026 against the browser’s live data index.

Explore the Published Site

What a Record Represents

An ingredient record identifies a practical reagent or formulation used in media. Curation distinguishes salts, hydrates, mixtures, and other forms, preserves raw names as synonyms, and records the convention used when sources are ambiguous. Mapping quality and curation status are separate fields. Mapping semantics.

The current model includes:

See the schema reference, environmental-context model, and component model.

Curation Workflow

Curators compare upstream recipe changes with the tracked ingredient corpus, make scoped updates with provenance, and validate ontology identifiers and labels through OAK/OLS. Ingredient occurrence counts help prioritize unmapped records. Validated mapping artifacts can then support coordinated downstream updates to CultureMech.

The former CultureMech collection writers are retired because their aggregate projections could overwrite ingredient curation. Current guidance is to review upstream changes and apply scoped updates to MIM-owned records. Current workflow and migration status.

Getting Started

Development and CI use Python 3.13, uv, and just:

git clone https://github.com/CultureBotAI/MediaIngredientMech.git
cd MediaIngredientMech
just install
just gen-schema
just validate-all

For an interactive curation session:

just snapshot
just curate
just report

These are the repository’s documented commands. See the curation guide and role-curation workflow for record editing and validation.

Exports and Integration

The project provides YAML records, browser JSON, generated inventories, and SSSOM mappings. SSSOM predicates preserve distinctions between exact, close, broader, and narrower matches. Registry identity mappings and ontology assertions have different meanings; downstream consumers should follow the mapping contract.

Deep-research tools help select providers and prepare ingredient research. Their results remain curation proposals until identity and evidence are validated. Provider workflow.

Repository & Documentation



Contact & Collaboration

For questions about MediaIngredientMech or to contribute:


Bibliography

  1. Santangelo BE, Hegde H, Caufield JH, Reese J, Kliegr T, Hunter LE, Lozupone CA, Mungall CJ, Joachimiak MP. KG-Microbe — Building Modular and Scalable Knowledge Graphs for Microbiome and Microbial Sciences. GigaScience. 2026;giag077. doi:10.1093/gigascience/giag077
  2. Caufield JH, Hegde H, Emonet V, Harris NL, Joachimiak MP, et al. Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning. Bioinformatics. 2024;40(3):btae104. doi:10.1093/bioinformatics/btae104 · free full text
  3. METPO: Microbial Ecophysiological Trait and Phenotype Ontology. BioPortal · GitHub

Full publication list →