MediaIngredientMech: Autonomous Knowledge Factory for Media Ingredients
Overview
MediaIngredientMech is an autonomous knowledge factory for culture-media ingredient identity and ontology mappings, with LLM-assisted curation and human oversight. It maintains ingredient records, synonyms, mapping quality, environmental context, and an audit trail for curation decisions. Repository overview.
The published browser contains 2,951 ingredients: 2,616 MAPPED, 261 UNMAPPED, and 74 REJECTED, giving 89% mapped coverage after rounding. These figures were checked on September 20, 2026 against the browser’s live data index.
Explore the Published Site
- Ingredient browser — search names, synonyms, and ontology identifiers; filter by source, mapping status, and mapping quality.
- Mapped ingredients — browse the mapped subset.
- Embedding map and graph layout — explore ingredient relationships in KG-Microbe embedding space.
What a Record Represents
An ingredient record identifies a practical reagent or formulation used in media. Curation distinguishes salts, hydrates, mixtures, and other forms, preserves raw names as synonyms, and records the convention used when sources are ambiguous. Mapping quality and curation status are separate fields. Mapping semantics.
The current model includes:
- Ingredient records with identifiers, synonyms, mapping status, and curation history.
- Ontology mappings to ChEBI and FOODON, with quality ratings; the browser also exposes NCIT and CAS identifiers.
- Environmental context linked to ENVO terms with relevance qualifiers.
- Curation events that record provenance and LLM assistance.
- Component relationships for ingredients made of other components, with evidence and validation.
See the schema reference, environmental-context model, and component model.
Curation Workflow
Curators compare upstream recipe changes with the tracked ingredient corpus, make scoped updates with provenance, and validate ontology identifiers and labels through OAK/OLS. Ingredient occurrence counts help prioritize unmapped records. Validated mapping artifacts can then support coordinated downstream updates to CultureMech.
The former CultureMech collection writers are retired because their aggregate projections could overwrite ingredient curation. Current guidance is to review upstream changes and apply scoped updates to MIM-owned records. Current workflow and migration status.
Getting Started
Development and CI use Python 3.13, uv, and just:
git clone https://github.com/CultureBotAI/MediaIngredientMech.git
cd MediaIngredientMech
just install
just gen-schema
just validate-all
For an interactive curation session:
just snapshot
just curate
just report
These are the repository’s documented commands. See the curation guide and role-curation workflow for record editing and validation.
Exports and Integration
The project provides YAML records, browser JSON, generated inventories, and SSSOM mappings. SSSOM predicates preserve distinctions between exact, close, broader, and narrower matches. Registry identity mappings and ontology assertions have different meanings; downstream consumers should follow the mapping contract.
Deep-research tools help select providers and prepare ingredient research. Their results remain curation proposals until identity and evidence are validated. Provider workflow.
Repository & Documentation
- Repository and published site
- Current mapping inventory
- Workflow guide
- License: CC0-1.0, as stated in the repository
Related Tools
- X-Mech Suite overview - All ten Mechs, their shared vocabulary and cross-references, and the culturebotai-claw orchestrator
- TaxonMech - Microbial taxa and strains grounded in NCBI Taxonomy, harmonized with GTDB, LPSN and BacDive
- HabitatMech - Habitats harmonized from GOLD, BacDive, PREGO and Madin et al. into ENVO-grounded records
- CommunityMech - Microbial community interaction modeling
- TraitMech - Autonomous knowledge factory for microbial ecophysiological traits
- CellStructureMech - Microbial cell structures, between the trait and protein layers
- ProteinTraitsMech - Protein sequence, structure, and function traits
- NaturalProductMech - Natural product structures with their producer organisms and gene clusters
- AntibioticMech - Antimicrobial structures harmonizing ChEBI and CARD/ARO
- CultureMech - Chemical entity extraction from media recipes (6,286 canonical media)
- MicroMediaParam - Chemical compound standardization (78% ChEBI coverage)
- kg-microbe - Central knowledge graph for microbial cultivation
Contact & Collaboration
For questions about MediaIngredientMech or to contribute:
- Principal Investigator: Dr. Marcin P. Joachimiak
- Email: mjoachimiak@lbl.gov
- Organization: CultureBotAI
- Laboratory: Environmental Genomics and Systems Biology Division, Lawrence Berkeley National Laboratory
Bibliography
- Santangelo BE, Hegde H, Caufield JH, Reese J, Kliegr T, Hunter LE, Lozupone CA, Mungall CJ, Joachimiak MP. KG-Microbe — Building Modular and Scalable Knowledge Graphs for Microbiome and Microbial Sciences. GigaScience. 2026;giag077. doi:10.1093/gigascience/giag077
- Caufield JH, Hegde H, Emonet V, Harris NL, Joachimiak MP, et al. Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning. Bioinformatics. 2024;40(3):btae104. doi:10.1093/bioinformatics/btae104 · free full text
- METPO: Microbial Ecophysiological Trait and Phenotype Ontology. BioPortal · GitHub