Resources & Tools β KG-Microbe Knowledge Graph and AI Tools
CultureBotAI led by Dr. Marcin P. Joachimiak develops and maintains various computational resources, databases, and tools for the microbial research community, including the comprehensive KG-Microbe knowledge graph.
Quick Navigation
New to CultureBotAI? Start with Project Ecosystem & Workflows to understand how tools work together.
Looking for specific tools?
- AI Curation Tools - The X-Mech suite of ten autonomous knowledge factories: CultureMech, MediaIngredientMech, CommunityMech, TraitMech, ProteinTraitsMech, AntibioticMech, CellStructureMech, HabitatMech, NaturalProductMech, TaxonMech (suite overview)
- Growth Media Prediction - MicroGrowLink, MicroGrowAgents, KOGUT
- Chemical Data Processing - CultureMech, MicroMediaParam
- Genome Analysis - eggnog_runner, eggnogtable
- Literature Mining - MATE-LLM
- Specialized Research - PFAS, Lanthanide bioprocessing
- Advanced Research Tools - Neurosymbolic reasoning, term extraction (NEW!)
- Developer Resources - Claude Code skills, APIs (NEW!)
Want to see workflows? Jump to Common Workflows
𧬠KG-Microbe: Microbial Knowledge Graph
Overview
KG-Microbe is our flagship resource developed by Dr. Marcin P. Joachimiak - a comprehensive knowledge graph that integrates diverse microbial data sources to enable AI-driven insights and predictions.
π Read the Publication - GigaScience article detailing kg-microbe development and applications.
π KG-Registry entry - Registry record with distributions and metadata. The graph carries 3,000+ organismal and 30,000+ genomic traits.
π METPO Ontology Integration
The Microbial Ecophysiological Trait and Phenotype Ontology (METPO) plays a crucial role in kg-microbe by providing standardized terminology for microbial phenotypes and ecological characteristics.
Key Benefits:
- Knowledge Organization - METPO terms provide semantic structure to organize diverse microbial data within the kg-microbe knowledge graph
- Text Extraction - Standardized ontology terms power automated literature mining and text extraction processes
- Semantic Consistency - Ensures consistent representation of microbial characteristics across different data sources
Links:
- BioPortal: https://bioportal.bioontology.org/ontologies/METPO
- GitHub Repository: github.com/berkeleybop/metpo
Key Features
- Multi-source integration from major biological databases
- Semantic consistency through ontology-driven organization
- Machine-readable formats (RDF, Neo4j, JSON-LD)
- Regular updates with automated data refresh pipelines
- API access for programmatic data retrieval
Data Sources
kg-microbe integrates data from:
- NCBI Taxonomy - Microbial taxonomy and phylogeny
- UniProt - Protein sequences and functional annotations
- GO (Gene Ontology) - Functional gene classifications
- Environmental ontologies - Habitat and growth condition data
- Literature sources - Manually curated cultivation data
Applications
- Growth condition prediction for uncultured organisms
- Taxonomic relationship exploration and phylogenetic analysis
- Literature mining for cultivation protocols
- Cross-organism comparison of growth preferences
Getting Started
# Clone the repository
git clone https://github.com/Knowledge-Graph-Hub/kg-microbe.git
# Install dependencies
cd kg-microbe
pip install -r requirements.txt
# Build latest knowledge graph
kg download
kg transform
kg merge
π Project Ecosystem & Workflows
Understanding the CultureBotAI Ecosystem
The CultureBotAI toolkit consists of interconnected projects organized into a data processing pipeline with kg-microbe as the foundational knowledge graph.
Architecture Overview
βββββββββββββββββββ
β kg-microbe β
β (Foundation) β
ββββββββββ¬βββββββββ
β
ββββββββββββββββββββββββββΌβββββββββββββββββββββββββ
β β β
ββββββββΌβββββββ βββββββββββΌββββββββββ ββββββββββΌβββββββββ
β Data β β Chemical β β Genome β
β Ingestion β β Processing β β Analysis β
β β β β β β
β β’ assay- β β β’ CultureMech β β β’ eggnog_runner β
β metadata β β β’ MicroMediaParam β β β’ eggnogtable β
β β’ MATE-LLM β β β β β
ββββββββ¬βββββββ βββββββββββ¬ββββββββββ ββββββββββ¬βββββββββ
β β β
βββββββββββββββββββββββββΌβββββββββββββββββββββββββ
β
ββββββββββββββΌβββββββββββββ
β AI Agent Systems β
β β
β β’ MicroGrowAgents β
β β’ MicroGrowLink β
β β’ PFASCommunityAgents β
ββββββββββββββ¬βββββββββββββ
β
βββββββββββββββββββββββββΌββββββββββββββββββββββββ
β β β
ββββββββΌβββββββ βββββββββββΌββββββββββ ββββββββββΌβββββββββ
β Specialized β β Web Services β β Analysis β
β Apps β β β β Tools β
β β β β’ MicroGrowLink β β β
β β’ PFAS-AI β β Service β β β’ microbe-rules β
β β’ CMM-AI β β β β β
βββββββββββββββ βββββββββββββββββββββ βββββββββββββββββββ
π€ AI Curation Tools
The X-Mech Suite is a fleet of ten ontology-grounded autonomous knowledge factories (CultureMech, MediaIngredientMech, CommunityMech, TraitMech, ProteinTraitsMech, AntibioticMech, CellStructureMech, HabitatMech, NaturalProductMech, TaxonMech; see the suite overview and relationship graph), coordinated by the culturebotai-claw orchestrator. Their curation workflows transform unstructured microbial cultivation data from literature, laboratory records, and sequence data into standardized, machine-readable knowledge graphs.
Pipeline Overview
Raw Cultivation Records (Literature, Lab Protocols)
β
CultureMech β MediaIngredientMech ββ ingredient mappings βββ KG-Microbe Knowledge Graph
Recipes Ingredient identity β
β β ontologies, mappings,
CommunityMech + other domain-specific Mechs βββββββββββββββββββββββ embeddings
Community, trait, taxon, habitat and molecular evidence
β
AI Predictions (MicroGrowAgents, MicroGrowLink), drawing on the Mechs and KG-Microbe
CultureMech - Autonomous Knowledge Factory for Culture Media
Dedicated Page | GitHub Repository | Web Interface | CC0-1.0 License
15,878 curated culture media recipes from major international repositories, deduplicated into 6,288 canonical media, with LinkML schema, ingredient ontology grounding, and browser-based exploration.
What it does: Curates source-specific recipes, grounds ingredient identifiers, validates records, and generates deduplicated media and browser outputs.
β Learn more on the dedicated CultureMech page
MediaIngredientMech - LLM-Assisted Ingredient Curation
Dedicated Page | GitHub Repository | Web Interface | CC0-1.0 License
2,953 curated ingredient records, 2,611 of them mapped (88% coverage), with LLM-assisted workflows for standardizing microbial cultivation ingredient data. Uses Large Language Models to intelligently map ingredient names to standardized ontology terms.
What it does: Curates ingredient identity, ontology mappings (mostly ChEBI, then MeSH, NCIT, MicrO, FOODON and ENVO; where curation found no term, about 175 kg-microbe registry ids and about 85 CAS numbers, plus about 55 placeholders pending curation), ENVO environmental context, and provenance through validated workflows with human oversight. Scoped updates preserve MIM-owned curation; see the dedicated page for supported commands.
β Learn more on the dedicated MediaIngredientMech page
CommunityMech - Microbial Community Interaction Modeling
Dedicated Page | GitHub Repository | Web Interface | BSD-3-Clause License
422 curated communities across 15 categories, modeled in LinkML with evidence-based ecological interactions for consortium design and multi-organism cultivation.
What it does: Provides structured representation of community composition, syntrophic interactions, and cultivation requirements for multi-species systems.
Related: Provides community evidence for PFASCommunityAgents consortium research.
β Learn more on the dedicated CommunityMech page
TraitMech - Microbial Ecophysiological Traits
GitHub Repository | Web Interface | CC0-1.0 License
Autonomous knowledge factory for microbial ecophysiological traits, seeded from METPO and curated incrementally β 763 trait records across 10 categories; 427 are marked REVIEWED and 519 carry causal graphs.
What it does: Standardizes the trait vocabulary used to describe microbial growth and ecology, and links traits to their evidence and, where a match exists, to kg-microbe.
ProteinTraitsMech - Protein Sequence & Structure Traits
GitHub Repository | Web Interface | CC0-1.0 License
Autonomous knowledge factory for protein sequence, structure, and function traits β 429,293 LinkML-validated records from 34 sources, one YAML per trait. Most are imported from those sources and not yet reviewed. Reviewed records carry evidence-backed causal graphs.
What it does: Extends trait curation from the organism level to the molecular level, connecting protein features to the phenotypes they help explain. Explore the corpus, protein and ESM-2 sequence maps.
CellStructureMech - Microbial Cell Structures
GitHub Repository | Web Interface | CC0-1.0 License (authored content; redistributed UniProt and Complex Portal material is CC BY 4.0)
542 structure records across 13 categories β organelles, envelope layers, appendages, microcompartments and multi-protein complexes β 475 of them grounded in GO cellular component.
What it does: Occupies the layer between traits and proteins, recording what a structure is made of, which organisms have it, what it does, and the causal mechanism by which it does so.
Related: Confers phenotypes recorded as TraitMech terms; hands single proteins off to ProteinTraitsMech.
HabitatMech - Microbial Habitats
GitHub Repository | Web Interface | CC0-1.0 License
3,206 habitat records harmonized from GOLD, BacDive, PREGO and Madin et al., grounded in ENVO, UBERON, FOODON, BTO and PO, with every sourceβs own attestation retained.
What it does: Gives the fleet one record per habitat concept, so that isolation sources expressed differently by each upstream database resolve to a single identity.
Related: Its causal graphs reuse most of TraitMechβs node types; TaxonMech leaves a taxonβs habitats and isolation sources to it; its site generator became AntibioticMechβs. See the relationship graph.
AntibioticMech - Antimicrobial Structures
GitHub Repository | Web Interface | CC0-1.0 License (code and schema) Β· CC BY 4.0 (record content)
2,939 antimicrobial chemical structures, 2,669 ontology-grounded, harmonizing ChEBIβs antimicrobial roles with CARD/ARO molecules, targets and resistance determinants.
What it does: Records one entry per antimicrobial structure, carrying its mode of action, molecular targets and the evidence for both where sources or curation have supplied them β 454 records have a mode of action and 282 a molecular target so far β and places all 2,939 on a chemical map by molecular fingerprint.
NaturalProductMech - Natural Product Structures
GitHub Repository | Web Interface | CC0-1.0 License (code and schema) Β· CC BY 4.0 (record content)
3,115 natural product structures, one per Standard InChIKey, seeded from nine sources and grounded in ChEBI, MIBiG and NCBI Taxonomy.
What it does: Links every structure to the biosynthetic gene cluster it comes from (3,115 of 3,115) and to its producer organisms (3,076), and carries cited occurrences (2,342) and measured bioactivities (176) where a source reports them. Producer claims are graded: of 3,407, only 805 rest on evidence that addressed the organism. The current records have no REVIEWED status entries; mechanism and evidence coverage vary by record.
TaxonMech - Microbial Taxa and Strains
GitHub Repository | Web Interface | CC0-1.0 License (project content; upstream sources retain their own terms)
625,960 taxon records at species level and below, with 100,745 distinct listed strains, resolving NCBI Taxonomy, GTDB, LPSN and BacDive onto one NCBI-grounded identity.
What it does: Serves as the taxonomic counterpart of the other Mechs. Its 21,071 listed strains with genome links carry NCBI, GTDB, BV-BRC/PATRIC, IMG or AllTheBacteria identifiers; taxon pages also show StrainInfo references and the supporting evidence. Higher taxa appear only in a recordβs lineage, never as records of their own.
Related: Leaves a taxonβs traits to TraitMech, its habitats to HabitatMech and its growth media to CultureMech, keeping only attestation counts. Most taxa the other Mechs cite are TaxonMech records, which the relationship graph shows as shared NCBI Taxonomy identifiers.
Common Workflows
Workflow 1: Novel Organism Media Prediction
- Start with organism taxonomy/genome
- Run eggnog_runner + eggnogtable (functional annotation)
- Query kg-microbe (related organisms, known preferences)
- Use MicroGrowAgents (integrate genome, literature, analogies)
- Get media recommendations with evidence
Workflow 2: Chemical Compound Knowledge Graph Integration
- Start with media composition text
- Run CultureMech (extract chemical entities)
- Run MicroMediaParam (map to ChEBI/PubChem)
- Integrate into kg-microbe (standardized chemical data)
- Enable downstream media predictions
Workflow 3: PFAS Biodegradation Consortia Design
- Query PFAS-AI database (identify candidate organisms)
- Extract genome features (eggnog_runner/eggnogtable)
- Query kg-microbe (environmental compatibility)
- Use PFASCommunityAgents (design optimized consortia)
- Get consortium composition + rationale
Workflow 4: Literature-Driven Culture Optimization
- Run MATE-LLM (extract cultivation protocols from papers)
- Integrate into kg-microbe (structured cultivation data)
- Use MicroGrowAgents LiteratureAgent (mine similar organisms)
- Get evidence-based media recommendations
Getting Started Guide
For Growth Media Prediction:
- Start with: MicroGrowAgents or MicroGrowLink
- Prerequisites: Access to kg-microbe knowledge graph
- Recommended workflow: Workflow 1
For Chemical Data Processing:
- Start with: CultureMech or MicroMediaParam
- Prerequisites: Media composition text data
- Recommended workflow: Workflow 2
For Specialized Research:
- PFAS biodegradation: Start with PFAS-AI, then PFASCommunityAgents
- Lanthanide bioprocessing: Start with CMM-AI
For Web-Based Access:
- API users: Start with MicroGrowLinkService
- Prerequisites: HTTP client, REST API knowledge
π§ CultureBotAI Software & Tools
Growth Media Prediction & Design
MicroGrowLink
Private repository β public release planned | Python
Knowledge graph-based framework for predicting microbial growth media using advanced graph and transformer models. Integrates microbial, chemical, and environmental data into a heterogeneous knowledge graph and applies link prediction to forecast which media enable growth of given taxa.
Supported Models:
- RGT (Relational Graph Transformer)
- HGT (Heterogeneous Graph Transformer)
- NBFNet (Neural Bellman-Ford Network)
Key Features:
- Heterogeneous knowledge graph integration
- Advanced transformer-based link prediction
- Multi-modal data integration (microbial, chemical, environmental)
Related Projects:
- Depends on: kg-microbe (knowledge graph foundation), MicroMediaParam (chemical compound mappings)
- Feeds into: Media formulation recommendations, MicroGrowLinkService (API deployment)
- Works with: MicroGrowAgents (complementary multi-agent predictions)
MicroGrowAgents
Dedicated Page | GitHub Repository β private repository, public release planned | Python | BSD-3-Clause
Agent-based system for AI-driven microbial cultivation and growth media design. Bridges the microbial cultivation gap through AI-powered multi-agent systems that integrate knowledge graphs, machine learning, and experimental automation.
Specialized Agents:
- LiteratureAgent - Mining 245+ papers for cultivation protocols
- AnalogyReasoningAgent - Cross-organism comparison and reasoning
- GenomeFunctionAgent - Auxotrophy detection from 57 Bakta-annotated genomes (667K features)
- MediaFormulationAgent - Schema-driven media recommendation with evidence-based ingredient suggestions
Key Achievements:
- 864,363 validated species across bacteria, archaea, fungi, and protozoa (GTDB + LPSN + NCBI)
- Multi-modal reasoning combining literature mining, metabolic modeling (FBA/gap-filling), chemical similarity (208K+ embeddings)
- Genome-guided design for organism-specific media formulation
Related Projects:
- Depends on: kg-microbe (knowledge graph foundation), MicroMediaParam (chemical mappings), eggnogtable (genome annotations), MATE-LLM (literature extraction)
- Feeds into: Media formulation recommendations, PFASCommunityAgents (consortium design)
- Works with: MicroGrowLink (complementary prediction approach)
KOGUT Transformer
DOE CODE 175162 | doi:10.11578/dc.20260210.3 | Released 2025-12-18
KOGUT (Knowledge Oriented Graph Unified Transformer) adapts the Relational Graph Transformer (RelGT) architecture β originally designed for relational tables and multi-table databases β to heterogeneous biological knowledge graphs, for link prediction over kg-microbe. Its primary task is predicting which growth media support a given microbial taxon.
Training data:
- Merged kg-microbe knowledge graph: 1,392,337 nodes and 2,960,472 edges
- 24 Biolink relation types spanning taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships
- Primary task: growth media suitability (
biolink:occurs_in, ~50K edges); the model can predict any of the 24 relation types - Trained on NVIDIA A100 GPUs at NERSC Perlmutter
Adaptations beyond the original RelGT:
- Multimodal node encoding from KG metadata β labels, categories, descriptions, and synonyms
- Extended k-hop subgraph sampling (3-hop default) tuned for sparse biological networks
- Biolink predicate preservation, with type-specific transformations for the 24 edge semantics
- Inductive learning, enabling zero-shot prediction for novel and uncultured taxa from feature-based embeddings (temperature, oxygen requirement, gram stain, cell shape)
Reported performance on growth media prediction: MRR 0.9966, Precision@1 0.9932, Hit@10 1.0000.
Status: registered in DOE CODE; no public source repository yet.
Related Projects:
- Depends on: kg-microbe (training graph)
- Works with: MicroGrowLink and MicroGrowAgents (complementary media-prediction approaches); explainable rule mining (interpretable counterpart)
MicroMediaParam
GitHub Repository | Python
Comprehensive chemical compound knowledge graph mapping pipeline for microbial growth media analysis. Extracts chemical compounds from media compositions and maps them to knowledge graph entities with standardized chemical properties.
Features:
- Processes 23,181 chemical entries from 1,807 microbial growth media
- 78% ChEBI coverage (18,088 compounds mapped)
- Multi-database mapping to ChEBI, PubChem, and CAS-RN identifiers
- Intelligent hydrate parsing and molecular weight calculation
- Solution expansion for DSMZ solution references
- 99.99% chemical mapping accuracy
Related Projects:
- Depends on: CultureMech (chemical entity extraction)
- Feeds into: kg-microbe (standardized chemical data), MicroGrowAgents (chemical mappings), MicroGrowLink (knowledge graph integration)
- Works with: assay-metadata (compound identification)
CultureMech
Dedicated Page | GitHub Repository | Web Interface | CC0-1.0 License
15,878 curated culture media recipes deduplicated into 6,288 canonical media, with chemical entity extraction and ontology grounding. Part of the X-Mech suite of autonomous knowledge factories.
β See the dedicated CultureMech page for full documentation, use cases, and examples.
Related Projects:
- Depends on: Text-based media composition data
- Feeds into: MicroMediaParam (entity mapping), kg-microbe (chemical data integration), MediaIngredientMech (ingredient curation)
- Works with: assay-metadata (standardized substrate processing)
Specialized Research Pipelines
CMM-AI: Lanthanide Bioprocessing Data Pipeline
Private repository β public release planned | Python
Automated data pipeline for lanthanide bioprocessing research, focusing on rare earth element-dependent biological processes in microorganisms. Integrates multiple biological databases to create comprehensive research datasets.
Scientific Focus:
- XoxF methanol dehydrogenase systems (lanthanide-dependent enzymes)
- Methylotrophic bacteria (Methylobacterium, Methylorubrum, Paracoccus)
- Environmental metal cycling and biogeochemistry
- Siderophore/lanthanophore transport mechanisms
- PQQ-dependent enzyme complexes
Related Projects:
- Depends on: kg-microbe (organism data), eggnogtable (functional annotations)
- Feeds into: Specialized lanthanide bioprocessing research
- Works with: MicroGrowAgents (media optimization for lanthanide-dependent organisms)
PFAS-AI: Machine Learning-Enabled PFAS Biodegradation Pipeline
GitHub Repository | Python
ML-enabled data pipeline for PFAS biodegradation research, focusing on identification and characterization of microorganisms capable of degrading per- and polyfluoroalkyl substances (PFAS).
Research Objectives:
- ML-Powered Database - Semantically-aware database using KG-Microbe platform to identify putative PFAS biodegradation genes, pathways, taxa, and environments
- Intelligent Consortia Design - Graph learning and LLMs to design optimized microbial consortia for PFAS remediation
Scientific Focus:
- C-F bond cleavage mechanisms (dehalogenases and defluorinases)
- Fluoride resistance systems
- Hydrocarbon degradation pathways
- Environmental context (AFFF-contaminated sites, groundwater, wastewater)
Related Projects:
- Depends on: kg-microbe (organism and gene identification)
- Feeds into: PFASCommunityAgents (consortium design)
- Works with: eggnogtable (functional gene annotations)
PFASCommunityAgents
Private repository β public release planned | Python
Multi-agent system for designing optimized microbial consortia for PFAS biodegradation. Uses AI-powered reasoning to compose consortia with complementary metabolic capabilities and syntrophic relationships.
Key Features:
- Consortium composition optimization
- Syntrophic relationship prediction
- Environmental context-aware design
- Multi-species compatibility assessment
Related Projects:
- Depends on: PFAS-AI (candidate organism database), MicroGrowAgents (agent architecture), kg-microbe (organism relationships)
- Feeds into: PFAS remediation research and consortium cultivation
- Works with: MicroGrowAgents (media design for consortia)
Data Processing & Analysis
assay-metadata: BacDive API Assay Metadata Extractor
GitHub Repository | Python
Extracts API assay metadata from BacDive JSON data with comprehensive identifier mappings to CHEBI, EC, RHEA, and PubChem databases.
Capabilities:
- Parses 99,392 bacterial strain records from BacDive
- Extracts 17 unique API kit types (API zym, API 50CHac, etc.)
- Maps substrate codes to CHEBI and PubChem identifiers
- Maps enzyme EC numbers to RHEA reaction databases
- Generates consolidated JSON metadata files
Related Projects:
- Depends on: BacDive API data
- Feeds into: kg-microbe (phenotypic assay data integration)
- Works with: MicroMediaParam (compound identification), CultureMech (substrate processing)
eggnog_runner
GitHub Repository | Python
Automated pipeline for running EggNOG-mapper functional annotation at scale. Processes genome assemblies in batch to generate functional annotations for downstream analysis.
Features:
- Batch genome processing with parallel execution
- Automated EggNOG-mapper execution
- Output standardization and organization
- Error handling and retry logic
Related Projects:
- Depends on: kg-microbe (genome data), EggNOG-mapper tool
- Feeds into: eggnogtable (annotation post-processing)
- Works with: MicroGrowAgents GenomeFunctionAgent (functional predictions)
eggnogtable
Private repository β public release planned | Python
Post-processing pipeline for EggNOG-mapper output into structured datasets. Extracts and organizes functional annotations including GO terms, EC numbers, and KEGG pathways.
Features:
- GO term extraction and organization
- EC number mapping
- KEGG pathway assignment
- Ontology term integration
- Structured dataset generation
Related Projects:
- Depends on: eggnog_runner (annotation output)
- Feeds into: MicroGrowAgents GenomeFunctionAgent (auxotrophy detection), kg-microbe (functional annotations), CMM-AI (enzyme identification)
- Works with: assay-metadata (enzyme EC number mapping)
microbe-rules: Machine Learning Models for Microbial Data
GitHub Repository | Python
Research code repository containing machine learning models and analysis pipelines for binary classification and comparative modeling of microbial datasets.
Features:
- Binary classification models for microbial data
- Model comparison and evaluation frameworks
- Automated data preparation pipelines
- Reproducible research workflows
Related Projects:
- Depends on: kg-microbe (training data), various microbial datasets
- Feeds into: Model optimization research
- Works with: MicroGrowLink (model comparison), MicroGrowAgents (ML component evaluation)
AI Agent Systems
MATE-LLM
Private repository β public release planned | Python
LLM-powered system for extracting structured microbial information from scientific literature. Automates the extraction of cultivation protocols, growth conditions, and microbial annotations from research papers.
Key Features:
- Entity extraction from scientific literature
- Automated cultivation protocol annotation
- Literature mining for growth conditions
- Knowledge graph integration preparation
- Structured data generation from unstructured text
Related Projects:
- Depends on: Scientific literature corpus, LLM APIs
- Feeds into: kg-microbe (literature-derived data), MicroGrowAgents LiteratureAgent (cultivation protocols)
- Works with: METPO ontology (standardized terminology)
Web Services & APIs
MicroGrowLinkService
GitHub Repository | Python | REST API | BSD-3-Clause
RESTful API service wrapper for MicroGrowLink prediction models. Provides HTTP endpoints for programmatic access to growth media predictions and enables integration with laboratory information management systems (LIMS).
Key Features:
- HTTP API endpoints for predictions
- Model serving infrastructure
- Batch prediction support
- LIMS integration capabilities
- Production deployment configuration
Related Projects:
- Depends on: MicroGrowLink (prediction models), kg-microbe (knowledge graph)
- Feeds into: External applications, LIMS integrations, web interfaces
- Works with: MicroGrowAgents (complementary API services)
π¬ Advanced Research Tools
neurosymbolreason - Neurosymbolic Analogy Reasoning
GitHub Repository | Python
Neurosymbolic analogy reasoning on microbial knowledge graph embeddings to analyze relationships between microbial taxa and their physical growth preferences. Combines neural network embeddings with symbolic reasoning for cross-organism inference.
Key Features:
- Knowledge graph embedding analysis
- Analogy-based reasoning for growth preferences
- Taxonomic relationship exploration
- Novel organism growth condition inference
Related Projects:
- Depends on: kg-microbe (knowledge graph embeddings), taxonomic data
- Feeds into: MicroGrowAgents (AnalogyReasoningAgent), growth prediction pipelines
- Works with: MicroGrowLink (complementary prediction approach)
auto-term-catalog - Automated Term Extraction
GitHub Repository | Python
Code for extracting AUTO terms from ontoGPT output. Processes ontology-based text mining results to create curated term catalogs for microbial cultivation research.
Key Features:
- OntoGPT output processing
- Automated term extraction and cataloging
- Integration with METPO ontology
- Standardized term generation
Related Projects:
- Depends on: OntoGPT output, METPO ontology
- Feeds into: kg-microbe (ontology terms), MATE-LLM (standardized vocabulary)
- Works with: Literature mining pipelines
π Developer Resources
culturebot-skills - Claude Code Skills
GitHub Repository | Skills/Configuration
Claude Code skills for CultureBot/KG-Microbe projects. Custom skills and workflows for AI-assisted development within the CultureBotAI ecosystem.
What it provides:
- Pre-configured Claude Code skills for common tasks
- Project-specific development workflows
- Integration helpers for CultureBotAI tools
- Best practices and code patterns
Use Cases:
- Automated code generation for kg-microbe integrations
- Data pipeline development assistance
- Documentation generation
- Testing and validation workflows
Getting Started:
# Install Claude Code skills
git clone https://github.com/CultureBotAI/culturebot-skills.git
# Follow setup instructions in repository README
π Datasets
Curated Cultivation Database
Curated collection of cultivation protocols for diverse microorganisms based on reference sources and literature.
Contents:
- Growth media compositions
- Environmental conditions (temperature, pH, atmosphere)
- Cultivation methods and protocols
- Literature references
Environmental Metadata Collection
Comprehensive dataset linking microorganisms to their natural habitats and environmental conditions.
π Related Organizations & Resources
Academic & Research Institutions
- ABPDU - Advanced Biofuels and Bioproducts Process Development Unit
- BacDive - Bacterial Diversity Metadatabase
- Cultivarium - Global microbial cultivation platform
- JBEI - Joint BioEnergy Institute
- JGI GOLD - Genomes Online Database
- KBase - Systems Biology Knowledgebase
- NMDC - National Microbiome Data Collaborative
- Palsson Lab - UC San Diego Systems Biology Research Group
Commercial Organizations
- Biolog - Microbial identification and characterization systems
- Isolation Bio - Microbial isolation and cultivation technology
π Documentation & Tutorials
API Documentation
Comprehensive documentation for programmatic access to kg-microbe and related tools:
- Neo4j graph database interface
- Python SDK usage examples
- Data schema specifications
Tutorials
Coming soon!
Example Notebooks
Jupyter notebooks demonstrating practical applications:
- Growth condition prediction workflows
- Literature mining pipelines
- Data visualization examples
π Data Access & APIs
Direct Downloads
- Knowledge Graph Dumps - Complete RDF/TTL files
- Processed Datasets - CSV/JSON formatted data tables
- Ontology Files - OWL/RDF ontology definitions
API Endpoints
Coming soon
Query Interfaces
Coming soon
π¦ Software Packages
Python Packages
Coming soon
π€ Community & Collaboration
Contributing
We welcome contributions from the research community:
- Data contributions - Share cultivation protocols and growth data
- Software development - Contribute to open source tools
- Literature curation - Help extract cultivation data from papers
- Validation - Test predictions against experimental results
Discussion Forums
- GitHub Discussions - Technical questions and feature requests
- Slack Community - Real-time collaboration and support
- Monthly Webinars - Updates and community presentations
Citation
If you use kg-microbe or other CultureBotAI resources in your research, please cite:
Santangelo, B.E., Hegde, H., Caufield, J.H., Reese, J., Kliegr, T., Hunter, L.E.,
Lozupone, C.A., Mungall, C.J., Joachimiak, M.P. (2026). KG-Microbe - Building
Modular and Scalable Knowledge Graphs for Microbiome and Microbial Sciences.
GigaScience, giag077. https://doi.org/10.1093/gigascience/giag077
Support & Contact
For technical support, collaboration inquiries, or questions about our resources:
- Email: MJoachimiak@lbl.gov
- GitHub Issues: Report bugs or request features
- Documentation: Comprehensive guides and API references
- Community Forums: Connect with other researchers and developers
Bibliography
- Santangelo BE, Hegde H, Caufield JH, Reese J, Kliegr T, Hunter LE, Lozupone CA, Mungall CJ, Joachimiak MP. KG-Microbe β Building Modular and Scalable Knowledge Graphs for Microbiome and Microbial Sciences. GigaScience. 2026;giag077. doi:10.1093/gigascience/giag077
- MΓ‘Ε‘a P, Kliegr T, Joachimiak MP. Explainable rule-based prediction of cultivation media for microbes. Computational and Structural Biotechnology Journal. 2025;27:5194β5206. doi:10.1016/j.csbj.2025.10.014 Β· free full text
- Naseem S, Miller MA, Martinez-Gomez NC, Sun N, Joachimiak MP. MicroGrowAgents: An Agentic AI System for Microbial Cultivation Engineering. bioRxiv. 2026. doi:10.64898/2026.06.04.729985
- Joachimiak MP. Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1 [software]. DOE CODE; 2025. doi:10.11578/dc.20260210.3 Β· DOE CODE 175162
- Joachimiak MP, Santangelo BE, Hegde H, Caufield JH, Reese J, Kliegr T, Hunter LE, Lozupone CA, Mungall CJ. kg-microbe: modular knowledge graph for microbiome and microbial sciences [software]. github.com/Knowledge-Graph-Hub/kg-microbe
- METPO: Microbial Ecophysiological Trait and Phenotype Ontology. BioPortal Β· GitHub
- Joachimiak MP. βRuleML/GOBLIN COST Action Lecture on Data Science: Teaching AI to Teach Humans About Microbiologyβ [talk]. RuleML / COST GOBLIN Action Seminar; 2026. Recording