GC mid1

METPO:1000430 · CLASS · REVIEWED

A GC-content phenotype with genome-wide GC composition above approximately 66.3% (the METPO `GC_>66.3` bin; note that the upstream label 'mid1' does not match this high-end numeric threshold, but the synonym is preserved as the authoritative bin definition).

GC-mid1 (METPO >66.3%) high-GC bin

DOI-backed graph linking strong GC-biased gene conversion to a GC content above ~66.3% (the threshold encoded by the METPO synonym GC_>66.3 on this record).

GC-mid1 (METPO >66.3%) high-GC bin Interactive directed graph showing evidence-backed causal relationships for GC mid1.

Edge evidence

  • strong GC-biased gene conversion confers GC mid1 METPO:2007700

    Strong GC-biased gene conversion yields high genome-wide GC composition.

    • DOI:10.1186/1471-2148-10-374 GC-biased gene conversion Supports GC-biased gene conversion as the driver of high-GC composition.
  • GC mid1 is a GC content rdfs:subClassOf

    GC mid1 is a quantitative bin of the GC-content phenotype.

    • DOI:10.1038/nrg2358 GC content Supports the >66.3% bin as a value within the GC-content distribution.
  • NHEJ pathway (Ku, LigD) positively correlated with GC mid1

    Ku/NHEJ presence is strongly associated with elevated genomic GC content.

    • DOI:10.1371/journal.pgen.1008493 Strong association between Ku presence and elevated GC content; Pearson r=0.54 (p<2.2e-16). Correlational across prokaryotes.
  • high double-strand break rate selects for GC mid1 METPO:2007401

    Higher DSB formation rate selects for increased GC content relative to genomic background.

    • DOI:10.1371/journal.pgen.1008493 Sites with higher DSB rates are under selection for increased GC content; proposed unifying driver of GC content across prokaryotes.
  • DNA replication and repair (DRR) system composition strongly correlated with GC mid1

    DRR-system KEGG-ortholog composition explains a large fraction of genomic GC variance.

    • DOI:10.1128/spectrum.02145-22 Linear model using 217 DRR-related KEGG orthologs explains up to 88% of variance in genomic GC (multiple correlation 0.94).
  • GC mid1 may increase NHEJ end-joining efficiency

    High GC may increase NHEJ repair efficiency by stabilizing short overhangs/microhomologies via extra hydrogen bonds.

    • DOI:10.1371/journal.pgen.1008493 High GC via increased hydrogen bonds may stabilize short overhang/microhomology pairing and increase NHEJ efficiency.

Provenance

Source
METPO (2025-11-25)
Definition source
DOI:10.1038/nrg2358

Parent traits (1)

Synonyms (1)

  • GC_>66.3 RELATED_SYNONYM · metpo.owl

kg-microbe context

Matched 1 kg-microbe node via direct_metpo.

  • METPO:1000430 [-2.804, -2.753, -0.396, +5.171, …]

512-dim DeepWalkSkipGramEnsmallen embedding from kg-microbe (2026-04-25).

Nearest neighbors in embedding space

Top-8 cosine-similar METPO traits from the 2026-04-25 deepwalk (512-D).

Deep research

Generated by just research-trait; source: research/traits/genomics/gc_mid1-deep-research-falcon.md

Unreviewed literature output — not curated TraitMech content Ontology identifiers suggested below have not been resolved against their ontologies, and some are known to be wrong. Check any CURIE against the source before using it.
# Curation report: **GC mid1** (`METPO:1000430`)

## Executive curation recommendation

`METPO:1000430` should be modeled as an **assay-derived whole-genome nucleotide-composition class**, not as a metabolic pathway, physiological capacity, or environmental preference. The operational phenotype is:

\[
GC_w=\frac{G+C}{A+T+G+C}>0.663\;\text{(approximately)}.
\]

The authoritative synonym `GC_>66.3` should govern interpretation. The upstream label **“GC mid1” is misleading**, because 66.3% is a high-end bin; preserve it only as the supplied label and add a curation note. Across bacteria, reported genomic GC content spans roughly <25% to 75%, while one broader prokaryotic compilation reported 8–75%, placing the threshold near the extreme high end rather than the middle (hershberg2015mutation—theengineof pages 6-7, hu2022apositivecorrelation pages 1-2).

The strongest defensible causal architecture is:

**DNA replication errors → nucleotide-specific mismatches → proofreading/MMR-dependent mutation spectrum → long-term substitution supply**, opposed or overridden by **recombination-associated GC-biased gene conversion (gBGC) and possibly selection → preferential persistence/fixation of G/C alleles → elevated whole-genome GC → `METPO:1000430`.**

However, no retrieved experiment directly drove a lineage across the **66.3% threshold**. Therefore, molecular repair edges can be curated strongly at the mutation-spectrum level, whereas final edges into `METPO:1000430` require evolutionary-timescale and uncertainty qualifiers.

## 1. Trait scope and boundary cases

### Included

- **Unit:** preferably a complete or sufficiently unbiased draft chromosome/genome.
- **Observation:** percentage of guanine plus cytosine among called genomic DNA bases.
- **Classification:** positive when whole-genome GC is above approximately 66.3%.
- **Examples:** *Deinococcus radiodurans* at 66.61% is just above the boundary; *Streptomyces* genomes at approximately 72% are clearly within the bin (long2018specificityofthe pages 1-2, dagva2024correctionofnonrandom pages 1-2).

### Excluded or separately modeled

1. **GC3 or fourfold-degenerate-site GC.** These are informative about weakly selected substitutions but are not equivalent to whole-genome GC.
2. **Coding-region, core-genome, accessory-genome, plasmid, or intergenic GC.** These can differ materially within one organism.
3. **rRNA/tRNA GC.** Structural-RNA GC may respond to temperature differently from whole-genome GC; older work found structural-RNA associations even when whole-genome associations were absent (hu2022apositivecorrelation pages 1-2).
4. **Local GC islands or horizontally transferred segments.** A local high-GC region does not establish the genome-wide phenotype.
5. **GC skew.** Strand asymmetry, generally measured as `(G−C)/(G+C)`, is a different property.
6. **Immediate regulatory phenotype.** Genomic GC is an accumulated evolutionary outcome, not generally an acutely inducible cellular state.
7. **Thermophily or habitat preference.** These may correlate with GC but are not definitions of the trait.

Assembly contamination, incomplete recovery, untrimmed plasmids, ambiguous bases, and metagenome-bin compositional bias can all move an estimate near 66.3%; threshold-adjacent assignments should retain the assembly method and confidence interval where possible.

## 2. Current mechanistic understanding

### Mutation pressure is generally antagonistic to high GC

Recent expert synthesis continues to describe bacterial mutation as biased rather than uniform and notes the apparent paradox that genomes can be GC-rich despite a broadly GC→AT mutational bias (Horton and Taylor, published 9 November 2023) (horton2023mutationbiasand pages 1-2). Earlier synthesis similarly concluded that mutation is generally AT-biased and that an additional evolutionary force is needed to maintain intermediate- and high-GC genomes (hershberg2015mutation—theengineof pages 6-7, lassalle2015gccontentevolutionin pages 1-4).

This supports an **inhibitory**, not activating, edge from baseline AT-biased mutation pressure to the high-GC trait. It also means that merely identifying a DNA-repair gene in a high-GC genome does not establish that the gene created the composition.

### Recombination-associated gBGC is the leading broad counterforce, but bacterial evidence remains indirect

Lassalle and colleagues found higher GC in recombining genes across broad bacterial clades: significant effects occurred in 11 of 14 groups and were stronger at GC3. Their within-genome analysis found recombination–GC associations with reported `R²` values of 0.24–0.68 across 11 groups; in *Streptococcus pyogenes*, unbinned gene-level values included `R²=0.034` and `0.087`, rising to `0.60` after binning (lassalle2015gccontentevolutionin pages 4-6, lassalle2015gccontentevolutionin pages 6-9, lassalle2015gccontentevolutionin pages 11-14). Intergenic regions flanked by recombining genes were also usually GC-richer, although individual significance was weak—only 1 of 14 comparisons—with 11 of 14 effects in the predicted direction (`p=0.03`) (lassalle2015gccontentevolutionin pages 6-9).

The interpretation is that homologous recombination creates heteroduplex mismatches and a repair bias preferentially transmits G/C alleles. The expected strength depends on effective population size, recombination rate, conversion-tract length, and repair-bias intensity (lassalle2015gccontentevolutionin pages 9-11). Nevertheless, these are comparative signatures, not a bacterial perturbation proving that gBGC causes a genome to exceed 66.3%. Exceptions include *Helicobacter pylori* and members of the *Bacillus anthracis/cereus* group (lassalle2015gccontentevolutionin pages 4-6).

### 2024 development: NucS directly reshapes mutation supply in a high-GC bacterium

Dagva et al. studied the approximately 72%-GC linear chromosome of *Streptomyces ambofaciens*. Their biochemical and mutation-accumulation experiments showed that NucS cooperates with the replication clamp and cleaves G/T, G/G, and T/T mismatches by producing double-strand breaks; the authors concluded that NucS-dependent MMR eliminates G/T mismatches generated during replication (published 6 March 2024; DOI below) (dagva2024correctionofnonrandom pages 1-2).

Deleting `nucS` caused:

- a **32-fold** average increase in total mutation rate;

Showing the first 60 of 258 lines of findings; the linked file also carries the run's front matter and the prompt it was given — read the full report.

Curation history

  1. · SEEDED_FROM_METPO · seed_from_metpo

    imported from data/raw/metpo.owl (CLASS)

  2. · CURATED_CAUSAL_GRAPH · claude

    Added DOI-backed definition (derived from METPO synonym GC_>66.3) and causal graph linking strong GC-biased gene conversion to this GC bin. Documented the upstream label-vs-threshold inconsistency.

  3. · GROUND_CAUSAL_PREDICATES · claude

    Grounded 2 causal-edge predicate_id field(s) via mappings/predicate_grounding.tsv (METPO:2000202×1, rdfs:subClassOf×1).

  4. · ENRICH_CAUSAL_GRAPH · claude

    Added 4 evidence-backed generic edges (4 new nodes) from the deep-research report.

  5. · GROUND_CAUSAL_PREDICATES · claude

    Grounded 1 causal-edge predicate_id field(s) via mappings/predicate_grounding.tsv (METPO:2007401×1).

  6. · MIGRATE_MICROBE_DOMAIN_EDGES · claude

    Re-grounded 1 causal edge(s) off microbe-domain METPO predicates (1 to confers), issue 301. The previous predicates are transitively rdfs:subPropertyOf METPO:2000001, whose rdfs:domain is METPO:1000525 (microbe), so a causal-graph subject entailed that the subject IS a microbe; CausalNodeTypeEnum has no organism member, so no such edge could ever satisfy the domain. Edge directions are unchanged - this pass only relabels and re-grounds. RO:0002234 (has output) is used where the subject is an activity, since biolink gives it the domain 'biological process or activity'; the METPO replacements are proposed in proposals/metpo_traitmech_v8 and v9 and are placeholder ids until METPO mints them.