GC mid1
METPO:1000430 · CLASS · REVIEWED
A GC-content phenotype with genome-wide GC composition above approximately 66.3% (the METPO `GC_>66.3` bin; note that the upstream label 'mid1' does not match this high-end numeric threshold, but the synonym is preserved as the authoritative bin definition).
GC-mid1 (METPO >66.3%) high-GC bin
Edge evidence
-
strong GC-biased gene conversion
confers
GC mid1
METPO:2007700Strong GC-biased gene conversion yields high genome-wide GC composition.
-
DOI:10.1186/1471-2148-10-374GC-biased gene conversion
-
-
GC mid1
is a
GC content
rdfs:subClassOfGC mid1 is a quantitative bin of the GC-content phenotype.
-
DOI:10.1038/nrg2358GC content
-
-
NHEJ pathway (Ku, LigD)
positively correlated with
GC mid1
Ku/NHEJ presence is strongly associated with elevated genomic GC content.
-
DOI:10.1371/journal.pgen.1008493
-
-
high double-strand break rate
selects for
GC mid1
METPO:2007401Higher DSB formation rate selects for increased GC content relative to genomic background.
-
DOI:10.1371/journal.pgen.1008493
-
-
DNA replication and repair (DRR) system composition
strongly correlated with
GC mid1
DRR-system KEGG-ortholog composition explains a large fraction of genomic GC variance.
-
DOI:10.1128/spectrum.02145-22
-
-
GC mid1
may increase
NHEJ end-joining efficiency
High GC may increase NHEJ repair efficiency by stabilizing short overhangs/microhomologies via extra hydrogen bonds.
-
DOI:10.1371/journal.pgen.1008493
-
Provenance
- Source
- METPO (2025-11-25)
- Definition source
- DOI:10.1038/nrg2358
Parent traits (1)
Synonyms (1)
- GC_>66.3
kg-microbe context
Matched 1 kg-microbe node via direct_metpo.
METPO:1000430[-2.804, -2.753, -0.396, +5.171, …]
Nearest neighbors in embedding space
- environment temperature optimum mid2 0.568
- environment NaCl range low 0.459
- environment pH range low 0.429
- environment pH range mid1 0.422
- environment NaCl range mid1 0.420
- environment pH range mid3 0.417
- environment temperature range mid1 0.413
- environment temperature range mid2 0.411
Deep research
# Curation report: **GC mid1** (`METPO:1000430`)
## Executive curation recommendation
`METPO:1000430` should be modeled as an **assay-derived whole-genome nucleotide-composition class**, not as a metabolic pathway, physiological capacity, or environmental preference. The operational phenotype is:
\[
GC_w=\frac{G+C}{A+T+G+C}>0.663\;\text{(approximately)}.
\]
The authoritative synonym `GC_>66.3` should govern interpretation. The upstream label **“GC mid1” is misleading**, because 66.3% is a high-end bin; preserve it only as the supplied label and add a curation note. Across bacteria, reported genomic GC content spans roughly <25% to 75%, while one broader prokaryotic compilation reported 8–75%, placing the threshold near the extreme high end rather than the middle (hershberg2015mutation—theengineof pages 6-7, hu2022apositivecorrelation pages 1-2).
The strongest defensible causal architecture is:
**DNA replication errors → nucleotide-specific mismatches → proofreading/MMR-dependent mutation spectrum → long-term substitution supply**, opposed or overridden by **recombination-associated GC-biased gene conversion (gBGC) and possibly selection → preferential persistence/fixation of G/C alleles → elevated whole-genome GC → `METPO:1000430`.**
However, no retrieved experiment directly drove a lineage across the **66.3% threshold**. Therefore, molecular repair edges can be curated strongly at the mutation-spectrum level, whereas final edges into `METPO:1000430` require evolutionary-timescale and uncertainty qualifiers.
## 1. Trait scope and boundary cases
### Included
- **Unit:** preferably a complete or sufficiently unbiased draft chromosome/genome.
- **Observation:** percentage of guanine plus cytosine among called genomic DNA bases.
- **Classification:** positive when whole-genome GC is above approximately 66.3%.
- **Examples:** *Deinococcus radiodurans* at 66.61% is just above the boundary; *Streptomyces* genomes at approximately 72% are clearly within the bin (long2018specificityofthe pages 1-2, dagva2024correctionofnonrandom pages 1-2).
### Excluded or separately modeled
1. **GC3 or fourfold-degenerate-site GC.** These are informative about weakly selected substitutions but are not equivalent to whole-genome GC.
2. **Coding-region, core-genome, accessory-genome, plasmid, or intergenic GC.** These can differ materially within one organism.
3. **rRNA/tRNA GC.** Structural-RNA GC may respond to temperature differently from whole-genome GC; older work found structural-RNA associations even when whole-genome associations were absent (hu2022apositivecorrelation pages 1-2).
4. **Local GC islands or horizontally transferred segments.** A local high-GC region does not establish the genome-wide phenotype.
5. **GC skew.** Strand asymmetry, generally measured as `(G−C)/(G+C)`, is a different property.
6. **Immediate regulatory phenotype.** Genomic GC is an accumulated evolutionary outcome, not generally an acutely inducible cellular state.
7. **Thermophily or habitat preference.** These may correlate with GC but are not definitions of the trait.
Assembly contamination, incomplete recovery, untrimmed plasmids, ambiguous bases, and metagenome-bin compositional bias can all move an estimate near 66.3%; threshold-adjacent assignments should retain the assembly method and confidence interval where possible.
## 2. Current mechanistic understanding
### Mutation pressure is generally antagonistic to high GC
Recent expert synthesis continues to describe bacterial mutation as biased rather than uniform and notes the apparent paradox that genomes can be GC-rich despite a broadly GC→AT mutational bias (Horton and Taylor, published 9 November 2023) (horton2023mutationbiasand pages 1-2). Earlier synthesis similarly concluded that mutation is generally AT-biased and that an additional evolutionary force is needed to maintain intermediate- and high-GC genomes (hershberg2015mutation—theengineof pages 6-7, lassalle2015gccontentevolutionin pages 1-4).
This supports an **inhibitory**, not activating, edge from baseline AT-biased mutation pressure to the high-GC trait. It also means that merely identifying a DNA-repair gene in a high-GC genome does not establish that the gene created the composition.
### Recombination-associated gBGC is the leading broad counterforce, but bacterial evidence remains indirect
Lassalle and colleagues found higher GC in recombining genes across broad bacterial clades: significant effects occurred in 11 of 14 groups and were stronger at GC3. Their within-genome analysis found recombination–GC associations with reported `R²` values of 0.24–0.68 across 11 groups; in *Streptococcus pyogenes*, unbinned gene-level values included `R²=0.034` and `0.087`, rising to `0.60` after binning (lassalle2015gccontentevolutionin pages 4-6, lassalle2015gccontentevolutionin pages 6-9, lassalle2015gccontentevolutionin pages 11-14). Intergenic regions flanked by recombining genes were also usually GC-richer, although individual significance was weak—only 1 of 14 comparisons—with 11 of 14 effects in the predicted direction (`p=0.03`) (lassalle2015gccontentevolutionin pages 6-9).
The interpretation is that homologous recombination creates heteroduplex mismatches and a repair bias preferentially transmits G/C alleles. The expected strength depends on effective population size, recombination rate, conversion-tract length, and repair-bias intensity (lassalle2015gccontentevolutionin pages 9-11). Nevertheless, these are comparative signatures, not a bacterial perturbation proving that gBGC causes a genome to exceed 66.3%. Exceptions include *Helicobacter pylori* and members of the *Bacillus anthracis/cereus* group (lassalle2015gccontentevolutionin pages 4-6).
### 2024 development: NucS directly reshapes mutation supply in a high-GC bacterium
Dagva et al. studied the approximately 72%-GC linear chromosome of *Streptomyces ambofaciens*. Their biochemical and mutation-accumulation experiments showed that NucS cooperates with the replication clamp and cleaves G/T, G/G, and T/T mismatches by producing double-strand breaks; the authors concluded that NucS-dependent MMR eliminates G/T mismatches generated during replication (published 6 March 2024; DOI below) (dagva2024correctionofnonrandom pages 1-2).
Deleting `nucS` caused:
- a **32-fold** average increase in total mutation rate;
Curation history
-
·
SEEDED_FROM_METPO · seed_from_metpo
imported from data/raw/metpo.owl (CLASS)
-
·
CURATED_CAUSAL_GRAPH · claude
Added DOI-backed definition (derived from METPO synonym GC_>66.3) and causal graph linking strong GC-biased gene conversion to this GC bin. Documented the upstream label-vs-threshold inconsistency.
-
·
GROUND_CAUSAL_PREDICATES · claude
Grounded 2 causal-edge predicate_id field(s) via mappings/predicate_grounding.tsv (METPO:2000202×1, rdfs:subClassOf×1).
-
·
ENRICH_CAUSAL_GRAPH · claude
Added 4 evidence-backed generic edges (4 new nodes) from the deep-research report.
-
·
GROUND_CAUSAL_PREDICATES · claude
Grounded 1 causal-edge predicate_id field(s) via mappings/predicate_grounding.tsv (METPO:2007401×1).
-
·
MIGRATE_MICROBE_DOMAIN_EDGES · claude
Re-grounded 1 causal edge(s) off microbe-domain METPO predicates (1 to confers), issue 301. The previous predicates are transitively rdfs:subPropertyOf METPO:2000001, whose rdfs:domain is METPO:1000525 (microbe), so a causal-graph subject entailed that the subject IS a microbe; CausalNodeTypeEnum has no organism member, so no such edge could ever satisfy the domain. Edge directions are unchanged - this pass only relabels and re-grounds. RO:0002234 (has output) is used where the subject is an activity, since biolink gives it the domain 'biological process or activity'; the METPO replacements are proposed in proposals/metpo_traitmech_v8 and v9 and are placeholder ids until METPO mints them.