GC high
METPO:1000432 · CLASS · REVIEWED
A GC-content phenotype with genome-wide GC composition at or below approximately 42.65% (the METPO `GC_<=42.65` bin; note that the upstream label 'high' does not match this numeric threshold, but the synonym is preserved as the authoritative bin definition).
GC-high (METPO ≤42.65%) low-GC bin
Edge evidence
-
AT-biased mutation pressure
confers
GC high
METPO:2007700AT-biased mutation pressure produces low genome-wide GC composition.
-
DOI:10.1186/1471-2148-10-374mutation bias
-
-
GC high
is a
GC content
rdfs:subClassOfGC high is a quantitative bin of the GC-content phenotype.
-
DOI:10.1038/nrg2358GC content
-
-
Cytosine deamination
contributes to
GC high
RO:0002326Cytosine deamination drives AT-enriching G->A transitions lowering genome-wide GC.
-
DOI:10.1038/s41467-026-71228-y
-
-
Loss of uracil-DNA glycosylase (UDG family)
causally promotes
GC high
Loss of UDG repair leaves deaminated cytosine unrepaired, promoting GC erosion.
-
DOI:10.1038/s41467-026-71228-y
-
-
Loss of BER glycosylases (Tag/AlkA/MPG/Nei)
causally promotes
GC high
Concerted loss of BER glycosylases biases mutation spectra toward GC-eroding changes.
-
DOI:10.1038/s41467-026-71228-y
-
-
Loss of MutT nucleotide-pool sanitization
causally promotes
GC high
Loss of MutT nucleotide-pool sanitization contributes to GC-eroding mutational dysregulation.
-
DOI:10.1038/s41467-026-71228-y
-
-
Hypermutator phenotype
causally promotes
GC high
Hypermutator state provides a mechanistic path to genome-wide AT enrichment.
-
DOI:10.1038/s41467-026-71228-y
-
-
Third-codon-position AT enrichment
contributes to
GC high
RO:0002326AT-rich substitution at third codon positions is a principal driver of decreased GC.
-
DOI:10.1038/s41467-026-71228-y
-
-
Purifying selection and biased gene conversion
opposes
GC high
Purifying selection and biased gene conversion counteract AT bias, resisting GC decline.
-
DOI:10.63635/mrj.v1i4.188
-
Provenance
- Source
- METPO (2025-11-25)
- Definition source
- DOI:10.1038/nrg2358
Parent traits (1)
Synonyms (1)
- GC_<=42.65
kg-microbe context
Matched 1 kg-microbe node via direct_metpo.
METPO:1000432[+0.921, +1.707, +0.625, +3.297, …]
Nearest neighbors in embedding space
- morphology cell length large 0.471
- environment temperature delta mid2 0.385
- morphology cell width very small 0.383
- environment pH range mid2 0.382
- environment pH range mid3 0.380
- environment temperature range mid2 0.380
- environment pH range mid1 0.376
- environment temperature range mid4 0.372
Deep research
# Curation report: microbial genomic low-GC phenotype
## Executive summary
The trait identifier must be quoted exactly as **METPO:1000432**. Despite its upstream label, **“GC high,”** the authoritative synonym and numerical definition describe the opposite phenotype: **whole-genome GC content ≤ approximately 42.65% (`GC_<=42.65`)**. For `data/traits/genomics/gc_high.yaml`, the numeric bin should control interpretation, while the misleading label should be retained only as provenance.
This is a **genome-composition class**, not a physiological activity. The strongest causal route supported by experiments is:
> spontaneous cytosine deamination and guanine oxidation → GC-to-AT/TA substitutions; loss of the corresponding repair capacity amplifies these substitutions → long-term decline in genome-wide GC.
The broader literature supports a systems-level model in which the composition of DNA replication and repair (DRR) machinery, phylogenetic history, mutation bias, recombination-associated GC-biased gene conversion, drift, and selection jointly determine genomic GC. Direct environmental adaptation to low GC is not established as a universal mechanism.
## 1. Trait scope and current understanding
### 1.1 Operational definition
**Recommended curation definition:** “A microbial genome-composition phenotype in which G+C bases constitute no more than approximately 42.65% of the complete or representative genome sequence.”
The phenotype should be calculated as:
\[
GC\% = 100\times\frac{G+C}{A+T+G+C}
\]
preferably over a complete, high-quality whole-genome assembly. The bin is somewhat stricter than the broad “low-GC” grouping used in recent comparative work, where most low-mode genomes occur below 45%. In 11,083 representative bacterial genomes, GC ranged from about 16% to 77% and was bimodal, with most genomes below 45% or above 60%; more than 60% of variance was explained at phylum level, with Blomberg’s K=1.47 and Pagel’s λ=0.998. Thus, 42.65% is an ontology-specific discretization, not a universal biological breakpoint (teng2023genomiclegaciesof pages 2-5).
### 1.2 Boundaries and nearby traits
The trait is **not equivalent to**:
- **GC3:** GC fraction at third codon positions. GC3 is especially responsive to synonymous substitution and codon usage and can differ substantially from whole-genome GC. Recombination studies often analyze GC3 rather than genomic GC (lassalle2015gccontentevolutionin pages 11-14).
- **Coding-sequence GC, noncoding GC, or local GC windows:** these can identify islands, horizontally transferred regions, or strand effects but do not alone establish the whole-genome bin. Teng et al. explicitly separated whole-genome GC, coding GC, noncoding GC, amino-acid-contributed GC, and codon-contributed GC (teng2023genomiclegaciesof pages 2-5).
- **GC skew:** strand asymmetry such as `(G−C)/(G+C)`; this concerns replication/transcription asymmetry rather than total composition. Whole bacterial genomes can be compositionally homogeneous while retaining strand-specific biases (lind2008wholegenomemutationalbiases pages 1-1).
- **An AT-biased mutation spectrum:** mutation bias is an upstream process. A currently high-GC genome may have AT-biased new mutations because equilibrium composition changes over long evolutionary periods (teng2023genomiclegaciesof pages 8-10, lassalle2015gccontentevolutionin pages 11-14).
- **Genome reduction:** reduced genomes are frequently AT-rich, especially in endosymbionts, but genome size and GC percentage are separate traits. Neither implies the other universally.
- **Low-GC Gram-positive bacteria:** an historical taxonomic description, not a mechanistic or phylogenetically exclusive class. Low-GC clades occur in multiple bacterial groups (teng2023genomiclegaciesof pages 2-5).
### 1.3 Expert synthesis
Recent authoritative analysis favors **indirect evolution through replication/repair systems and historical contingency**, rather than a single adaptive advantage of low GC. A phylogenetically informed model based on 217 DRR-related KEGG orthologs explained 88% of observed GC variance (multiple correlation 0.94); however, this is predictive comparative evidence, not proof that each correlated gene causes the phenotype (teng2023genomiclegaciesof pages 2-5). Figure 3 of that study shows the model fit and opposing associations of DnaE2 and MutS2, as well as pathway-level correlations involving BER, MMR, replication, recombination, and translesion synthesis (teng2023genomiclegaciesof media 5c6ce460).
## 2. Candidate graph nodes
### Trait and measurement nodes
- **Low whole-genome GC content:** **METPO:1000432**
- Parent trait: **METPO:1000127**
- Whole-genome GC percentage — label-only assay/measurement node
- GC3, coding-sequence GC, noncoding-sequence GC, local GC, and GC skew — label-only boundary/measurement nodes
### Genes, proteins, and complexes
- **MutM/Fpg DNA glycosylase**, **MutY adenine glycosylase** — repair oxidized guanine-associated lesions
- **Ung** and **Mug** uracil-DNA glycosylases — remove uracil arising from cytosine deamination
- **Vsr endonuclease** — very-short-patch repair of G:T mismatches
- **MutS/MutL mismatch-repair system** — canonical MMR; direction of compositional effect is taxon dependent
- **MutS2** — MutS homologue; do not conflate automatically with canonical MutS-directed MMR
- **NucS/EndoMS** — noncanonical mismatch-repair endonuclease in certain archaea and actinobacteria; potentially relevant but presently not supported as a universal low-GC determinant
- **DnaE/Pol III α**, **PolC**, **DnaE2**, **DinB/Pol IV**, and **Pol V** — replicative or error-prone/translesion polymerases
- **RecA/RuvC and homologous-recombination machinery** — candidates connecting recombination and gene conversion
Curation history
-
·
SEEDED_FROM_METPO · seed_from_metpo
imported from data/raw/metpo.owl (CLASS)
-
·
CURATED_CAUSAL_GRAPH · claude
Added DOI-backed definition (derived from METPO synonym GC_<=42.65) and causal graph linking AT-biased mutation pressure to this GC bin. Documented the upstream label-vs-threshold inconsistency.
-
·
GROUND_CAUSAL_PREDICATES · claude
Grounded 2 causal-edge predicate_id field(s) via mappings/predicate_grounding.tsv (METPO:2000202×1, rdfs:subClassOf×1).
-
·
ENRICH_CAUSAL_GRAPH · claude
Added 7 evidence-backed generic edges (7 new nodes) from the deep-research report.
-
·
GROUND_CAUSAL_PREDICATES · claude
Grounded 2 causal-edge predicate_id field(s) via mappings/predicate_grounding.tsv (RO:0002326×2).
-
·
MIGRATE_MICROBE_DOMAIN_EDGES · claude
Re-grounded 1 causal edge(s) off microbe-domain METPO predicates (1 to confers), issue 301. The previous predicates are transitively rdfs:subPropertyOf METPO:2000001, whose rdfs:domain is METPO:1000525 (microbe), so a causal-graph subject entailed that the subject IS a microbe; CausalNodeTypeEnum has no organism member, so no such edge could ever satisfy the domain. Edge directions are unchanged - this pass only relabels and re-grounds. RO:0002234 (has output) is used where the subject is an activity, since biolink gives it the domain 'biological process or activity'; the METPO replacements are proposed in proposals/metpo_traitmech_v8 and v9 and are placeholder ids until METPO mints them.