
Grow Glycans with Enzymes
grow_glycans_step.RdThis function simulates the action of enzymes on glycans. Think of it like a primordial soup where you put in a few glycans and enzymes, and let them interact to generate new glycans.
grow_glycans_step() performs one round of enzyme action,
while grow_glycans() performs multiple rounds.
The only difference between grow_glycans_step() and grow_glycans(n_steps = 1)
is that the latter returns the original input glycans as well.
For both, a vector of unique glycan structures is returned.
The number of glycans generated by grow_glycans() is typically exponential,
and can quickly become very large.
Therefore, it is recommended to use a small number
of steps and carefully select the enzymes.
Also, you can use the filter argument to prune the results after each round.
Usage
grow_glycans_step(glycans, enzymes)
grow_glycans(glycans, enzymes, n_steps = 5, filter = NULL)Arguments
- glycans
A
glyrepr::glycan_structure(), or a character vector of glycan structure strings supported byglyparse::auto_parse().- enzymes
A character vector of gene symbols, or a list of
enzyme()objects.- n_steps
The maximum number of rounds to perform. The actual number of rounds may be less if no new glycans can be generated.
- filter
A function to filter the generated glycans. It should have a single
glyrepr::glycan_structure()vector as input, and return a logical vector of the same length. The function will be called on the results after each round.
Value
A glyrepr::glycan_structure() vector of all unique glycans generated.
Important notes
Here are some important notes for all functions in the glyenzy package.
Applicability
Known-enzyme algorithms and enzyme information in glyenzy are applicable only to humans. Curated coverage is strongest for N-glycans and O-glycans and also includes selected glycosphingolipid headgroups and other glycan contexts. Lipid and protein aglycones are not represented, so glycolipid rules model the carbohydrate headgroup with ceramide omitted. Results may be inaccurate for unsupported glycan contexts or other species (e.g., plants, insects).
Inclusiveness
The algorithm takes an intentionally inclusive approach, assuming that all possible isoenzymes capable of catalyzing a given reaction may be involved. Therefore, results should be interpreted with caution.
For example, in humans, detection of the motif "Neu5Ac(a2-3)Gal(b1-" will return both "ST3GAL3" and "ST3GAL4". In reality, only one of them might be active, depending on factors such as tissue specificity.
Concrete glycans by default
Most functions only work for glycans containing concrete residues
(e.g., "Glc", "GalNAc"), and not for glycans with generic
residues (e.g., "Hex", "HexNAc").
Inputs with generic or mixed residues are supported where explicitly
documented, such as trace_biosynthesis() and path_biosynthesis().
Substituents
Sulfate substituents are supported. Other substituents, such as
phosphorylation and methylation, are not supported. Use
glyrepr::remove_substituents() when unsupported substituents are present.
Incomplete or non-concrete glycan structures
If the glycan structure is incomplete, partially degraded, or contains
generic or mixed residues, the result may be misleading. Glycans with a
glyrepr::get_structure_level() other than "intact", or with a
glyrepr::get_mono_type() other than "concrete", are matched with the
lenient motif matching mode in glymotif. A warning is raised because enzyme
predictions may be less reliable.
Starting points
For known-enzyme path inference:
For N-glycans, the starting structure is assumed to be "Glc(3)Man(9)GlcNAc(2)", the N-glycan precursor transferred to Asn by OST.
For O-GalNAc glycans, the starting structure is assumed to be "GalNAc(a1-".
For O-GlcNAc glycans, the starting structure is assumed to be "GlcNAc(b1-".
For O-Man glycans, the starting structure is assumed to be "Man(a1-".
For O-Fuc glycans, the starting structure is assumed to be "Fuc(a1-".
For O-Glc glycans, the starting structure is assumed to be "Glc(b1-".
For GlcCer glycans, the starting structure is assumed to be "Glc(b1-",
For GalCer glycans, the starting structure is assumed to be "Gal(b1-"
Examples
# Use `grow_glycans_step()` to build glycans step by step
glycan <- "GlcNAc(b1-2)Man(a1-3)[Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-"
glycan |>
grow_glycans_step("MGAT2") |>
grow_glycans_step("B4GALT1") |>
grow_glycans_step("ST3GAL3")
#> <glycan_structure[0]>
#> # Unique structures: 0
# Use `grow_glycans()` to simulate a primordial soup
glycans <- c(
"GlcNAc(b1-2)Man(a1-3)[Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-",
"GlcNAc(b1-2)Man(a1-3)[GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-"
)
enzymes <- c("B4GALT1", "ST3GAL3")
grow_glycans(glycans, enzymes, n_steps = 5)
#> <glycan_structure[6]>
#> [1] GlcNAc(b1-2)Man(a1-3)[Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [2] GlcNAc(b1-2)Man(a1-3)[GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [3] Gal(b1-4)GlcNAc(b1-2)Man(a1-3)[Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [4] Gal(b1-4)GlcNAc(b1-2)Man(a1-3)[GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [5] Gal(b1-4)GlcNAc(b1-2)Man(a1-6)[GlcNAc(b1-2)Man(a1-3)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [6] Gal(b1-4)GlcNAc(b1-2)Man(a1-3)[Gal(b1-4)GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> # Unique structures: 6
# Use `filter` to prune the results after each round
# Here we keep only glycans that are synthesized by MGAT2
grow_glycans(glycans, enzymes, n_steps = 5, filter = ~ have_enzyme(.x, "MGAT2"))
#> <glycan_structure[4]>
#> [1] GlcNAc(b1-2)Man(a1-3)[GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [2] Gal(b1-4)GlcNAc(b1-2)Man(a1-3)[GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [3] Gal(b1-4)GlcNAc(b1-2)Man(a1-6)[GlcNAc(b1-2)Man(a1-3)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> [4] Gal(b1-4)GlcNAc(b1-2)Man(a1-3)[Gal(b1-4)GlcNAc(b1-2)Man(a1-6)]Man(b1-4)GlcNAc(b1-4)GlcNAc(b1-
#> # Unique structures: 4