
This function integrates pathway name/description annotations, ten of the most advanced differential abundance (DA) methods, and visualization of DA results.
Source:R/ggpicrust2.R
ggpicrust2.RdThis function integrates pathway name/description annotations, ten of the most advanced differential abundance (DA) methods, and visualization of DA results.
Usage
ggpicrust2(
file = NULL,
data = NULL,
metadata,
group,
pathway,
daa_method = "ALDEx2",
ko_to_kegg = FALSE,
filter_for_prokaryotes = TRUE,
p_adjust_method = "BH",
order = "group",
p_values_bar = TRUE,
x_lab = NULL,
select = NULL,
reference = NULL,
colors = NULL,
p_values_threshold = 0.05,
p.adjust = NULL
)Arguments
- file
A character string representing the file path of the input file containing KO abundance data in picrust2 export format. The input file should have KO identifiers in the first column and sample identifiers in the first row. The remaining cells should contain the abundance values for each KO-sample pair.
- data
An optional data.frame or matrix containing KO/pathway abundance data. Data frames may use the same format as the input file, with feature identifiers in the first non-numeric column and samples in the remaining columns. Already-normalized matrix/data.frame inputs with feature identifiers in row names and samples in columns are also accepted. If provided, the function will use this data instead of reading from the file. By default, this parameter is set to NULL.
- metadata
A tibble, consisting of sample information
- group
A character, name of the group
- pathway
A character, consisting of "EC", "KO", "MetaCyc"
- daa_method
a character specifying the method for differential abundance analysis, default is "ALDEx2", choices are: - "ALDEx2": ANOVA-Like Differential Expression tool for high throughput sequencing data - "DESeq2": Differential expression analysis based on the negative binomial distribution using DESeq2 - "edgeR": Exact test for differences between two groups of negative-binomially distributed counts using edgeR - "limma voom": Limma-voom framework for the analysis of RNA-seq data - "metagenomeSeq": Fit logistic regression models to test for differential abundance between groups using metagenomeSeq - "LinDA": Linear models for differential abundance analysis of microbiome compositional data - "Maaslin2": Multivariate Association with Linear Models (MaAsLin2) for differential abundance analysis - "Lefser": Linear discriminant analysis effect size using the lefser package
- ko_to_kegg
Logical or logical-like string controlling conversion of KO abundance to KEGG pathway abundance.
- filter_for_prokaryotes
Logical. If TRUE (default), filters out KEGG pathways that are specific to eukaryotes (e.g., human diseases, organismal systems) when ko_to_kegg = TRUE. Set to FALSE to include all KEGG pathways.
- p_adjust_method
A character specifying the method for p-value adjustment, default is "BH".
- order
A character to control the order of the main plot rows
- p_values_bar
A character to control if the main plot has the p_values bar
- x_lab
A character to control the x-axis label name, you can choose from "feature","pathway_name" and "description"
- select
A vector consisting of pathway names to be selected
- reference
A character, a reference group level for several DA methods
- colors
A vector consisting of colors number
- p_values_threshold
A numeric value specifying the threshold for statistical significance of differential abundance. Pathways with adjusted p-values below this threshold will be displayed in the plot. Default is 0.05. Must be in the range (0, 1].
- p.adjust
Deprecated alias for
p_adjust_method. Do not supply both parameters with different values.
Value
A list containing:
Numbered elements (1, 2, ...): Sub-lists for each DA method, each containing:
plot: A ggplot2 error bar plot visualizing the differential abundance resultsresults: A data frame of differential abundance results for that method
abundance: The processed abundance data (KEGG pathway or original) for downstream analysismetadata: The metadata data framegroup: The group variable name used in the analysisdaa_results_df: The complete annotated DAA results data frameko_to_kegg: Logical indicating whether KO to KEGG conversion was performed
These additional fields allow seamless integration with pathway_pca and
pathway_heatmap for further visualization without re-preparing data.
Examples
if (FALSE) { # \dontrun{
# Requires MicrobiomeStat and KEGGREST; KEGG annotation uses the internet.
data("ko_abundance")
data("metadata")
results <- ggpicrust2(
data = ko_abundance, metadata = metadata, group = "Environment",
pathway = "KO", daa_method = "LinDA", ko_to_kegg = TRUE,
x_lab = "pathway_name"
)
results[[1]]$plot
head(results[[1]]$results)
# Reuse the aligned abundance and metadata for exploratory plots.
pathway_pca(results$abundance, results$metadata, results$group)
sig_features <- unique(results$daa_results_df$feature[
!is.na(results$daa_results_df$p_adjust) &
results$daa_results_df$p_adjust < 0.05
])
if (length(sig_features) > 0) {
pathway_heatmap(results$abundance[sig_features, , drop = FALSE],
results$metadata, results$group)
}
# For your own files, replace data = ko_abundance with file = "your_file.tsv"
# and supply matching metadata. See vignette("using_ggpicrust2").
} # }