Skip to contents

This function integrates pathway name/description annotations, ten of the most advanced differential abundance (DA) methods, and visualization of DA results.

Usage

ggpicrust2(
  file = NULL,
  data = NULL,
  metadata,
  group,
  pathway,
  daa_method = "ALDEx2",
  ko_to_kegg = FALSE,
  filter_for_prokaryotes = TRUE,
  p_adjust_method = "BH",
  order = "group",
  p_values_bar = TRUE,
  x_lab = NULL,
  select = NULL,
  reference = NULL,
  colors = NULL,
  p_values_threshold = 0.05,
  p.adjust = NULL
)

Arguments

file

A character string representing the file path of the input file containing KO abundance data in picrust2 export format. The input file should have KO identifiers in the first column and sample identifiers in the first row. The remaining cells should contain the abundance values for each KO-sample pair.

data

An optional data.frame or matrix containing KO/pathway abundance data. Data frames may use the same format as the input file, with feature identifiers in the first non-numeric column and samples in the remaining columns. Already-normalized matrix/data.frame inputs with feature identifiers in row names and samples in columns are also accepted. If provided, the function will use this data instead of reading from the file. By default, this parameter is set to NULL.

metadata

A tibble, consisting of sample information

group

A character, name of the group

pathway

A character, consisting of "EC", "KO", "MetaCyc"

daa_method

a character specifying the method for differential abundance analysis, default is "ALDEx2", choices are: - "ALDEx2": ANOVA-Like Differential Expression tool for high throughput sequencing data - "DESeq2": Differential expression analysis based on the negative binomial distribution using DESeq2 - "edgeR": Exact test for differences between two groups of negative-binomially distributed counts using edgeR - "limma voom": Limma-voom framework for the analysis of RNA-seq data - "metagenomeSeq": Fit logistic regression models to test for differential abundance between groups using metagenomeSeq - "LinDA": Linear models for differential abundance analysis of microbiome compositional data - "Maaslin2": Multivariate Association with Linear Models (MaAsLin2) for differential abundance analysis - "Lefser": Linear discriminant analysis effect size using the lefser package

ko_to_kegg

Logical or logical-like string controlling conversion of KO abundance to KEGG pathway abundance.

filter_for_prokaryotes

Logical. If TRUE (default), filters out KEGG pathways that are specific to eukaryotes (e.g., human diseases, organismal systems) when ko_to_kegg = TRUE. Set to FALSE to include all KEGG pathways.

p_adjust_method

A character specifying the method for p-value adjustment, default is "BH".

order

A character to control the order of the main plot rows

p_values_bar

A character to control if the main plot has the p_values bar

x_lab

A character to control the x-axis label name, you can choose from "feature","pathway_name" and "description"

select

A vector consisting of pathway names to be selected

reference

A character, a reference group level for several DA methods

colors

A vector consisting of colors number

p_values_threshold

A numeric value specifying the threshold for statistical significance of differential abundance. Pathways with adjusted p-values below this threshold will be displayed in the plot. Default is 0.05. Must be in the range (0, 1].

p.adjust

Deprecated alias for p_adjust_method. Do not supply both parameters with different values.

Value

A list containing:

  • Numbered elements (1, 2, ...): Sub-lists for each DA method, each containing:

    • plot: A ggplot2 error bar plot visualizing the differential abundance results

    • results: A data frame of differential abundance results for that method

  • abundance: The processed abundance data (KEGG pathway or original) for downstream analysis

  • metadata: The metadata data frame

  • group: The group variable name used in the analysis

  • daa_results_df: The complete annotated DAA results data frame

  • ko_to_kegg: Logical indicating whether KO to KEGG conversion was performed

These additional fields allow seamless integration with pathway_pca and pathway_heatmap for further visualization without re-preparing data.

Examples

if (FALSE) { # \dontrun{
# Requires MicrobiomeStat and KEGGREST; KEGG annotation uses the internet.
data("ko_abundance")
data("metadata")
results <- ggpicrust2(
  data = ko_abundance, metadata = metadata, group = "Environment",
  pathway = "KO", daa_method = "LinDA", ko_to_kegg = TRUE,
  x_lab = "pathway_name"
)
results[[1]]$plot
head(results[[1]]$results)

# Reuse the aligned abundance and metadata for exploratory plots.
pathway_pca(results$abundance, results$metadata, results$group)
sig_features <- unique(results$daa_results_df$feature[
  !is.na(results$daa_results_df$p_adjust) &
    results$daa_results_df$p_adjust < 0.05
])
if (length(sig_features) > 0) {
  pathway_heatmap(results$abundance[sig_features, , drop = FALSE],
                  results$metadata, results$group)
}
# For your own files, replace data = ko_abundance with file = "your_file.tsv"
# and supply matching metadata. See vignette("using_ggpicrust2").
} # }