Skip to contents

The function pathway_errorbar() is used to visualize the results of functional pathway differential abundance analysis as error bar plots.

Usage

pathway_errorbar(
  abundance,
  daa_results_df,
  Group,
  ko_to_kegg = FALSE,
  p_values_threshold = 0.05,
  order = "group",
  select = NULL,
  p_value_bar = TRUE,
  colors = NULL,
  x_lab = NULL,
  log2_fold_change_color = "#87ceeb",
  max_features = 30,
  color_theme = "default",
  pathway_class_colors = NULL,
  smart_colors = FALSE,
  accessibility_mode = FALSE,
  legend_position = "top",
  legend_direction = "horizontal",
  legend_title = NULL,
  legend_title_size = 12,
  legend_text_size = 10,
  legend_key_size = 0.8,
  legend_ncol = NULL,
  legend_nrow = NULL,
  pvalue_format = "numeric",
  pvalue_stars = TRUE,
  pvalue_colors = FALSE,
  pvalue_size = "auto",
  pvalue_angle = 0,
  pvalue_thresholds = c(0.001, 0.01, 0.05),
  pvalue_star_symbols = c("***", "**", "*"),
  pathway_class_text_size = "auto",
  pathway_class_text_color = "black",
  pathway_class_text_face = "bold",
  pathway_class_text_angle = 0,
  pathway_class_position = "right",
  pathway_names_text_size = "auto"
)

Arguments

abundance

A data frame with row names representing pathways and column names representing samples. Values must be non-negative abundances. The function normalizes each sample by its total over all supplied features before computing group means and standard deviations. Supply the full feature matrix and use select to limit displayed rows; pre-filtering the matrix changes the denominator. Data frames may also provide a leading non-numeric feature ID column (for example #NAME, feature, or pathway); it is converted to row names before sample-count validation.

daa_results_df

A data frame containing the results of the differential abundance analysis of the pathways, generated by the pathway_daa function. x_lab should be a column name of daa_results_df. Within the selected method and group pair, feature identifiers must be unique.

Group

A vector assigning each sample to a group, not a metadata column name. Prefer setNames(metadata$Environment, metadata$sample_name) with your grouping and sample-ID columns. The groups are used to color the samples in the figure. Values must be non-missing and non-empty for all abundance columns, and must include the DAA result's group1 and group2 labels. A named vector is aligned to abundance column names; otherwise the vector is interpreted in abundance-column order.

ko_to_kegg

A logical parameter indicating whether there was a convertion that convert ko abundance to kegg abundance.

p_values_threshold

A numeric parameter specifying the threshold for statistical significance of differential abundance. Pathways with adjusted p-values below this threshold will be considered significant. Must be in the range (0, 1].

order

A single character string controlling the ordering of the rows in the figure. The options are: "p_values" (order by p-values), "name" (order by pathway name), "group" (order by the group with the highest mean relative abundance), or "pathway_class" (order by the pathway category).

select

A vector of pathway names to be included in the figure. This can be used to limit the number of pathways displayed. If NULL, all pathways will be displayed.

p_value_bar

A logical parameter indicating whether to display a bar showing the p-value threshold for significance. If TRUE, the bar will be displayed.

colors

A vector of colors to be used to represent the groups in the figure. Each color corresponds to a group. If NULL, colors will be selected based on the color_theme.

x_lab

A character string to be used as the x-axis label in the figure. The default value is "description" for KOs'descriptions and "pathway_name" for KEGG pathway names.

log2_fold_change_color

A character string specifying the color for log2 fold change bars. Default is "#87ceeb" (light blue). Can also be "auto" to use theme-based colors.

max_features

A numeric parameter specifying the maximum number of features to display before issuing a warning. Default is 30. Set to a higher value to display more features, or Inf to disable the limit entirely.

color_theme

A character string specifying the color theme to use. Options include: "default", "nature", "science", "cell", "nejm", "lancet", "colorblind_friendly", "viridis", "plasma", "minimal", "high_contrast", "pastel", "bold". Default is "default".

pathway_class_colors

A vector of colors for pathway class annotations. If NULL, colors will be selected from the theme.

smart_colors

A logical parameter indicating whether to use intelligent color selection based on data characteristics. Default is FALSE.

accessibility_mode

A logical parameter indicating whether to use accessibility-friendly colors. Default is FALSE.

legend_position

A character string specifying legend position. Options: "top", "bottom", "left", "right", "none". Default is "top".

legend_direction

A character string specifying legend direction. Options: "horizontal", "vertical". Default is "horizontal".

legend_title

A character string for legend title. If NULL, no title is displayed.

legend_title_size

A numeric value specifying legend title font size. Default is 12.

legend_text_size

A numeric value specifying legend text font size. Default is 10.

legend_key_size

A numeric value specifying legend key size in cm. Default is 0.8.

legend_ncol

A numeric value specifying number of columns in legend. If NULL, automatic layout is used.

legend_nrow

A numeric value specifying number of rows in legend. If NULL, automatic layout is used.

pvalue_format

A character string specifying p-value format. Options: "numeric", "scientific", "smart", "stars_only", "combined". Default is "numeric".

pvalue_stars

A logical parameter indicating whether to display significance stars. Default is TRUE.

pvalue_colors

A logical parameter indicating whether to use color coding for significance levels. Default is FALSE.

pvalue_size

A numeric value or "auto" for p-value text size. Default is "auto".

pvalue_angle

A numeric value specifying p-value text angle in degrees. Default is 0.

pvalue_thresholds

A numeric vector of significance thresholds. Default is c(0.001, 0.01, 0.05).

pvalue_star_symbols

A character vector of star symbols for significance levels. Default is c("***", "**", "*").

pathway_class_text_size

A numeric value or "auto" for pathway class text size. Default is "auto".

pathway_class_text_color

A character string for pathway class text color. Use "auto" for theme-based color. Default is "black".

pathway_class_text_face

A character string for pathway class text face. Options: "plain", "bold", "italic". Default is "bold".

pathway_class_text_angle

A numeric value specifying pathway class text angle in degrees. Default is 0.

pathway_class_position

A character string specifying pathway class position. Options: "left", "right", "none". Default is "right".

pathway_names_text_size

A numeric value or "auto" for pathway names (y-axis labels) text size. Default is "auto".

Value

A ggplot2 (patchwork) plot showing the differential abundance results, or NULL with a warning when all pathway annotations are missing. Check for significant results before plotting; an empty significant set is a valid analysis outcome.

Details

The abundance panel contains descriptive means and standard deviations of per-sample relative abundance. If DAA supplies log2_fold_change, the effect panel preserves that model estimate. Otherwise it uses a pseudocount-stabilized log2 ratio of group mean relative abundances. A model coefficient need not equal the descriptive abundance ratio.

Examples

if (FALSE) { # \dontrun{
# Requires the optional MicrobiomeStat and KEGGREST packages.
data("ko_abundance")
data("metadata")
kegg_abundance <- ko2kegg_abundance(data = ko_abundance)
sample_groups <- setNames(metadata$Environment, metadata$sample_name)

daa_results_df <- pathway_daa(
  abundance = kegg_abundance,
  metadata = metadata,
  group = "Environment",
  daa_method = "LinDA"
)
# For ALDEx2, select one test (e.g. ALDEx2_Welch's t test) first.
daa_annotated_results_df <- pathway_annotation(
  pathway = "KO",
  daa_results_df = daa_results_df,
  ko_to_kegg = TRUE
)

if (any(daa_annotated_results_df$p_adjust < 0.05, na.rm = TRUE)) {
  p <- pathway_errorbar(
    abundance = kegg_abundance,
    daa_results_df = daa_annotated_results_df,
    Group = sample_groups,
    ko_to_kegg = TRUE,
    order = "pathway_class",
    x_lab = "pathway_name"
  )
}
# See vignette("using_ggpicrust2") for the complete stepwise workflow.
} # }