Skip to contents

Core aggregation function that bridges PICRUSt2 contribution data with differential abundance analysis results. Optionally maps ASV/OTU IDs to taxonomic names and filters to significant pathways.

Usage

aggregate_taxa_contributions(
  contrib_data,
  taxonomy = NULL,
  tax_level = "Genus",
  top_n = 10,
  daa_results_df = NULL,
  pathway_ids = NULL,
  p_threshold = 0.05,
  contribution_col = "auto"
)

Arguments

contrib_data

A data.frame from read_contrib_file or read_strat_file.

taxonomy

Optional data frame or matrix mapping taxon IDs to taxonomy. Taxon IDs may be stored in a standard ID column or in explicit row names. Supports QIIME2 format (semicolon-delimited taxonomy strings) or DADA2 format (separate columns for each rank).

tax_level

Character. Taxonomic rank for aggregation. One of "Kingdom", "Phylum", "Class", "Order", "Family", "Genus", "Species". Default "Genus".

top_n

Integer. Number of top taxa to keep by total contribution; remaining are lumped as "Other". Default 10.

daa_results_df

Optional data.frame from pathway_daa, used to filter contributions to significant pathways.

pathway_ids

Optional character vector of pathway IDs to filter. Alternative to daa_results_df.

p_threshold

Numeric. Significance cutoff when using daa_results_df. Default 0.05. Must be in the range (0, 1].

contribution_col

Character. Column to aggregate. Use "auto" to select the first available column from norm_taxon_function_contrib, taxon_function_abun, taxon_rel_function_abun, and abundance.

Value

A tidy data.frame with columns: sample, function_id, taxon_label, contribution.

Details

With daa_results_df, significant feature IDs are used as filters. KEGG pathway IDs are expanded to member KOs only when the contribution input is KO-level. This selects KO rows; output function_id values remain KO IDs and are not reconstructed pathway contributions. With pathway-level input (for example MetaCyc), IDs are matched directly. Aggregation sums the selected contribution column within each sample/function/taxonomic label. Percentage plots normalize those sums afterward; aggregation itself does not convert raw abundance to fractions.

Taxonomy can be provided in two formats:

  • QIIME2: A column named Taxon or taxonomy containing semicolon-delimited strings (e.g., "k__Bacteria;p__Firmicutes;...")

  • DADA2: Separate columns for each rank (Kingdom, Phylum, etc.)

Contribution table identifier columns (sample, function_id, and taxon) and requested pathway/function IDs must be non-empty and non-missing, because R's aggregation functions otherwise omit missing keys without preserving contribution totals. Contribution tables must not mix gene-family-level and pathway-level identifiers in the same aggregation.

Examples

# \donttest{
# Basic usage with synthetic data
contrib <- data.frame(
  sample = rep(c("S1", "S2"), each = 6),
  function_id = rep(c("K00001", "K00002", "K00003"), 4),
  taxon = rep(c("ASV1", "ASV2"), each = 3, times = 2),
  taxon_function_abun = seq_len(12)
)
agg <- aggregate_taxa_contributions(contrib, top_n = 2)
head(agg)
# }