
Aggregate taxa contributions for visualization
Source:R/taxa_contribution.R
aggregate_taxa_contributions.RdCore aggregation function that bridges PICRUSt2 contribution data with differential abundance analysis results. Optionally maps ASV/OTU IDs to taxonomic names and filters to significant pathways.
Usage
aggregate_taxa_contributions(
contrib_data,
taxonomy = NULL,
tax_level = "Genus",
top_n = 10,
daa_results_df = NULL,
pathway_ids = NULL,
p_threshold = 0.05,
contribution_col = "auto"
)Arguments
- contrib_data
A data.frame from
read_contrib_fileorread_strat_file.- taxonomy
Optional data frame or matrix mapping taxon IDs to taxonomy. Taxon IDs may be stored in a standard ID column or in explicit row names. Supports QIIME2 format (semicolon-delimited taxonomy strings) or DADA2 format (separate columns for each rank).
- tax_level
Character. Taxonomic rank for aggregation. One of
"Kingdom","Phylum","Class","Order","Family","Genus","Species". Default"Genus".- top_n
Integer. Number of top taxa to keep by total contribution; remaining are lumped as "Other". Default 10.
- daa_results_df
Optional data.frame from
pathway_daa, used to filter contributions to significant pathways.- pathway_ids
Optional character vector of pathway IDs to filter. Alternative to
daa_results_df.- p_threshold
Numeric. Significance cutoff when using
daa_results_df. Default 0.05. Must be in the range (0, 1].- contribution_col
Character. Column to aggregate. Use
"auto"to select the first available column fromnorm_taxon_function_contrib,taxon_function_abun,taxon_rel_function_abun, andabundance.
Details
With daa_results_df, significant feature IDs are used as filters.
KEGG pathway IDs are expanded to member KOs only when the contribution
input is KO-level. This selects KO rows; output function_id values
remain KO IDs and are not reconstructed pathway contributions. With
pathway-level input (for example MetaCyc), IDs are matched directly.
Aggregation sums the selected contribution column within each
sample/function/taxonomic label. Percentage plots normalize those sums
afterward; aggregation itself does not convert raw abundance to fractions.
Taxonomy can be provided in two formats:
QIIME2: A column named
Taxonortaxonomycontaining semicolon-delimited strings (e.g., "k__Bacteria;p__Firmicutes;...")DADA2: Separate columns for each rank (
Kingdom,Phylum, etc.)
Contribution table identifier columns (sample, function_id,
and taxon) and requested pathway/function IDs must be non-empty and
non-missing, because R's aggregation functions otherwise omit missing keys
without preserving contribution totals. Contribution tables must not mix
gene-family-level and pathway-level identifiers in the same aggregation.
Examples
# \donttest{
# Basic usage with synthetic data
contrib <- data.frame(
sample = rep(c("S1", "S2"), each = 6),
function_id = rep(c("K00001", "K00002", "K00003"), 4),
taxon = rep(c("ASV1", "ASV2"), each = 3, times = 2),
taxon_function_abun = seq_len(12)
)
agg <- aggregate_taxa_contributions(contrib, top_n = 2)
head(agg)
# }