Skip to contents

Parses PICRUSt2 contribution files such as pred_metagenome_contrib.tsv. It also accepts the long contribution schema used by path_abun_contrib.tsv; for clarity, use read_pathway_contrib_file when reading pathway-level contribution output.

Usage

read_contrib_file(
  file = NULL,
  data = NULL,
  type = c("auto", "gene_family", "pathway")
)

Arguments

file

Path to the contribution file (.tsv, .txt, .csv, or gzipped variants).

data

A data.frame already loaded from the contribution file. If both file and data are provided, data is used.

type

Character. Contribution file level. One of "auto", "gene_family", or "pathway". Default "auto".

Value

A data.frame with columns: sample, function_id, taxon, contribution abundance columns from the original file, and feature_level.

Details

The contribution file records how much each ASV/OTU contributes to the predicted abundance of each gene family or pathway in each sample. PICRUSt2 versions differ in whether they include norm_taxon_function_contrib; when that column is absent downstream aggregation can use taxon_function_abun or taxon_rel_function_abun. Contribution tables must describe one feature level at a time. Mixed gene-family identifiers such as KOs/ECs and pathway identifiers are rejected, because direct pathway matching and KEGG pathway-to-KO expansion have different biological meanings. Identifier columns (sample, function_id/function, and taxon) must contain non-empty values without NA; otherwise downstream aggregation would silently drop or mislabel contributions.

Examples

# \donttest{
# From a data.frame
contrib_df <- data.frame(
  sample = rep(c("S1", "S2"), each = 4),
  `function` = rep(c("K00001", "K00002"), 4),
  taxon = rep(c("ASV1", "ASV2"), each = 2, times = 2),
  taxon_function_abun = seq_len(8),
  check.names = FALSE
)
result <- read_contrib_file(data = contrib_df)
head(result)
# }