Skip to contents

This function takes a file containing KO (KEGG Orthology) abundance data in picrust2 export format and converts it to KEGG pathway abundance data. The input file should be in .tsv, .txt, or .csv format.

Usage

ko2kegg_abundance(
  file = NULL,
  data = NULL,
  method = c("abundance", "sum"),
  filter_for_prokaryotes = TRUE,
  progress = interactive()
)

Arguments

file

A character string representing the file path of the input file containing KO abundance data in picrust2 export format. The input file should have unique KO identifiers in the first column and sample identifiers in the first row. The remaining cells should contain the abundance values for each KO-sample pair.

data

An optional data frame or numeric matrix containing KO abundance data with one row per unique KO identifier. Data frames may use the same format as the input file, with KO identifiers in the first column, or data frames and matrices may store KO identifiers in row names. Sample columns must be numeric, finite, non-missing, and non-negative. If provided, the function will use this data instead of reading from the file. By default, this parameter is set to NULL.

method

Method for calculating pathway abundance. One of:

  • "abundance": (Default) Upper-half mean aggregation matching the unstructured pathway abundance rule used by the PICRUSt2 pathway pipeline. This is a KO-to-KEGG pathway aggregation approximation, not a replacement for the full PICRUSt2 pathway pipeline with MinPath and structured MetaCyc pathway inference.

  • "sum": Simple summation of all KO abundances. This is the legacy method and may double-count KOs belonging to multiple pathways.

filter_for_prokaryotes

Logical. If TRUE (default), filters out KEGG pathways that are not relevant to prokaryotic (bacterial/archaeal) analysis. The function always removes non-pathway KEGG buckets before this filter is applied. The prokaryote filter removes pathways in categories such as:

  • Human diseases (cancer, neurodegenerative diseases, addiction, etc.)

  • Organismal systems (immune system, nervous system, endocrine system, etc.)

Bacterial infection pathways and antimicrobial resistance pathways are retained. Set to FALSE to include all KEGG pathways (for eukaryotic analysis or custom filtering).

progress

Logical. Whether to show a progress bar while aggregating pathways. Defaults to interactive() so non-interactive scripts and tests stay quiet.

Value

A data frame with KEGG pathway abundance values. Rows represent KEGG pathways, identified by their KEGG pathway IDs. Columns represent samples, identified by their sample IDs from the input file.

Details

The default "abundance" method follows the unstructured pathway abundance rule in PICRUSt2's pathway pipeline:

  1. For each pathway, collect abundances of all associated KOs present in the data

  2. Sort the abundances in ascending order

  3. Take the upper half of the sorted values using the PICRUSt2 indexing rule sorted[int(n / 2):] (equivalent to floor(n / 2) + 1 through n in R's one-based indexing)

  4. Calculate the mean as the pathway abundance

This approach has several advantages over simple summation:

  • Does not inflate abundances for pathways containing more KOs

  • More robust to missing or low-abundance KOs

  • Provides a more accurate representation of pathway activity

The "sum" method is provided for backward compatibility and simply sums all KO abundances for each pathway.

Input sample columns must contain finite non-missing non-negative values. Missing values are rejected rather than ignored because the upper-half mean aggregation depends on the number and ordering of KO abundances available for each sample.

Input KO identifiers must be unique after cleaning optional ko: prefixes. Duplicate KO rows are rejected because the pathway aggregation rule treats each KO as one feature/reaction entry; repeated rows would duplicate evidence and distort upper-half mean or sum aggregation.

This function does not run MinPath, does not perform PICRUSt2's structured MetaCyc pathway inference, and does not estimate pathway coverage. It is intended as a practical offline KO-to-KEGG pathway aggregation step for downstream comparison and visualization.

Pathway Filtering

Before abundance calculation, KEGG BRITE hierarchies and "Not Included in Pathway or Brite" pseudo-pathways are removed because they are not KEGG pathway maps and cannot be consistently annotated as pathways (for example, ko99980).

When filter_for_prokaryotes = TRUE, the function excludes KEGG pathways that are biologically irrelevant to prokaryotic organisms. KEGG reference pathways include pathways from all domains of life, and many human/animal-specific pathways would appear in bacterial analysis simply because some KOs are shared across organisms.

The following KEGG Level 2 categories are excluded:

  • Cancer pathways (overview and specific types)

  • Neurodegenerative diseases (Alzheimer's, Parkinson's, etc.)

  • Substance dependence (addiction pathways)

  • Cardiovascular diseases

  • Endocrine and metabolic diseases

  • Immune diseases

  • Organismal systems (immune, nervous, endocrine, digestive, etc.)

The following are RETAINED even with filtering:

  • Infectious disease: bacterial (Salmonella, E. coli, Tuberculosis, etc.)

  • Drug resistance: antimicrobial (antibiotic resistance)

  • All Metabolism pathways

  • Genetic/Environmental Information Processing

  • Cellular Processes

Examples

if (FALSE) { # \dontrun{
library(ggpicrust2)
library(readr)

# Example 1: Default - filtered for prokaryotic analysis
data(ko_abundance)
kegg_abundance <- ko2kegg_abundance(data = ko_abundance)

# Example 2: Include all pathways (for eukaryotic analysis)
kegg_abundance_all <- ko2kegg_abundance(data = ko_abundance, filter_for_prokaryotes = FALSE)

# Example 3: Using legacy sum method with filtering
kegg_abundance_sum <- ko2kegg_abundance(data = ko_abundance, method = "sum")

# Example 4: From file
input_file <- tempfile(fileext = ".tsv")
utils::write.table(ko_abundance, input_file, sep = "\t",
                   quote = FALSE, row.names = FALSE)
kegg_abundance <- ko2kegg_abundance(file = input_file)
unlink(input_file)
} # }