Chapter 3

16S/ITS Amplicon Pipeline

Microbiome community profiling from amplicon sequencing data — ASV inference, taxonomy classification, diversity metrics, phylogenetic analysis, and functional prediction.

Pipeline Overview

The amplicon pipeline is designed for microbiome community samples — mixed microbial populations where the goal is to identify which organisms are present and in what proportions. It supports both 16S rRNA (bacterial communities) and ITS (fungal communities) marker gene sequencing.

📤 FASTQ (R1+R2)
→
✂️ Cutadapt
→
🧬 DADA2
→
📖 SILVA / UNITE
→
🌳 MAFFT + FastTree
→
📈 Alpha Diversity
→
⚗️ PICRUSt2
→
🤖 AI
→
📊 Report

16S vs. ITS — User-Selected Target Gene

Select the target gene (16S or ITS) during upload. The pipeline automatically routes to the correct reference database: SILVA v138.2 for 16S rRNA (bacteria/archaea) or UNITE v9.0 for ITS (fungi). Primer presets for common amplicon kits (e.g., V3-V4 region: 341F/785R) are available, or custom primer sequences can be entered manually.

Step 1: Primer Trimming (Cutadapt v5.2)

Cutadapt removes PCR primer sequences from both the forward and reverse reads. This is a critical preprocessing step because primer sequences are artificial and should not be included in the biological analysis.

Output MetricDescription
Reads in / Reads outTotal read pairs before and after primer trimming
Primer detection ratePercentage of reads where primers were found. Should be >90% for well-prepared libraries
Reads too shortReads discarded after trimming because they fell below the minimum length threshold

If a sample shows <70% primer detection rate, this may indicate incorrect primer sequences were used during library preparation, or the sample contains significant non-target DNA.

Step 2: ASV Inference (DADA2 v1.34.0)

DADA2 is the core of the amplicon pipeline. It performs error-corrected denoising to resolve Amplicon Sequence Variants (ASVs) — exact biological sequences that replace the older OTU (Operational Taxonomic Unit) approach. ASVs provide single-nucleotide resolution and are reproducible across studies.

DADA2 Processing Steps

Sub-stepDescription
Quality filteringTruncates reads at positions where quality drops below a threshold. The pipeline uses quality-adaptive truncation — it analyzes the quality profile and automatically selects optimal truncation lengths.
Error learningDADA2 learns the error model from the data itself, accounting for the specific error profile of each sequencing run.
DenoisingApplies the error model to distinguish true biological variants from sequencing errors. This is what makes DADA2 superior to OTU clustering.
Paired-end mergingMerges forward and reverse reads based on overlap. Requires at least 12 bp overlap (default).
Chimera removalIdentifies and removes chimeric sequences — artificial reads created when two different template molecules join during PCR amplification.

Report Outputs

Input reads
Total reads entering DADA2
Filtered
After quality truncation
Merged
After paired-end merging
Non-chimeric
Final reads used for analysis
ASVs
Unique biological sequences

Step 3: Taxonomy Classification (SILVA v138.2 / UNITE v9.0)

Each ASV is classified against a curated reference database using DADA2's naive Bayesian classifier. Classification is performed at all taxonomic ranks: Kingdom → Phylum → Class → Order → Family → Genus → Species.

DatabaseTarget GeneCoverageVersion
SILVA16S rRNABacteria + Archaea — the most comprehensive 16S reference with >500,000 curated sequencesv138.2 (2024)
UNITEITSFungi — the gold standard for fungal taxonomy based on ITS1/ITS2 sequencesv9.0 (July 2023)

Understanding "Unclassified" Taxa

The classifier uses an 80% confidence threshold by default. When an ASV sequence doesn't match any reference at a given rank with sufficient confidence, it is reported as unclassified at that level. This is common at genus and species level — especially for environmental samples and Proteobacteria-dominant communities. It does not mean the analysis failed; it means the database lacks a confident reference for that particular sequence.

Report Outputs

Step 4: Phylogenetic Analysis (MAFFT v7.525 + FastTree v2.2.0)

MAFFT performs multiple sequence alignment of all ASV representative sequences, then FastTree constructs a phylogenetic tree using approximate maximum-likelihood methods. This tree is used for:

Step 5: Alpha Diversity Metrics

Alpha diversity quantifies the diversity within a single sample. BioAnalysis.ca calculates six complementary metrics:

MetricWhat It MeasuresInterpretation
Shannon Index (H')Combination of richness and evenness. Accounts for both the number of species and how evenly they are distributed.<2.0 = low diversity (dysbiosis); 2.0–3.0 = moderate; >3.0 = high (healthy gut)
Simpson Index (1-D)Probability that two randomly selected individuals belong to different species. Less sensitive to rare species.0 = no diversity; 1 = maximum diversity. >0.9 is typical for healthy samples
Chao1 EstimatorEstimates the true total number of species, accounting for unobserved rare species based on singleton/doubleton ratios.Higher than Observed ASVs suggests many rare species remain undetected
Observed ASVsSimple count of unique ASV sequences detected. The most intuitive richness metric.Typical range: 50–500 for human microbiome; depends on environment type
Faith's PDSum of branch lengths on the phylogenetic tree connecting all observed ASVs. Phylogenetically-weighted richness.Higher values indicate more phylogenetically diverse communities
Good's CoverageEstimates how completely the sequencing captured the true diversity. Based on the proportion of singleton ASVs.>0.99 = excellent coverage; >0.97 = good; <0.95 = insufficient sequencing depth

Rarefaction Curves

The report includes a rarefaction curve showing how the number of observed ASVs increases with sequencing depth. A curve that reaches a plateau indicates sufficient sequencing depth — all detectable diversity has been captured. A still-rising curve suggests deeper sequencing might reveal additional taxa.

Step 6: Functional Prediction (PICRUSt2)

PICRUSt2 (Phylogenetic Investigation of Communities by Reconstruction of Unobserved States) predicts the functional potential of the microbial community based on 16S marker gene data. It maps ASVs to their closest sequenced genome relatives and infers KEGG metabolic pathway abundances.

Important Limitation

PICRUSt2 provides predicted functional potential, not measured gene expression or metabolite levels. Results should be interpreted as "this community likely has the genetic capacity for these functions" rather than "this community is actively performing these functions." For confirmed functional analysis, metagenomics (shotgun sequencing) or metatranscriptomics is recommended.

Top predicted pathways are displayed in a ranked bar chart in the report, with pathway abundances given as relative proportions of total predicted functional content.

Supported Amplicon Regions

RegionCommon PrimersTarget OrganismsNotes
V3-V4 (16S)341F / 785RBacteria, ArchaeaMost commonly used region for human microbiome studies. Good balance of resolution and database coverage.
V4 (16S)515F / 806RBacteria, ArchaeaEarth Microbiome Project standard. Shorter amplicon, compatible with 2×150 bp runs.
V1-V2 (16S)27F / 338RBacteriaGood resolution for some clinical genera (Staphylococcus, Streptococcus).
ITS1 / ITS2ITS1F / ITS2RFungiVariable-length region. DADA2 handles length variation through ITS-specific processing.

Custom primer sequences can be entered during upload. The pipeline will automatically use the provided sequences for Cutadapt trimming regardless of region.