Pipeline Overview
The amplicon pipeline is designed for microbiome community samples — mixed microbial populations where the goal is to identify which organisms are present and in what proportions. It supports both 16S rRNA (bacterial communities) and ITS (fungal communities) marker gene sequencing.
16S vs. ITS — User-Selected Target Gene
Select the target gene (16S or ITS) during upload. The pipeline automatically routes to the correct reference database: SILVA v138.2 for 16S rRNA (bacteria/archaea) or UNITE v9.0 for ITS (fungi). Primer presets for common amplicon kits (e.g., V3-V4 region: 341F/785R) are available, or custom primer sequences can be entered manually.
Step 1: Primer Trimming (Cutadapt v5.2)
Cutadapt removes PCR primer sequences from both the forward and reverse reads. This is a critical preprocessing step because primer sequences are artificial and should not be included in the biological analysis.
| Output Metric | Description |
|---|---|
| Reads in / Reads out | Total read pairs before and after primer trimming |
| Primer detection rate | Percentage of reads where primers were found. Should be >90% for well-prepared libraries |
| Reads too short | Reads discarded after trimming because they fell below the minimum length threshold |
If a sample shows <70% primer detection rate, this may indicate incorrect primer sequences were used during library preparation, or the sample contains significant non-target DNA.
Step 2: ASV Inference (DADA2 v1.34.0)
DADA2 is the core of the amplicon pipeline. It performs error-corrected denoising to resolve Amplicon Sequence Variants (ASVs) — exact biological sequences that replace the older OTU (Operational Taxonomic Unit) approach. ASVs provide single-nucleotide resolution and are reproducible across studies.
DADA2 Processing Steps
| Sub-step | Description |
|---|---|
| Quality filtering | Truncates reads at positions where quality drops below a threshold. The pipeline uses quality-adaptive truncation — it analyzes the quality profile and automatically selects optimal truncation lengths. |
| Error learning | DADA2 learns the error model from the data itself, accounting for the specific error profile of each sequencing run. |
| Denoising | Applies the error model to distinguish true biological variants from sequencing errors. This is what makes DADA2 superior to OTU clustering. |
| Paired-end merging | Merges forward and reverse reads based on overlap. Requires at least 12 bp overlap (default). |
| Chimera removal | Identifies and removes chimeric sequences — artificial reads created when two different template molecules join during PCR amplification. |
Report Outputs
Step 3: Taxonomy Classification (SILVA v138.2 / UNITE v9.0)
Each ASV is classified against a curated reference database using DADA2's naive Bayesian classifier. Classification is performed at all taxonomic ranks: Kingdom → Phylum → Class → Order → Family → Genus → Species.
| Database | Target Gene | Coverage | Version |
|---|---|---|---|
| SILVA | 16S rRNA | Bacteria + Archaea — the most comprehensive 16S reference with >500,000 curated sequences | v138.2 (2024) |
| UNITE | ITS | Fungi — the gold standard for fungal taxonomy based on ITS1/ITS2 sequences | v9.0 (July 2023) |
Understanding "Unclassified" Taxa
The classifier uses an 80% confidence threshold by default. When an ASV sequence doesn't match any reference at a given rank with sufficient confidence, it is reported as unclassified at that level. This is common at genus and species level — especially for environmental samples and Proteobacteria-dominant communities. It does not mean the analysis failed; it means the database lacks a confident reference for that particular sequence.
Report Outputs
- Phylum composition — interactive donut chart + stacked bar chart showing relative abundance of all detected phyla
- Top genera table — ranked list of the most abundant genera with relative abundance percentages and visual proportion bars
- Species-level identification — where resolution allows, species assignments with the percentage of ASVs that could be classified to species level
- Dominant genus banner — prominently displayed in the report hero section
Step 4: Phylogenetic Analysis (MAFFT v7.525 + FastTree v2.2.0)
MAFFT performs multiple sequence alignment of all ASV representative sequences, then FastTree constructs a phylogenetic tree using approximate maximum-likelihood methods. This tree is used for:
- Faith's Phylogenetic Diversity (PD) — a diversity metric that accounts for evolutionary relationships, not just species counts
- UniFrac distances — for future beta diversity comparisons that weight phylogenetic distance
- Visual exploration — the phylogenetic tree is available as a downloadable Newick file for visualization in iTOL, FigTree, or other tree viewers
Step 5: Alpha Diversity Metrics
Alpha diversity quantifies the diversity within a single sample. BioAnalysis.ca calculates six complementary metrics:
| Metric | What It Measures | Interpretation |
|---|---|---|
| Shannon Index (H') | Combination of richness and evenness. Accounts for both the number of species and how evenly they are distributed. | <2.0 = low diversity (dysbiosis); 2.0–3.0 = moderate; >3.0 = high (healthy gut) |
| Simpson Index (1-D) | Probability that two randomly selected individuals belong to different species. Less sensitive to rare species. | 0 = no diversity; 1 = maximum diversity. >0.9 is typical for healthy samples |
| Chao1 Estimator | Estimates the true total number of species, accounting for unobserved rare species based on singleton/doubleton ratios. | Higher than Observed ASVs suggests many rare species remain undetected |
| Observed ASVs | Simple count of unique ASV sequences detected. The most intuitive richness metric. | Typical range: 50–500 for human microbiome; depends on environment type |
| Faith's PD | Sum of branch lengths on the phylogenetic tree connecting all observed ASVs. Phylogenetically-weighted richness. | Higher values indicate more phylogenetically diverse communities |
| Good's Coverage | Estimates how completely the sequencing captured the true diversity. Based on the proportion of singleton ASVs. | >0.99 = excellent coverage; >0.97 = good; <0.95 = insufficient sequencing depth |
Rarefaction Curves
The report includes a rarefaction curve showing how the number of observed ASVs increases with sequencing depth. A curve that reaches a plateau indicates sufficient sequencing depth — all detectable diversity has been captured. A still-rising curve suggests deeper sequencing might reveal additional taxa.
Step 6: Functional Prediction (PICRUSt2)
PICRUSt2 (Phylogenetic Investigation of Communities by Reconstruction of Unobserved States) predicts the functional potential of the microbial community based on 16S marker gene data. It maps ASVs to their closest sequenced genome relatives and infers KEGG metabolic pathway abundances.
Important Limitation
PICRUSt2 provides predicted functional potential, not measured gene expression or metabolite levels. Results should be interpreted as "this community likely has the genetic capacity for these functions" rather than "this community is actively performing these functions." For confirmed functional analysis, metagenomics (shotgun sequencing) or metatranscriptomics is recommended.
Top predicted pathways are displayed in a ranked bar chart in the report, with pathway abundances given as relative proportions of total predicted functional content.
Supported Amplicon Regions
| Region | Common Primers | Target Organisms | Notes |
|---|---|---|---|
| V3-V4 (16S) | 341F / 785R | Bacteria, Archaea | Most commonly used region for human microbiome studies. Good balance of resolution and database coverage. |
| V4 (16S) | 515F / 806R | Bacteria, Archaea | Earth Microbiome Project standard. Shorter amplicon, compatible with 2×150 bp runs. |
| V1-V2 (16S) | 27F / 338R | Bacteria | Good resolution for some clinical genera (Staphylococcus, Streptococcus). |
| ITS1 / ITS2 | ITS1F / ITS2R | Fungi | Variable-length region. DADA2 handles length variation through ITS-specific processing. |
Custom primer sequences can be entered during upload. The pipeline will automatically use the provided sequences for Cutadapt trimming regardless of region.