← BioAnalysis.ca · Documentation home

16S/ITS Amplicon Analysis Pipeline

BioAnalysis.ca — Comprehensive Microbiome Profiling Workflow
Pipeline v2.0 • Updated April 2026

0 System Architecture

BioAnalysis.ca uses a distributed architecture with dedicated compute resources for each analysis job. No shared resources means zero cross-contamination risk and consistent performance.

🌐
Master Server
Web app, API, PostgreSQL, MinIO object storage, job orchestration
🖥️
Analysis VPS
Ubuntu 22.04, all bioinformatics tools, databases, GPU-ready
📊
Result Delivery
Interactive web report, downloadable PDF, raw data ZIP
Security: Files are transferred via private network (10.0.0.x). Analysis servers have no public internet access. All data is encrypted at rest (MinIO server-side encryption).

→ Pipeline Flow Overview

📤 Upload FASTQ
→
🔍 Input Validation
→
✂️ Primer Trimming
→
🧬 DADA2 Denoising
→
🔬 Taxonomy
→
🌳 Phylogenetics
→
📈 Diversity
→
⚡ Functional Prediction
→
📋 Quality Assessment
→
📄 Report
Supported regions: V3-V4 (341F/805R), V4 (515F/806R), V1-V2 (27F/338R), V1-V3 (27F/534R), ITS1 (ITS1F/ITS2), ITS2 (ITS3/ITS4)

1 Input Validation (3-Layer System)

Layer 1 — Upload-Time Validation (instant)

Runs during file upload, before any compute resources are used:

Layer 2 — Pre-Pipeline Validation (~30 seconds)

Runs on the analysis server after download, before starting the pipeline:

Layer 3 — Quality Assessment (post-analysis)

After analysis completes, every result is scored 0–100:

Check❌ Fail⚠️ Warn✅ PassWeight
Read Count< 500< 5,000≥ 10,0002×
Primer Trim Rate< 20%< 60%≥ 80%2×
ASV Count< 5< 20≥ 502×
Read Survival< 10%< 30%≥ 50%1×
Shannon Diversity< 0.5< 1.5≥ 2.51×
Good's Coverage< 90%< 95%≥ 98%1×
Chimera Rate> 30%> 10%< 10%1×

Verdict: Excellent (all pass) → Good (1 warning) → Fair (2+ warnings) → Poor (any fail)

2 Quality-Adaptive Truncation Detection

NEW in v2.0 Instead of hardcoded truncation values, the system adapts to actual read quality:

Algorithm

  1. Scan first 10,000 reads from R1 and R2 files
  2. Compute median Phred quality score at each base position
  3. Find the first position where median quality drops below Q25
  4. Set truncation at that position (minimum floor: 150bp)
  5. Safety caps: use 10th percentile read length (≥90% reads survive), never exceed region presets
V3-V4 Preset
R1: 280bp / R2: 200bp
V4 Preset
R1: 240bp / R2: 160bp
V1-V2 Preset
R1: 260bp / R2: 200bp
ITS (variable)
No truncation (trunc=0)
Rationale: Adaptive truncation can only shorten (for degraded quality), never exceed presets. This prevents data loss from over-aggressive truncation while protecting against quality-related errors.

3 Primer Trimming

Cutadapt v4.9

Removes primer sequences from the 5' end of reads. Reads without detectable primers are discarded.

ParameterValueRationale
Error rate10% (default)Allows 1-2 mismatches in primer matching
Min read length100 bpDiscard very short fragments post-trim
Quality trimQ20 (3' end)Trim low-quality tails before denoising
Threads4 (auto)Parallel processing for speed
Paired-endAuto-detectedLinked adapters for R1+R2

Output: Trimmed FASTQ files + statistics (reads in/out, trim rate %)

4 Denoising & ASV Generation

DADA2 v1.30.0 (R/Bioconductor)

DADA2 infers exact Amplicon Sequence Variants (ASVs) from the quality-filtered reads. Unlike OTU clustering (97% similarity), ASVs resolve single-nucleotide differences.

Sub-steps

  1. filterAndTrim() — truncate reads (quality-adaptive values), discard reads with >2 expected errors
  2. learnErrors() — learn error model from the data (parametric error model)
  3. dada() — denoise reads using the learned error model
  4. mergePairs() — merge R1+R2 (paired-end), requiring ≥12bp overlap
  5. makeSequenceTable() — create ASV × sample count matrix
  6. removeBimeraDenovo() — remove chimeric sequences (consensus method)
  7. assignTaxonomy() — classify ASVs using naive Bayesian classifier
  8. addSpecies() — species-level assignment via exact matching
ParameterValue
maxEE (max expected errors)Fwd: 2, Rev: 2
truncQ2 (truncate at first Q≤2 base)
Chimera methodconsensus
Min overlap (paired-end)12 bp

Output: ASV count table, representative sequences (FASTA), taxonomy assignments, denoising stats (reads at each stage)

5 Taxonomic Classification

SILVA v138.2 UNITE v9.0

GeneClassification DBSpecies DBMethod
16S (bacterial)SILVA NR99 v138.2SILVA species v138.2Naive Bayes + exact match
ITS (fungal)UNITE v9.0N/ANaive Bayes

Taxonomy is assigned at 7 ranks: Kingdom → Phylum → Class → Order → Family → Genus → Species. Species assignment requires 100% identity to a reference sequence.

6 Phylogenetic Tree Construction

MAFFT v7.520 FastTree v2.1.11

16S only — skipped for ITS (high indel rates make alignments unreliable).

  1. MAFFT — multiple sequence alignment of all ASV representative sequences (FFT-NS-2 algorithm)
  2. FastTree — approximate maximum-likelihood tree (GTR+CAT model)

Output: Newick tree file, used for Faith's Phylogenetic Diversity and UniFrac distances.

Note: FastTree uses heuristic optimization, so minor branch length variations (±1-2%) are expected between runs. This is normal and does not affect biological interpretation.

7 Alpha Diversity Analysis

scikit-bio v0.6.2

MetricWhat It MeasuresInterpretation
Shannon Index (H')Richness + evenness>3.0 = healthy gut; <2.0 = low diversity
Simpson Index (1-D)Dominance probability→ 1.0 = high diversity; → 0.0 = dominated by one taxon
Chao1Estimated total richnessPredicts unseen species from singletons/doubletons
Observed ASVsRaw count of unique sequencesMinimum richness estimate
Faith's PDPhylogenetic branch length sumMeasures evolutionary diversity (requires tree)
Good's CoverageSampling completeness>0.99 = adequately sequenced

8 Functional Prediction

PICRUSt2 v2.5.3 (16S only)

Predicts metagenomic functional content from 16S marker gene data:

  1. Phylogenetic placement — places ASVs into a reference tree (EPA-ng)
  2. Gene content prediction — predicts gene families using ancestral state reconstruction
  3. Pathway inference — maps predicted genes to MetaCyc metabolic pathways (MinPath)
Disclaimer: PICRUSt2 predictions are computational estimates based on reference genome annotations. They should be validated with shotgun metagenomics for definitive functional profiling. Not applicable to ITS data.

9 Report Generation

matplotlib v3.9 WeasyPrint v63.1 Gemini AI

Charts Generated

AI Interpretation

Each report section receives a contextual AI interpretation via Google Gemini (gemini-2.0-flash), providing clinical/ecological context for non-specialist users.

Deliverables

⚙ Complete Software Stack

ToolVersionPurpose
Cutadapt4.9Primer trimming, quality filtering
DADA21.30.0Denoising, ASV inference, chimera removal, taxonomy
MAFFT7.520Multiple sequence alignment
FastTree2.1.11Phylogenetic tree construction (GTR+CAT)
scikit-bio0.6.2Alpha diversity metrics
PICRUSt22.5.3Functional prediction from 16S
matplotlib3.9Composition and rarefaction charts
WeasyPrint63.1PDF report generation
R4.4.3DADA2 runtime environment
Python3.12Pipeline orchestration

Reference Databases

DatabaseVersionUse
SILVANR99 v138.216S taxonomy classification + species assignment
UNITEv9.0ITS (fungal) taxonomy classification
MetaCycv26.5Metabolic pathway reference (PICRUSt2)