← BioAnalysis.ca · Documentation home

Whole Genome Sequencing (WGS) Pipeline

BioAnalysis.ca — Bacterial Isolate Analysis Workflow
Pipeline v2.0 • Updated April 2026

0 System Architecture

Each WGS job receives a dedicated analysis environment. Files are transferred over a private network, analyzed in isolation, and results are encrypted at rest.

🌐
Master Server
Web app, API, PostgreSQL, MinIO, Celery job queue
🖥️
Analysis VPS
Ubuntu 22.04, 8 CPU, 16GB RAM, all tools pre-installed
📊
Result Delivery
Interactive report, PDF, research data ZIP

→ Pipeline Flow & Analysis Tiers

📤 Upload FASTQ
→
🔍 Validation
→
✂️ QC (Fastp)
→
🦠 Species ID
→
🧬 Assembly
→
📝 Annotation
→
💊 AMR Detection
→
🔗 MLST
→
📄 Report

Basic

  • Quality Control
  • Species ID
  • PDF Report

Standard

  • All Basic steps
  • Genome Assembly
  • Gene Annotation
  • MLST Typing

Advanced

  • All Standard steps
  • AMR Detection
  • Resistance Profiling

1 Input Validation

Same 3-layer validation as the 16S pipeline:

2 Quality Control

Fastp v1.1.0

All-in-one FASTQ preprocessing: adapter removal, quality trimming, read filtering, and QC reporting.

ParameterValuePurpose
Adapter detectionAuto (overlap analysis)Removes Illumina adapters without specifying sequences
Quality thresholdQ20 (sliding window)Trims 3' ends below Q20 in 4bp windows
Min read length50 bpDiscards very short fragments
Complexity filterEnabled (30%)Removes low-complexity reads (poly-N, repeats)
DeduplicationDisabledPreserves coverage depth for assembly
Threads4Parallel processing

Output: Trimmed FASTQ, HTML QC report, JSON stats (total reads, GC content, quality distribution, adapter rates)

3 Species Identification

Kraken2 v2.17.1 Bracken v2.9

Kraken2 — k-mer Classification

Classifies each read by matching k-mers (k=35) against a pre-built database. Uses a lowest common ancestor (LCA) algorithm to handle ambiguous matches.

Bracken — Bayesian Abundance Re-estimation

Refines Kraken2's output using Bayesian re-estimation of species abundance, redistributing reads assigned to higher taxonomic levels.

ParameterValue
DatabaseStandard (archaea, bacteria, viral, human)
Confidence threshold0.1
Bracken levelSpecies (S)
Min reads for Bracken10

Output: Primary species, confidence %, full taxonomy report, species distribution

4 Genome Assembly

SPAdes v4.2.0 QUAST v5.3.0

SPAdes — De Novo Assembly

Multi-k-mer de Bruijn graph assembler optimized for bacterial isolates. Uses the --isolate mode for single-genome assembly.

ParameterValuePurpose
Mode--isolateOptimized for pure isolate sequencing
k-mer sizesAuto (21,33,55,77,99,127)Multi-k-mer improves contiguity
Threads4Parallel graph processing
Memory limit12 GBPrevents OOM on limited VPS

QUAST — Assembly Quality Assessment

Evaluates assembly quality without a reference genome:

Output: Scaffolds FASTA, QUAST report (TSV + HTML), assembly stats

5 Gene Annotation

Bakta v1.12.0

Rapid, comprehensive annotation of bacterial genomes. Bakta uses a hierarchical annotation approach with multiple databases.

Feature TypeDetection Method
CDS (protein-coding)Prodigal + UniProt/RefSeq alignment
tRNAtRNAscan-SE
rRNAInfernal + Rfam
ncRNAInfernal + Rfam
CRISPR arraysPILER-CR
Signal peptidesDeepSig

Output: GFF3 annotation, protein FASTA (.faa), nucleotide FASTA (.ffn), GenBank format, annotation summary

6 AMR Detection

AMRFinder+ v4.2.7 (NCBI)

Identifies antimicrobial resistance genes, stress response genes, and virulence factors. Uses both nucleotide and protein-level searches.

ParameterValue
Search modeProtein + Nucleotide (both)
OrganismAuto-detected from species ID (enables intrinsic resistance)
DatabaseNCBI AMR Reference Database (auto-updated)
Coverage threshold≥ 80%
Identity threshold≥ 90%

Output Categories

Clinical note: AMRFinder+ results are for research use only. Phenotypic antimicrobial susceptibility testing (AST) remains the gold standard for clinical decision-making.

7 MLST Typing

mlst v2.33.1 (Torsten Seemann)

Multi-Locus Sequence Typing identifies the sequence type (ST) of the isolate by matching 7 housekeeping gene alleles against PubMLST schemas.

Output: Scheme name, sequence type (ST), 7 allele profiles

8 Report Generation

matplotlib v3.9 WeasyPrint v63.1 Gemini AI

Charts Generated

AI Interpretation

Google Gemini provides section-specific interpretations for Species, Assembly, AMR, MLST, and QC — translating technical metrics into plain-language clinical context.

Deliverables

⚙ Complete Software Stack

ToolVersionPurpose
Fastp1.1.0Read QC, adapter removal, quality trimming
Kraken22.17.1k-mer taxonomic classification
Bracken2.9Bayesian abundance re-estimation
SPAdes4.2.0De novo genome assembly
QUAST5.3.0Assembly quality assessment
Bakta1.12.0Genome annotation (CDS, rRNA, tRNA)
AMRFinder+4.2.7AMR gene and point mutation detection
mlst2.33.1Multi-locus sequence typing
matplotlib3.9Scientific chart generation
WeasyPrint63.1PDF report generation
Python3.12Pipeline orchestration

Reference Databases

DatabasePurpose
Kraken2 Standard DBTaxonomic classification (archaea, bacteria, viral, human)
Bakta Light DBGene annotation (UniProt, RefSeq, Pfam, COG)
NCBI AMR Reference DBResistance gene detection
PubMLST schemasMLST allele matching (130+ species)