Chapter 2

WGS Pipeline

Complete walkthrough of the whole-genome sequencing analysis pipeline — from raw reads to assembled, annotated, and resistance-profiled bacterial genomes.

Pipeline Overview

The WGS pipeline is designed for pure bacterial isolates — single-organism samples obtained from culture plates, blood cultures, or other isolation techniques. It performs a sequential series of analyses, each building on the previous step's output.

📤 FASTQ Input
→
⚡ Fastp (QC)
→
🧫 ConFindr
→
🔬 Kraken2
→
🧩 SPAdes
→
📐 QUAST
→
📝 Bakta
→
🏷️ MLST
→
🛡️ AMRFinder+
→
🤖 AI
→
📊 Report

Step 1: Quality Control (Fastp v1.1.0)

Fastp performs all-in-one preprocessing of the raw sequencing reads:

OperationDescription
Quality filteringRemoves reads below quality threshold (default Q15). Trims low-quality bases from 3' ends using a sliding window approach.
Adapter trimmingAutomatically detects and removes Illumina sequencing adapters. No adapter sequence input required — fastp identifies them from overlap analysis of paired reads.
Length filteringDiscards reads shorter than minimum length after trimming (default: 50 bp).
GC content analysisReports GC distribution, which can reveal contamination (bimodal distributions indicate mixed organisms).
Duplication rateEstimates PCR duplicate level. High duplication (>50%) suggests low library complexity.

Report outputs: Total reads before/after filtering, quality score distributions, adapter content, per-base quality heatmaps, and pass/fail quality assessment.

Step 2: Contamination Detection (ConFindr)

ConFindr detects intra-species contamination by analyzing ribosomal MLST (rMLST) genes. Unlike Kraken2 (which identifies species), ConFindr identifies strain-level mixing — for example, two different strains of E. coli mixed together, which Kraken2 would report as a single species.

Why Contamination Detection Matters

A contaminated isolate will produce a chimeric genome assembly with inflated contig counts, incorrect AMR gene calls, and unreliable MLST typing. ConFindr catches this before the assembly step, saving the researcher from interpreting invalid results.

Step 3: Species Identification (Kraken2 v2.17.1)

Kraken2 performs ultrafast taxonomic classification by matching k-mers (short DNA subsequences) against a comprehensive reference database. BioAnalysis.ca uses the PlusPF-16 database (~16 GB), which includes bacteria, archaea, viruses, plasmids, human, fungi, and protozoa.

OutputDescription
Species classificationTop species with confidence percentage. A clean isolate should show >90% reads mapping to a single species.
Genus-level breakdownDistribution of reads across genera — useful for detecting contamination at the genus level.
Classification ratePercentage of reads that mapped to any reference. Low classification rates (<50%) may indicate novel organisms or poor data quality.

Step 4: Genome Assembly (SPAdes + QUAST)

SPAdes performs de novo genome assembly — constructing contiguous sequences (contigs) from the short reads without a reference genome. QUAST then evaluates the assembly quality.

MetricDescriptionGood Values
Total contigsNumber of assembled sequences<200 for most bacteria
N5050% of the assembly is in contigs this size or larger. Higher = better continuity>50,000 bp
Largest contigLength of the longest assembled sequence>200,000 bp
Total lengthSum of all contig lengths — approximates genome sizeMatches expected genome size ±10%
GC contentGuanine-Cytosine percentage — characteristic of each speciesSpecies-specific

Step 5: Gene Annotation (Bakta v1.12.0)

Bakta is a rapid, standardized annotation tool for bacterial genomes. It identifies all genomic features including:

Bakta uses a lightweight reference database (db-light, ~1.3 GB) that provides standardized, taxonomy-aware annotation using UniProt, Pfam, and NCBI databases.

Step 6: MLST Typing (mlst v2.33.1)

mlst performs multi-locus sequence typing by scanning the assembled genome against the PubMLST database. MLST assigns a sequence type (ST) number based on allelic profiles of 7 housekeeping genes.

Clinical Significance

MLST typing is essential for: (1) tracking pathogen lineages during outbreak investigations, (2) identifying high-risk clones (e.g., E. coli ST131, Klebsiella ST258), (3) linking isolates to known epidemiological clusters, and (4) comparing with international surveillance databases.

Step 7: AMR Detection (AMRFinderPlus v4.2.7)

NCBI AMRFinderPlus identifies antimicrobial resistance genes, stress response genes, and virulence factors using the curated NCBI AMR reference database. This is the gold standard tool used by public health agencies worldwide.

Detection CategoryExamples
AMR genesblaNDM-1, blaCTX-M-15, mecA, vanA, mcr-1 — genes conferring resistance to specific antibiotic classes
Stress responseBiocide resistance, heavy metal tolerance, acid/heat resistance genes
Virulence factorsToxin genes, adhesion factors, invasion-related genes, capsule biosynthesis
Point mutationsChromosomal mutations in drug targets (e.g., gyrA mutations causing fluoroquinolone resistance)

For each detected gene, the report shows: gene name, gene family, resistance mechanism, antibiotic class affected, sequence identity (%), and coverage (%). This enables researchers to assess whether resistance is acquired (plasmid-borne) or intrinsic (chromosomal).