Pipeline Overview
The WGS pipeline is designed for pure bacterial isolates — single-organism samples obtained from culture plates, blood cultures, or other isolation techniques. It performs a sequential series of analyses, each building on the previous step's output.
Step 1: Quality Control (Fastp v1.1.0)
Fastp performs all-in-one preprocessing of the raw sequencing reads:
| Operation | Description |
|---|---|
| Quality filtering | Removes reads below quality threshold (default Q15). Trims low-quality bases from 3' ends using a sliding window approach. |
| Adapter trimming | Automatically detects and removes Illumina sequencing adapters. No adapter sequence input required — fastp identifies them from overlap analysis of paired reads. |
| Length filtering | Discards reads shorter than minimum length after trimming (default: 50 bp). |
| GC content analysis | Reports GC distribution, which can reveal contamination (bimodal distributions indicate mixed organisms). |
| Duplication rate | Estimates PCR duplicate level. High duplication (>50%) suggests low library complexity. |
Report outputs: Total reads before/after filtering, quality score distributions, adapter content, per-base quality heatmaps, and pass/fail quality assessment.
Step 2: Contamination Detection (ConFindr)
ConFindr detects intra-species contamination by analyzing ribosomal MLST (rMLST) genes. Unlike Kraken2 (which identifies species), ConFindr identifies strain-level mixing — for example, two different strains of E. coli mixed together, which Kraken2 would report as a single species.
Why Contamination Detection Matters
A contaminated isolate will produce a chimeric genome assembly with inflated contig counts, incorrect AMR gene calls, and unreliable MLST typing. ConFindr catches this before the assembly step, saving the researcher from interpreting invalid results.
Step 3: Species Identification (Kraken2 v2.17.1)
Kraken2 performs ultrafast taxonomic classification by matching k-mers (short DNA subsequences) against a comprehensive reference database. BioAnalysis.ca uses the PlusPF-16 database (~16 GB), which includes bacteria, archaea, viruses, plasmids, human, fungi, and protozoa.
| Output | Description |
|---|---|
| Species classification | Top species with confidence percentage. A clean isolate should show >90% reads mapping to a single species. |
| Genus-level breakdown | Distribution of reads across genera — useful for detecting contamination at the genus level. |
| Classification rate | Percentage of reads that mapped to any reference. Low classification rates (<50%) may indicate novel organisms or poor data quality. |
Step 4: Genome Assembly (SPAdes + QUAST)
SPAdes performs de novo genome assembly — constructing contiguous sequences (contigs) from the short reads without a reference genome. QUAST then evaluates the assembly quality.
| Metric | Description | Good Values |
|---|---|---|
| Total contigs | Number of assembled sequences | <200 for most bacteria |
| N50 | 50% of the assembly is in contigs this size or larger. Higher = better continuity | >50,000 bp |
| Largest contig | Length of the longest assembled sequence | >200,000 bp |
| Total length | Sum of all contig lengths — approximates genome size | Matches expected genome size ±10% |
| GC content | Guanine-Cytosine percentage — characteristic of each species | Species-specific |
Step 5: Gene Annotation (Bakta v1.12.0)
Bakta is a rapid, standardized annotation tool for bacterial genomes. It identifies all genomic features including:
- Coding sequences (CDS) — protein-coding genes with functional descriptions
- tRNA genes — transfer RNA genes (typically 50-80 per genome)
- rRNA genes — ribosomal RNA operons (16S, 23S, 5S)
- ncRNA — non-coding regulatory RNAs
- CRISPR arrays — adaptive immune system elements
- Signal peptides — secretion signals indicating exported proteins
- Hypothetical proteins — predicted genes with no known function
Bakta uses a lightweight reference database (db-light, ~1.3 GB) that provides standardized, taxonomy-aware annotation using UniProt, Pfam, and NCBI databases.
Step 6: MLST Typing (mlst v2.33.1)
mlst performs multi-locus sequence typing by scanning the assembled genome against the PubMLST database. MLST assigns a sequence type (ST) number based on allelic profiles of 7 housekeeping genes.
Clinical Significance
MLST typing is essential for: (1) tracking pathogen lineages during outbreak investigations, (2) identifying high-risk clones (e.g., E. coli ST131, Klebsiella ST258), (3) linking isolates to known epidemiological clusters, and (4) comparing with international surveillance databases.
Step 7: AMR Detection (AMRFinderPlus v4.2.7)
NCBI AMRFinderPlus identifies antimicrobial resistance genes, stress response genes, and virulence factors using the curated NCBI AMR reference database. This is the gold standard tool used by public health agencies worldwide.
| Detection Category | Examples |
|---|---|
| AMR genes | blaNDM-1, blaCTX-M-15, mecA, vanA, mcr-1 — genes conferring resistance to specific antibiotic classes |
| Stress response | Biocide resistance, heavy metal tolerance, acid/heat resistance genes |
| Virulence factors | Toxin genes, adhesion factors, invasion-related genes, capsule biosynthesis |
| Point mutations | Chromosomal mutations in drug targets (e.g., gyrA mutations causing fluoroquinolone resistance) |
For each detected gene, the report shows: gene name, gene family, resistance mechanism, antibiotic class affected, sequence identity (%), and coverage (%). This enables researchers to assess whether resistance is acquired (plasmid-borne) or intrinsic (chromosomal).