← BioAnalysis.ca · Documentation home
Each WGS job receives a dedicated analysis environment. Files are transferred over a private network, analyzed in isolation, and results are encrypted at rest.
Same 3-layer validation as the 16S pipeline:
Fastp v1.1.0
All-in-one FASTQ preprocessing: adapter removal, quality trimming, read filtering, and QC reporting.
| Parameter | Value | Purpose |
|---|---|---|
| Adapter detection | Auto (overlap analysis) | Removes Illumina adapters without specifying sequences |
| Quality threshold | Q20 (sliding window) | Trims 3' ends below Q20 in 4bp windows |
| Min read length | 50 bp | Discards very short fragments |
| Complexity filter | Enabled (30%) | Removes low-complexity reads (poly-N, repeats) |
| Deduplication | Disabled | Preserves coverage depth for assembly |
| Threads | 4 | Parallel processing |
Output: Trimmed FASTQ, HTML QC report, JSON stats (total reads, GC content, quality distribution, adapter rates)
Kraken2 v2.17.1 Bracken v2.9
Classifies each read by matching k-mers (k=35) against a pre-built database. Uses a lowest common ancestor (LCA) algorithm to handle ambiguous matches.
Refines Kraken2's output using Bayesian re-estimation of species abundance, redistributing reads assigned to higher taxonomic levels.
| Parameter | Value |
|---|---|
| Database | Standard (archaea, bacteria, viral, human) |
| Confidence threshold | 0.1 |
| Bracken level | Species (S) |
| Min reads for Bracken | 10 |
Output: Primary species, confidence %, full taxonomy report, species distribution
SPAdes v4.2.0 QUAST v5.3.0
Multi-k-mer de Bruijn graph assembler optimized for bacterial isolates. Uses the --isolate mode for single-genome assembly.
| Parameter | Value | Purpose |
|---|---|---|
| Mode | --isolate | Optimized for pure isolate sequencing |
| k-mer sizes | Auto (21,33,55,77,99,127) | Multi-k-mer improves contiguity |
| Threads | 4 | Parallel graph processing |
| Memory limit | 12 GB | Prevents OOM on limited VPS |
Evaluates assembly quality without a reference genome:
Output: Scaffolds FASTA, QUAST report (TSV + HTML), assembly stats
Bakta v1.12.0
Rapid, comprehensive annotation of bacterial genomes. Bakta uses a hierarchical annotation approach with multiple databases.
| Feature Type | Detection Method |
|---|---|
| CDS (protein-coding) | Prodigal + UniProt/RefSeq alignment |
| tRNA | tRNAscan-SE |
| rRNA | Infernal + Rfam |
| ncRNA | Infernal + Rfam |
| CRISPR arrays | PILER-CR |
| Signal peptides | DeepSig |
Output: GFF3 annotation, protein FASTA (.faa), nucleotide FASTA (.ffn), GenBank format, annotation summary
AMRFinder+ v4.2.7 (NCBI)
Identifies antimicrobial resistance genes, stress response genes, and virulence factors. Uses both nucleotide and protein-level searches.
| Parameter | Value |
|---|---|
| Search mode | Protein + Nucleotide (both) |
| Organism | Auto-detected from species ID (enables intrinsic resistance) |
| Database | NCBI AMR Reference Database (auto-updated) |
| Coverage threshold | ≥ 80% |
| Identity threshold | ≥ 90% |
mlst v2.33.1 (Torsten Seemann)
Multi-Locus Sequence Typing identifies the sequence type (ST) of the isolate by matching 7 housekeeping gene alleles against PubMLST schemas.
~ prefix if closest match is inexactOutput: Scheme name, sequence type (ST), 7 allele profiles
matplotlib v3.9 WeasyPrint v63.1 Gemini AI
Google Gemini provides section-specific interpretations for Species, Assembly, AMR, MLST, and QC — translating technical metrics into plain-language clinical context.
| Tool | Version | Purpose |
|---|---|---|
| Fastp | 1.1.0 | Read QC, adapter removal, quality trimming |
| Kraken2 | 2.17.1 | k-mer taxonomic classification |
| Bracken | 2.9 | Bayesian abundance re-estimation |
| SPAdes | 4.2.0 | De novo genome assembly |
| QUAST | 5.3.0 | Assembly quality assessment |
| Bakta | 1.12.0 | Genome annotation (CDS, rRNA, tRNA) |
| AMRFinder+ | 4.2.7 | AMR gene and point mutation detection |
| mlst | 2.33.1 | Multi-locus sequence typing |
| matplotlib | 3.9 | Scientific chart generation |
| WeasyPrint | 63.1 | PDF report generation |
| Python | 3.12 | Pipeline orchestration |
| Database | Purpose |
|---|---|
| Kraken2 Standard DB | Taxonomic classification (archaea, bacteria, viral, human) |
| Bakta Light DB | Gene annotation (UniProt, RefSeq, Pfam, COG) |
| NCBI AMR Reference DB | Resistance gene detection |
| PubMLST schemas | MLST allele matching (130+ species) |