Chapter 1

Platform Overview

How BioAnalysis.ca works, what it does, and how researchers interact with the system from sample upload to final report.

What Is BioAnalysis.ca?

BioAnalysis.ca is a fully automated, cloud-based bioinformatics platform designed for researchers who need to analyze bacterial genomic data without maintaining their own computational infrastructure. The platform accepts raw sequencing data (FASTQ files) from Illumina sequencers and performs end-to-end analysis using peer-reviewed, publication-standard bioinformatics tools.

The platform supports two distinct analysis modes:

🧬 Whole-Genome Sequencing (WGS)

For pure bacterial isolates. Performs quality control, species identification, de novo genome assembly, gene annotation, antimicrobial resistance (AMR) gene detection, and multi-locus sequence typing (MLST). Ideal for clinical isolate characterization, outbreak investigation, and genomic surveillance.

πŸ”¬ 16S/ITS Amplicon Analysis

For microbiome community profiling. Performs primer trimming, amplicon sequence variant (ASV) inference via DADA2, taxonomy classification against SILVA v138.2 (16S) or UNITE v9.0 (ITS), alpha diversity metrics, phylogenetic tree construction, and functional prediction via PICRUSt2.

The Three-Step Workflow

BioAnalysis.ca was designed around the principle that bioinformatics should not be the bottleneck in microbiology research. The entire processβ€”from raw sequencing data to interpreted resultsβ€”requires three actions from the user:

Step 1: Upload FASTQ Files

Drag and drop .fastq.gz or .fastq files directly in the browser. Select the analysis type (WGS or 16S/ITS), choose the analysis tier, and optionally assign the sample to a project for comparative analysis later. Paired-end and single-end reads are both supported. Files must be under 1 GB each. You can upload up to 200 files (100 samples) per batch submission.

Step 2: Automated Analysis

Once uploaded, the platform runs the full bioinformatics pipeline automatically. Each step is tracked in real time β€” the dashboard shows exactly which tool is running, what percentage is complete, and how long the analysis has been running. No user intervention is needed during this phase.

Step 3: Explore Results

When the pipeline completes, results are available in two formats: an interactive web report accessible from any browser (with tooltips, hoverable charts, and dynamic visualizations) and a downloadable PDF formatted for printing and archiving. An AI interpretation section provides professional scientific commentary on every result section.

Dashboard & Job Management

After logging in, researchers see a unified dashboard that displays all their analyses in a single view:

ColumnDescription
SampleOriginal filename and file size of the uploaded FASTQ data
TypeAnalysis type badge: 🧬 WGS, πŸ”¬ 16S, or πŸ„ ITS
StatusCurrent state: Queued, Running (with spinner), Completed (βœ“), or Failed (βœ•)
ProgressVisual progress bar with percentage and current pipeline step label (e.g., "Taxonomy Classification")
SpeciesPrimary organism identified β€” dominant species (WGS) or dominant genus (amplicon)
Submitted / CompletedTimestamps for job creation and completion
DurationTotal wall-clock time from start to finish

The dashboard auto-refreshes every 10 seconds while jobs are running, so researchers can monitor progress in real time without manually reloading. Clicking any row navigates directly to that job's detailed report page.

Project System & Comparative Analysis

Samples can be grouped into projects for cross-sample comparison. This is particularly powerful for amplicon studies where researchers are comparing microbiome communities across conditions, time points, or treatment groups.

Projects require a minimum of 2 samples and support up to 100 samples per project. When a project contains multiple completed amplicon analyses, the platform generates a comparative project report that includes:

Typical Processing Times

3-5m
16S Basic
5-10m
16S Complete
~30m
WGS Basic
1-2h
WGS Standard
2-3h
WGS Advanced

Processing times depend on input file size and sequencing depth. Times shown are for typical Illumina MiSeq paired-end datasets. All pipelines include AI interpretation, which adds approximately 30-60 seconds to the total time.

Analysis Tiers

WGS Tiers

TierSteps IncludedTypical Use Case
BasicQC (Fastp) β†’ Species ID (Kraken2) β†’ AI Interpretation β†’ ReportQuick species identification from an isolate
StandardBasic + ConFindr contamination check β†’ SPAdes assembly β†’ QUAST β†’ Bakta annotation β†’ MLST typingFull genome characterization
AdvancedStandard + AMRFinderPlus AMR/virulence detection β†’ Research Data Package (ZIP)Clinical/AMR surveillance, publication-grade output

Amplicon Tiers

TierSteps IncludedTypical Use Case
Amplicon BasicCutadapt primer trimming β†’ DADA2 denoising β†’ SILVA/UNITE taxonomy β†’ Composition charts β†’ AI Interpretation β†’ ReportQuick microbiome profiling β€” "what's in my sample?"
Amplicon CompleteBasic + MAFFT/FastTree phylogeny β†’ Alpha diversity (Shannon, Simpson, Chao1, Faith's PD, rarefaction) β†’ PICRUSt2 functional prediction β†’ Research Data PackageFull microbiome study with diversity metrics and functional analysis