BioAnalysis.ca uses a distributed architecture with dedicated compute resources for each analysis job. No shared resources means zero cross-contamination risk and consistent performance.
🌐
Master Server
Web app, API, PostgreSQL, MinIO object storage, job orchestration
🖥️
Analysis VPS
Ubuntu 22.04, all bioinformatics tools, databases, GPU-ready
📊
Result Delivery
Interactive web report, downloadable PDF, raw data ZIP
Security: Files are transferred via private network (10.0.0.x). Analysis servers have no public internet access. All data is encrypted at rest (MinIO server-side encryption).
Runs during file upload, before any compute resources are used:
Compression check — .gz files must start with the gzip signature 0x1f 0x8b; uncompressed .fastq files are also accepted and compressed on the analysis server
FASTQ format check — decompresses first 4KB, verifies @header / sequence / + / quality 4-line format
FASTA detection — rejects FASTA files (headers starting with >) with clear error
NEW in v2.0 Instead of hardcoded truncation values, the system adapts to actual read quality:
Algorithm
Scan first 10,000 reads from R1 and R2 files
Compute median Phred quality score at each base position
Find the first position where median quality drops below Q25
Set truncation at that position (minimum floor: 150bp)
Safety caps: use 10th percentile read length (≥90% reads survive), never exceed region presets
V3-V4 Preset
R1: 280bp / R2: 200bp
V4 Preset
R1: 240bp / R2: 160bp
V1-V2 Preset
R1: 260bp / R2: 200bp
ITS (variable)
No truncation (trunc=0)
Rationale: Adaptive truncation can only shorten (for degraded quality), never exceed presets. This prevents data loss from over-aggressive truncation while protecting against quality-related errors.
3 Primer Trimming
Cutadapt v4.9
Removes primer sequences from the 5' end of reads. Reads without detectable primers are discarded.
assignTaxonomy() — classify ASVs using naive Bayesian classifier
addSpecies() — species-level assignment via exact matching
Parameter
Value
maxEE (max expected errors)
Fwd: 2, Rev: 2
truncQ
2 (truncate at first Q≤2 base)
Chimera method
consensus
Min overlap (paired-end)
12 bp
Output: ASV count table, representative sequences (FASTA), taxonomy assignments, denoising stats (reads at each stage)
5 Taxonomic Classification
SILVA v138.2UNITE v9.0
Gene
Classification DB
Species DB
Method
16S (bacterial)
SILVA NR99 v138.2
SILVA species v138.2
Naive Bayes + exact match
ITS (fungal)
UNITE v9.0
N/A
Naive Bayes
Taxonomy is assigned at 7 ranks: Kingdom → Phylum → Class → Order → Family → Genus → Species. Species assignment requires 100% identity to a reference sequence.
6 Phylogenetic Tree Construction
MAFFT v7.520FastTree v2.1.11
16S only — skipped for ITS (high indel rates make alignments unreliable).
MAFFT — multiple sequence alignment of all ASV representative sequences (FFT-NS-2 algorithm)
FastTree — approximate maximum-likelihood tree (GTR+CAT model)
Output: Newick tree file, used for Faith's Phylogenetic Diversity and UniFrac distances.
Note: FastTree uses heuristic optimization, so minor branch length variations (±1-2%) are expected between runs. This is normal and does not affect biological interpretation.
7 Alpha Diversity Analysis
scikit-bio v0.6.2
Metric
What It Measures
Interpretation
Shannon Index (H')
Richness + evenness
>3.0 = healthy gut; <2.0 = low diversity
Simpson Index (1-D)
Dominance probability
→ 1.0 = high diversity; → 0.0 = dominated by one taxon
Chao1
Estimated total richness
Predicts unseen species from singletons/doubletons
Observed ASVs
Raw count of unique sequences
Minimum richness estimate
Faith's PD
Phylogenetic branch length sum
Measures evolutionary diversity (requires tree)
Good's Coverage
Sampling completeness
>0.99 = adequately sequenced
8 Functional Prediction
PICRUSt2 v2.5.3(16S only)
Predicts metagenomic functional content from 16S marker gene data:
Phylogenetic placement — places ASVs into a reference tree (EPA-ng)
Gene content prediction — predicts gene families using ancestral state reconstruction
Disclaimer: PICRUSt2 predictions are computational estimates based on reference genome annotations. They should be validated with shotgun metagenomics for definitive functional profiling. Not applicable to ITS data.
9 Report Generation
matplotlib v3.9WeasyPrint v63.1Gemini AI
Charts Generated
Phylum composition — stacked bar chart of relative abundances
Genus composition — top 20 genera bar chart
Rarefaction curve — species accumulation curve (sampling depth vs. observed ASVs)
AI Interpretation
Each report section receives a contextual AI interpretation via Google Gemini (gemini-2.0-flash), providing clinical/ecological context for non-specialist users.
Deliverables
Interactive Web Report — full results with charts, quality assessment, parameter transparency
PDF Report — publication-ready document with all figures and tables
Research Data ZIP — ASV table, taxonomy table, representative sequences, tree file, raw stats
⚙ Complete Software Stack
Tool
Version
Purpose
Cutadapt
4.9
Primer trimming, quality filtering
DADA2
1.30.0
Denoising, ASV inference, chimera removal, taxonomy