Chapter 5

Technology & Infrastructure

Cloud architecture, security model, tool versions, reference databases, reproducibility, and the full technology stack powering BioAnalysis.ca.

System Architecture

BioAnalysis.ca uses a microservices architecture with dedicated services for the web interface, API backend, job orchestration, and bioinformatics computation. All services are containerized with Docker and orchestrated via Docker Compose on the master server.

Master Server (Application Host)

Runs all application services in Docker containers, managed by Dokploy (a self-hosted deployment platform). This server handles user authentication, file uploads, job queuing, result storage, and web serving.

ServiceTechnologyRole
FrontendNext.js 15 (App Router), TypeScriptServer-rendered web application with interactive client-side components
Backend APIFastAPI, Python 3.12RESTful API for authentication, job management, file handling, and result serving
WorkerCelery (Python)Asynchronous job orchestrator — dispatches analysis to cloud servers and monitors progress
DatabasePostgreSQL 16Relational database for users, jobs, results, projects
Queue / CacheRedis 7Job queue broker (Celery) and result cache
Object StorageMinIO (S3-compatible)Stores uploaded FASTQ files and pipeline output (reports, charts, data packages)
Reverse ProxyNginxRoutes traffic between frontend and backend, handles TLS termination

Analysis Server (Compute)

Bioinformatics computation runs on cloud-based analysis servers configured with all required bioinformatics tools and reference databases. Data is logically isolated per user and per job at the storage layer.

SpecificationConfiguration
Operating SystemUbuntu 22.04 LTS
EnvironmentConda (Miniconda) with isolated bioanalysis environment
Cloud ProviderMajor cloud infrastructure providers (Vultr, Hetzner, DigitalOcean)

Pipeline Tool Versions (as of April 2026)

All tools are pinned to specific versions on the analysis server snapshot to ensure reproducibility across all analyses.

ToolVersionPurposeCitation
fastp1.1.0Read QC and adapter trimmingChen et al., 2018, Bioinformatics
Cutadapt5.2Primer sequence removalMartin, 2011, EMBnet.journal
Kraken22.17.1Taxonomic classification (WGS)Wood et al., 2019, Genome Biology
SPAdes3.15+De novo genome assemblyBankevich et al., 2012, J Comput Biol
QUASTlatestAssembly quality assessmentGurevich et al., 2013, Bioinformatics
Bakta1.12.0Genome annotationSchwengers et al., 2021, Microbial Genomics
AMRFinderPlus4.2.7AMR gene detectionFeldgarden et al., 2021, Nature Microbiology
mlst2.33.1Multi-locus sequence typingSeemann T, github.com/tseemann/mlst
DADA21.34.0Amplicon denoisingCallahan et al., 2016, Nature Methods
R4.4.2DADA2 runtime environmentR Core Team, 2024
MAFFTv7.525Multiple sequence alignmentKatoh & Standley, 2013, Mol Biol Evol
FastTree2.2.0Phylogenetic tree constructionPrice et al., 2010, PLoS One
PICRUSt22.5.3Functional predictionDouglas et al., 2020, Nature Biotechnology
AI InterpretationLatestScientific interpretationProprietary AI service

Reference Databases

DatabaseVersionSizeUsed By
Kraken2 PlusPF-162024-06-05~16 GBSpecies classification (WGS). Includes bacteria, archaea, viruses, plasmids, human, fungi, protozoa.
SILVAv138.2~130 MB16S rRNA taxonomy classification. >500,000 curated sequences covering bacteria and archaea.
UNITEv9.0~80 MBITS taxonomy classification for fungal community profiling.
Bakta db-lightSchema v6~1.3 GBGenome annotation using UniProt, Pfam, and NCBI databases.
NCBI AMRFinderv4 format~500 MBCurated AMR gene, virulence factor, and stress response gene database.
PubMLSTContinuously updated~200 MBAllelic profiles for MLST typing across >100 bacterial species schemes.

Data Security & Privacy

🔐 Authentication

JWT-based authentication with bcrypt-hashed passwords. Access tokens expire after 60 minutes; refresh tokens after 30 days. All API endpoints require authentication except the public landing page.

🔒 Encryption in Transit

All traffic is encrypted via TLS/HTTPS. File uploads, API calls, and report viewing are all encrypted end-to-end. Certificate management is handled by Dokploy with automatic renewal.

🔒 Data Isolation

Each user can only access their own jobs and results. All uploaded files and results are stored in isolated, per-user storage paths with no cross-user access.

🔒 Secure & Isolated Compute

Each analysis job runs in an isolated environment. Data is separated at the storage level by user account, ensuring results from one user are never accessible to another.

Reproducibility & Provenance

Scientific reproducibility is a core design principle. Every analysis records:

This information is available in the report's Pipeline Parameters section and in the downloadable Research Data Package (JSON format), providing everything needed for a complete Methods section in a publication.

Application Technology Stack

LayerTechnologyPurpose
Frontend FrameworkNext.js 15 (App Router)Server-side rendering, routing, and API proxy
Frontend LanguageTypeScriptType-safe JavaScript for maintainability
StylingVanilla CSSNo framework — full control over scientific chart rendering and print layouts
ChartsCustom SVG (inline)All charts are custom-built SVG components — no chart library dependency. Enables precise control over scientific color palettes, print rendering, and animation.
Backend FrameworkFastAPI (Python 3.12)High-performance async API with automatic OpenAPI documentation
ORMSQLAlchemy 2.0 (async)Type-safe database queries with async PostgreSQL driver (asyncpg)
MigrationsAlembicDatabase schema versioning and migration management
Task QueueCelery + RedisDistributed task execution for pipeline orchestration
ContainerizationDocker + Docker ComposeReproducible multi-service deployment
DeploymentDokploy (self-hosted)Git-push auto-deploy with zero-downtime restarts
Bioinformatics RuntimeConda (Miniconda)Isolated bioinformatics tool environment on analysis servers

Design Philosophy

Scientific Color Palette

All reports — 16S/ITS amplicon, WGS, and comparative project reports — use the same curated 10-color scientific palette (teal, amber, rose, violet, emerald, sky, orange, indigo, cyan, red). The palette is designed for: (1) high contrast on white backgrounds, (2) print-safety for PDF output, and (3) colorblind accessibility. Consistent colors mean the same taxon or sample is always represented by the same color across individual and comparative reports.

The platform follows a consistent visual language across all reports: