System Architecture
BioAnalysis.ca uses a microservices architecture with dedicated services for the web interface, API backend, job orchestration, and bioinformatics computation. All services are containerized with Docker and orchestrated via Docker Compose on the master server.
Master Server (Application Host)
Runs all application services in Docker containers, managed by Dokploy (a self-hosted deployment platform). This server handles user authentication, file uploads, job queuing, result storage, and web serving.
| Service | Technology | Role |
|---|---|---|
| Frontend | Next.js 15 (App Router), TypeScript | Server-rendered web application with interactive client-side components |
| Backend API | FastAPI, Python 3.12 | RESTful API for authentication, job management, file handling, and result serving |
| Worker | Celery (Python) | Asynchronous job orchestrator — dispatches analysis to cloud servers and monitors progress |
| Database | PostgreSQL 16 | Relational database for users, jobs, results, projects |
| Queue / Cache | Redis 7 | Job queue broker (Celery) and result cache |
| Object Storage | MinIO (S3-compatible) | Stores uploaded FASTQ files and pipeline output (reports, charts, data packages) |
| Reverse Proxy | Nginx | Routes traffic between frontend and backend, handles TLS termination |
Analysis Server (Compute)
Bioinformatics computation runs on cloud-based analysis servers configured with all required bioinformatics tools and reference databases. Data is logically isolated per user and per job at the storage layer.
| Specification | Configuration |
|---|---|
| Operating System | Ubuntu 22.04 LTS |
| Environment | Conda (Miniconda) with isolated bioanalysis environment |
| Cloud Provider | Major cloud infrastructure providers (Vultr, Hetzner, DigitalOcean) |
Pipeline Tool Versions (as of April 2026)
All tools are pinned to specific versions on the analysis server snapshot to ensure reproducibility across all analyses.
| Tool | Version | Purpose | Citation |
|---|---|---|---|
| fastp | 1.1.0 | Read QC and adapter trimming | Chen et al., 2018, Bioinformatics |
| Cutadapt | 5.2 | Primer sequence removal | Martin, 2011, EMBnet.journal |
| Kraken2 | 2.17.1 | Taxonomic classification (WGS) | Wood et al., 2019, Genome Biology |
| SPAdes | 3.15+ | De novo genome assembly | Bankevich et al., 2012, J Comput Biol |
| QUAST | latest | Assembly quality assessment | Gurevich et al., 2013, Bioinformatics |
| Bakta | 1.12.0 | Genome annotation | Schwengers et al., 2021, Microbial Genomics |
| AMRFinderPlus | 4.2.7 | AMR gene detection | Feldgarden et al., 2021, Nature Microbiology |
| mlst | 2.33.1 | Multi-locus sequence typing | Seemann T, github.com/tseemann/mlst |
| DADA2 | 1.34.0 | Amplicon denoising | Callahan et al., 2016, Nature Methods |
| R | 4.4.2 | DADA2 runtime environment | R Core Team, 2024 |
| MAFFT | v7.525 | Multiple sequence alignment | Katoh & Standley, 2013, Mol Biol Evol |
| FastTree | 2.2.0 | Phylogenetic tree construction | Price et al., 2010, PLoS One |
| PICRUSt2 | 2.5.3 | Functional prediction | Douglas et al., 2020, Nature Biotechnology |
| AI Interpretation | Latest | Scientific interpretation | Proprietary AI service |
Reference Databases
| Database | Version | Size | Used By |
|---|---|---|---|
| Kraken2 PlusPF-16 | 2024-06-05 | ~16 GB | Species classification (WGS). Includes bacteria, archaea, viruses, plasmids, human, fungi, protozoa. |
| SILVA | v138.2 | ~130 MB | 16S rRNA taxonomy classification. >500,000 curated sequences covering bacteria and archaea. |
| UNITE | v9.0 | ~80 MB | ITS taxonomy classification for fungal community profiling. |
| Bakta db-light | Schema v6 | ~1.3 GB | Genome annotation using UniProt, Pfam, and NCBI databases. |
| NCBI AMRFinder | v4 format | ~500 MB | Curated AMR gene, virulence factor, and stress response gene database. |
| PubMLST | Continuously updated | ~200 MB | Allelic profiles for MLST typing across >100 bacterial species schemes. |
Data Security & Privacy
🔐 Authentication
JWT-based authentication with bcrypt-hashed passwords. Access tokens expire after 60 minutes; refresh tokens after 30 days. All API endpoints require authentication except the public landing page.
🔒 Encryption in Transit
All traffic is encrypted via TLS/HTTPS. File uploads, API calls, and report viewing are all encrypted end-to-end. Certificate management is handled by Dokploy with automatic renewal.
🔒 Data Isolation
Each user can only access their own jobs and results. All uploaded files and results are stored in isolated, per-user storage paths with no cross-user access.
🔒 Secure & Isolated Compute
Each analysis job runs in an isolated environment. Data is separated at the storage level by user account, ensuring results from one user are never accessible to another.
Reproducibility & Provenance
Scientific reproducibility is a core design principle. Every analysis records:
- Exact tool versions — pinned in the analysis server snapshot, recorded in the result JSON
- Database versions — SILVA v138.2, UNITE v9.0, Kraken2 PlusPF-16, etc.
- Pipeline parameters — DADA2 truncation lengths, quality thresholds, primer sequences, error model settings
- Audit trail — full command log showing every tool invocation with exact arguments
- Resource profiling — CPU time, peak memory, disk usage, and wall-clock duration per step
- Quality assessment — automated quality scoring with pass/warn/fail verdicts for each checkpoint
This information is available in the report's Pipeline Parameters section and in the downloadable Research Data Package (JSON format), providing everything needed for a complete Methods section in a publication.
Application Technology Stack
| Layer | Technology | Purpose |
|---|---|---|
| Frontend Framework | Next.js 15 (App Router) | Server-side rendering, routing, and API proxy |
| Frontend Language | TypeScript | Type-safe JavaScript for maintainability |
| Styling | Vanilla CSS | No framework — full control over scientific chart rendering and print layouts |
| Charts | Custom SVG (inline) | All charts are custom-built SVG components — no chart library dependency. Enables precise control over scientific color palettes, print rendering, and animation. |
| Backend Framework | FastAPI (Python 3.12) | High-performance async API with automatic OpenAPI documentation |
| ORM | SQLAlchemy 2.0 (async) | Type-safe database queries with async PostgreSQL driver (asyncpg) |
| Migrations | Alembic | Database schema versioning and migration management |
| Task Queue | Celery + Redis | Distributed task execution for pipeline orchestration |
| Containerization | Docker + Docker Compose | Reproducible multi-service deployment |
| Deployment | Dokploy (self-hosted) | Git-push auto-deploy with zero-downtime restarts |
| Bioinformatics Runtime | Conda (Miniconda) | Isolated bioinformatics tool environment on analysis servers |
Design Philosophy
Scientific Color Palette
All reports — 16S/ITS amplicon, WGS, and comparative project reports — use the same curated 10-color scientific palette (teal, amber, rose, violet, emerald, sky, orange, indigo, cyan, red). The palette is designed for: (1) high contrast on white backgrounds, (2) print-safety for PDF output, and (3) colorblind accessibility. Consistent colors mean the same taxon or sample is always represented by the same color across individual and comparative reports.
The platform follows a consistent visual language across all reports:
- Animated chart reveals — bars and donut segments animate on first view, creating an engaging report-reading experience
- Hover tooltips everywhere — every data point, chart segment, and table cell provides additional detail on hover
- Print-optimized CSS —
@media printrules ensure clean PDF output with proper page breaks, hidden navigation, and cover/back pages - Sticky section navigation — scroll-spy navigation bar that highlights the current report section
- Responsive layout — reports are readable on mobile, tablet, and desktop screens