Report Formats
Every completed analysis produces results in two complementary formats:
🌐 Interactive Web Report
A rich, browser-based report with hoverable charts, animated visualizations, sticky navigation, and dynamic data exploration. Optimized for on-screen review and presentation. Accessible from any device with a web browser — no software installation needed. Includes tooltips on every data point for detailed inspection.
📄 PDF Report
A print-optimized PDF with a professional cover page, all charts rendered as static images, and proper page breaks. Suitable for archiving, sharing with collaborators who lack platform access, and inclusion in lab notebooks or grant applications. Generated using browser-native print CSS.
WGS Report Sections
| Section | Contents |
|---|---|
| Quality Control | Read statistics (total reads, passed reads, fail rate), quality score distributions, adapter content, GC content, and duplication rate. Includes fastp pass/fail assessment. |
| Species Identification | Kraken2 classification results with top species, genus breakdown, and classification rate. Visual composition chart showing relative abundance of detected taxa. |
| Genome Assembly | QUAST metrics: total contigs, N50, largest contig, total length, GC content. Assembly quality assessment with interpretive benchmarks. |
| Gene Annotation | Bakta summary: total CDS, tRNA, rRNA, ncRNA, CRISPR arrays, and hypothetical proteins. Functional category breakdown chart. |
| MLST Typing | Sequence type (ST), scheme used, and allelic profile for all 7 loci. Links to PubMLST database for epidemiological context. |
| AMR Detection | Table of detected AMR genes with gene name, family, mechanism, target antibiotic class, % identity, and % coverage. Grouped by resistance mechanism. |
| AI Interpretation | Section-by-section AI interpretation covering significance, clinical context, and suggested follow-up actions. |
| Pipeline Provenance | Exact tool versions, database versions, command parameters, and timestamps for full reproducibility. |
Amplicon Report Sections
| Section | Contents |
|---|---|
| DADA2 Denoising Pipeline | Visual pipeline showing reads at each stage: Input → Filtered → Merged → Non-chimeric → Final ASVs. Allows assessment of data loss at each step. |
| Taxonomic Composition | Animated donut chart (phylum level), horizontal bar chart (top genera), genus abundance table with rank and relative proportion bars, species-level identification where resolvable. |
| Alpha Diversity | Six diversity metrics (Shannon, Simpson, Chao1, Observed ASVs, Faith's PD, Good's Coverage) with interpretive color-coded assessment (healthy/moderate/dysbiotic). Interactive rarefaction curve with data-point tooltips. |
| Functional Prediction | PICRUSt2 top metabolic pathways with relative abundance bars. Includes disclaimer about predictive nature. |
| Quality Assessment | Overall quality score (0-100), per-check pass/warn/fail results for sequencing depth, chimera rate, classification rate, and diversity metrics. |
| AI Interpretation | Section-specific AI commentary on taxonomy, diversity, functional potential, and overall community profile. |
| Pipeline Parameters | All DADA2 parameters (truncation lengths, error rates), database used, primer sequences, and tool versions. |
AI-Powered Scientific Interpretation
BioAnalysis.ca uses advanced AI to generate professional scientific interpretations for every report section. The AI receives the complete analysis results — not just summary statistics — and produces context-aware commentary.
How AI Interpretation Works
Architecture
- After the bioinformatics pipeline completes, an asynchronous task sends the full result set to the AI interpretation service
- Each report section (QC, taxonomy, diversity, AMR, overall) receives a separate, focused prompt engineered for that specific data type
- The AI is instructed to write for a PhD-level microbiologist audience — scientifically precise, data-driven, and concise (2–4 sentences per section)
- Interpretations include clinical/research significance and flag only what is statistically or clinically notable
- AI responses are stored in the database and rendered in the report as interpretation cards
Hybrid Model Strategy
- Gemini 2.5 Flash — used for high-volume, metrics-heavy sections: QC, pipeline parameters, alpha diversity, assembly statistics
- Gemini 3.1 Pro — used for deep-reasoning sections: taxonomy interpretation and the overall executive summary, which require integration of multiple data points into a coherent clinical narrative
- This routing strategy optimizes for both speed and scientific quality without increasing cost on straightforward numerical sections
What the AI Covers
| Section | AI Interpretation Focus |
|---|---|
| Quality Control | Assessment of sequencing quality, read loss at each step, whether depth is sufficient for the analysis type |
| Taxonomy / Species | Significance of the dominant taxa, known clinical/environmental associations, comparison to typical community profiles |
| Diversity | What the diversity metrics indicate about community health/dysbiosis, comparison to published reference ranges |
| AMR (WGS only) | Clinical significance of detected resistance genes, mechanism of action, known associations with treatment failure |
| Overall Summary | Integrated interpretation combining all sections into a cohesive narrative with key findings highlighted |
Important Disclaimer
AI interpretation is provided as a research aid — a high-quality first draft that saves time. It is not a substitute for expert microbiological review. All AI-generated text is clearly labeled in the report. Researchers should verify all interpretations against their domain knowledge and published literature before including them in publications or clinical decisions.
Comparative Project Reports
When multiple samples are grouped into a project, the platform generates an interactive cross-sample comparative report. Projects support 2–100 samples and work for both WGS and 16S/ITS amplicon data types. AI interpretation is automatically generated for the project report on first view, covering section-level analysis and an overall executive summary.
WGS Comparative Report
| Section | Description |
|---|---|
| Executive Summary + AI | Sample count, unique species, total reads, average genome size, AMR burden. AI summary: cohort composition, key cross-sample finding. |
| Species & MLST + AI | Species identification per sample with confidence %, MLST sequence types. AI flags clonal relationships (shared ST) and multi-species cohorts. |
| Genome Assembly + AI | N50, contig count, genome size comparison across all samples. AI identifies fragmented assemblies (N50 <100 Kbp) or outliers. |
| AMR Heatmap + AI | Gene-presence heatmap across all samples, drug class distribution chart. AI highlights shared genes (clonal spread / HGT) and MDR samples. |
| Read Quality + AI | Mean quality score and total reads per sample, QC pass/fail badges. AI flags samples below Q30. |
| Pipeline Charts | Full-resolution charts from the individual analysis pipelines, grouped by sample. |
Amplicon Comparative Report
| Section | Description |
|---|---|
| Executive Summary + AI | Sample count, total ASVs, average Shannon diversity, shared genera. AI overall summary (Pro model): cohort diversity profile and key ecological finding. |
| Alpha Diversity + AI | Animated bar charts comparing Shannon, Simpson, Chao1, Observed ASVs, Faith's PD, and Good's Coverage. AI flags samples with Shannon <2 as potential dysbiosis. |
| Community Composition | Stacked phylum bars, top genera per sample (mini charts), richness vs. evenness scatter plot, Firmicutes:Bacteroidetes ratio. |
| Genus Abundance Heatmap | Top 20 genera × samples heatmap. Dark cells = high abundance. Instantly reveals which genera are shared vs. sample-specific. |
| Beta Diversity | Bray-Curtis dissimilarity matrix computed from genus-level abundances. 0 = identical, 1 = completely different communities. |
| Core Microbiome | Genera classified as Core (present in 100% of samples), Common (≥50%), or Rare (unique to one sample). |
| Taxonomy + AI | Dominant genus per sample, genus frequency across cohort, shared genera highlighted. AI identifies whether the cohort shares a community type or has divergent profiles. |
| Pipeline Funnel | Read retention table: Raw → Filtered → Merged → Non-chimeric, with per-sample retention % and funnel bars. |
How AI is Triggered for Project Reports
Project AI interpretation is generated automatically on the first time the report page is viewed. This design handles two common scenarios: (1) projects created with newly uploaded samples that complete sequencing, and (2) projects assembled from previously-completed jobs. The AI generation runs in the background and the page polls for results every 15 seconds until complete. Once generated, results are cached in the database and load instantly on subsequent visits.
Research Data Package
Advanced and Complete tier analyses include a downloadable Research Data Package (ZIP) containing all raw analysis outputs suitable for downstream analysis and publication supplementary materials:
- ASV table (CSV) — samples × ASVs count matrix, ready for import into R/Python
- Taxonomy table (CSV) — full Kingdom-to-Species classification for every ASV
- Representative sequences (FASTA) — the actual DNA sequence of each ASV
- Phylogenetic tree (Newick) — importable into iTOL, FigTree, or ggtree
- Composition charts (PNG) — publication-ready phylum and genus composition figures
- Pipeline parameters (JSON) — all tool settings for Methods section writing
- Quality assessment (JSON) — structured quality scores for automated QC workflows