An OTU (operational taxonomic unit) is a cluster of sequencing reads that are at least a set percentage similar, traditionally 97%. An ASV (amplicon sequence variant) is an exact biological sequence inferred after sequencing errors are corrected, so two ASVs can differ by a single base. ASVs have become the standard for short-read 16S studies because they are more precise and, unlike OTUs, mean the same thing in every dataset (Callahan et al., 2017).
What is an OTU?
Early 16S pipelines could not tell sequencing errors from real variation, so they grouped similar reads together and treated each group as one taxon. The 97% similarity threshold became a convention because it roughly matched the species level for full-length 16S sequences.
OTUs can be built three ways:
- De novo: reads are clustered against each other. The clusters depend on which other reads are in the dataset, so the same organism can fall into different OTUs in different studies.
- Closed-reference: reads are matched to a reference database at 97%. Reads that match nothing are discarded, so novel organisms are lost.
- Open-reference: a mix of both.
What is an ASV?
Denoising methods such as DADA2 (Callahan et al., 2016), Deblur (Amir et al., 2017) and UNOISE model the errors produced by the sequencer. They then infer which sequences are truly present in the sample. The result is a set of exact sequences, each with a count per sample.
Because an ASV is a real DNA sequence rather than a cluster label, the same organism amplified with the same primers gives the same ASV in any study, on any run.
How do ASVs and OTUs compare?
| Property | OTUs (97%) | ASVs |
|---|---|---|
| Unit | Cluster of similar reads | Exact error-corrected sequence |
| Resolution | Differences under 3% are merged | Single-nucleotide differences kept |
| Comparable between studies | No (de novo) or only via a shared reference | Yes, directly |
| Novel taxa | Lost in closed-reference clustering | Kept |
| Sensitive to sequencing errors | Errors inflate OTU counts | Errors are modelled and removed |
| Typical tools | mothur, VSEARCH, older QIIME | DADA2, Deblur, UNOISE |
How does DADA2 infer ASVs?
DADA2 works in four stages:
- Learn error rates. It estimates how often each base is misread as each other base at each quality score, from the data itself.
- Partition reads. Starting from the most abundant sequence, it asks whether each less abundant sequence could plausibly be an error copy of a more abundant one. If not, it becomes a new ASV.
- Merge pairs. Forward and reverse reads are joined into full-length amplicon sequences, which requires them to overlap.
- Remove chimeras. Sequences that look like a hybrid of two more abundant parents, artifacts of PCR, are discarded.
Because the error model is learned per run, samples from the same run should be denoised together. In a multi-sample study, BioAnalysis.ca fits one shared error model across all samples for this reason.
Why did ASVs replace OTUs?
Callahan, McMurdie and Holmes (2017) set out the main arguments:
- Reproducibility: ASVs are consistent labels, so results can be compared and pooled across studies without re-clustering.
- Reusability: an ASV table from one study can be combined with another that used the same primers.
- Resolution: single-base differences can separate ecologically distinct strains that a 97% cluster would merge.
- Accuracy: denoising removes errors that would otherwise inflate richness.
Major tools followed. QIIME 2's denoising options are DADA2 and Deblur, and most published short-read 16S studies now report ASVs.
Do ASVs have drawbacks?
Some. Many bacteria carry several copies of the 16S gene, and those copies are not always identical, so one genome can produce more than one ASV (Schloss, 2021). ASVs are also stricter about rare sequences: a variant seen only once is usually indistinguishable from an error and is discarded. Neither is a reason to go back to OTUs, but both are worth remembering when you count ASVs as if they were species.
When do OTUs still make sense?
- Comparing new data with an older study that published only OTU tables.
- Some long-read or high-error workflows where exact denoising is less mature.
- Analyses that deliberately want coarser units. In that case, aggregating ASVs at genus level is usually clearer than re-clustering.
What does BioAnalysis.ca report?
BioAnalysis.ca denoises with DADA2 and reports ASVs. Every report includes the ASV table, the taxonomy of each ASV against SILVA v138.2 (or UNITE v9.0 for ITS), composition summarized at phylum and genus level, and read counts at each DADA2 stage. See the full 16S pipeline.
References
- Amir A, McDonald D, Navas-Molina JA, et al. Deblur rapidly resolves single-nucleotide community sequence patterns. mSystems. 2017;2:e00191-16.
- Callahan BJ, McMurdie PJ, Rosen MJ, et al. DADA2: High-resolution sample inference from Illumina amplicon data. Nature Methods. 2016;13:581–583.
- Callahan BJ, McMurdie PJ, Holmes SP. Exact sequence variants should replace operational taxonomic units in marker-gene data analysis. The ISME Journal. 2017;11:2639–2643.
- Schloss PD. Amplicon sequence variants artificially split bacterial genomes into separate clusters. mSphere. 2021;6:e00191-21.