Guide

ASVs vs OTUs: Why DADA2 Replaced OTU Clustering

The difference between amplicon sequence variants (ASVs) and operational taxonomic units (OTUs), why DADA2 ASVs are now the standard, and when OTUs still appear.

An OTU (operational taxonomic unit) is a cluster of sequencing reads that are at least a set percentage similar, traditionally 97%. An ASV (amplicon sequence variant) is an exact biological sequence inferred after sequencing errors are corrected, so two ASVs can differ by a single base. ASVs have become the standard for short-read 16S studies because they are more precise and, unlike OTUs, mean the same thing in every dataset (Callahan et al., 2017).

What is an OTU?

Early 16S pipelines could not tell sequencing errors from real variation, so they grouped similar reads together and treated each group as one taxon. The 97% similarity threshold became a convention because it roughly matched the species level for full-length 16S sequences.

OTUs can be built three ways:

  • De novo: reads are clustered against each other. The clusters depend on which other reads are in the dataset, so the same organism can fall into different OTUs in different studies.
  • Closed-reference: reads are matched to a reference database at 97%. Reads that match nothing are discarded, so novel organisms are lost.
  • Open-reference: a mix of both.

What is an ASV?

Denoising methods such as DADA2 (Callahan et al., 2016), Deblur (Amir et al., 2017) and UNOISE model the errors produced by the sequencer. They then infer which sequences are truly present in the sample. The result is a set of exact sequences, each with a count per sample.

Because an ASV is a real DNA sequence rather than a cluster label, the same organism amplified with the same primers gives the same ASV in any study, on any run.

How do ASVs and OTUs compare?

PropertyOTUs (97%)ASVs
UnitCluster of similar readsExact error-corrected sequence
ResolutionDifferences under 3% are mergedSingle-nucleotide differences kept
Comparable between studiesNo (de novo) or only via a shared referenceYes, directly
Novel taxaLost in closed-reference clusteringKept
Sensitive to sequencing errorsErrors inflate OTU countsErrors are modelled and removed
Typical toolsmothur, VSEARCH, older QIIMEDADA2, Deblur, UNOISE

How does DADA2 infer ASVs?

DADA2 works in four stages:

  1. Learn error rates. It estimates how often each base is misread as each other base at each quality score, from the data itself.
  2. Partition reads. Starting from the most abundant sequence, it asks whether each less abundant sequence could plausibly be an error copy of a more abundant one. If not, it becomes a new ASV.
  3. Merge pairs. Forward and reverse reads are joined into full-length amplicon sequences, which requires them to overlap.
  4. Remove chimeras. Sequences that look like a hybrid of two more abundant parents, artifacts of PCR, are discarded.

Because the error model is learned per run, samples from the same run should be denoised together. In a multi-sample study, BioAnalysis.ca fits one shared error model across all samples for this reason.

Why did ASVs replace OTUs?

Callahan, McMurdie and Holmes (2017) set out the main arguments:

  • Reproducibility: ASVs are consistent labels, so results can be compared and pooled across studies without re-clustering.
  • Reusability: an ASV table from one study can be combined with another that used the same primers.
  • Resolution: single-base differences can separate ecologically distinct strains that a 97% cluster would merge.
  • Accuracy: denoising removes errors that would otherwise inflate richness.

Major tools followed. QIIME 2's denoising options are DADA2 and Deblur, and most published short-read 16S studies now report ASVs.

Do ASVs have drawbacks?

Some. Many bacteria carry several copies of the 16S gene, and those copies are not always identical, so one genome can produce more than one ASV (Schloss, 2021). ASVs are also stricter about rare sequences: a variant seen only once is usually indistinguishable from an error and is discarded. Neither is a reason to go back to OTUs, but both are worth remembering when you count ASVs as if they were species.

When do OTUs still make sense?

  • Comparing new data with an older study that published only OTU tables.
  • Some long-read or high-error workflows where exact denoising is less mature.
  • Analyses that deliberately want coarser units. In that case, aggregating ASVs at genus level is usually clearer than re-clustering.

What does BioAnalysis.ca report?

BioAnalysis.ca denoises with DADA2 and reports ASVs. Every report includes the ASV table, the taxonomy of each ASV against SILVA v138.2 (or UNITE v9.0 for ITS), composition summarized at phylum and genus level, and read counts at each DADA2 stage. See the full 16S pipeline.

References

  1. Amir A, McDonald D, Navas-Molina JA, et al. Deblur rapidly resolves single-nucleotide community sequence patterns. mSystems. 2017;2:e00191-16.
  2. Callahan BJ, McMurdie PJ, Rosen MJ, et al. DADA2: High-resolution sample inference from Illumina amplicon data. Nature Methods. 2016;13:581–583.
  3. Callahan BJ, McMurdie PJ, Holmes SP. Exact sequence variants should replace operational taxonomic units in marker-gene data analysis. The ISME Journal. 2017;11:2639–2643.
  4. Schloss PD. Amplicon sequence variants artificially split bacterial genomes into separate clusters. mSphere. 2021;6:e00191-21.
FAQ

Frequently asked questions

What is the difference between an ASV and an OTU?

An OTU (operational taxonomic unit) is a cluster of reads that share at least a similarity threshold, usually 97%. An ASV (amplicon sequence variant) is an exact biological sequence inferred after correcting sequencing errors. ASVs resolve single-nucleotide differences and are directly comparable between studies; OTUs are not.

Does DADA2 produce ASVs or OTUs?

DADA2 produces ASVs. It models the error profile of each sequencing run and uses it to decide whether a less abundant sequence is a real variant or an error derived from a more abundant one.

Should I still use 97% OTUs?

For new Illumina 16S studies, no. ASVs are now the recommended unit because they are more precise, reproducible and reusable. OTU clustering is mainly useful for comparing against older OTU-based datasets.

Analyze your 16S or ITS data today

Your first 3 analyses are free, with no credit card. Upload FASTQ files and get a publication-ready report, processed and stored in Canada.