How to Analyze Oxford Nanopore (ONT) Data for Antibody Discovery
A complete workflow for processing, annotating, clonotyping, and analyzing full-length paired antibody sequences from ONT long-read sequencing.
Oxford Nanopore Technologies (ONT) sequencing is increasingly being used for antibody discovery because its long reads can capture full-length antibody sequences and, in the right experimental designs, preserve native VH–VL pairing in a single molecule.
To turn ONT antibody sequencing data into useful discovery insights, you need to go beyond reading the sequences. A typical workflow includes UMI-based consensus generation, antibody annotation, VH–VL pairing, clonotyping, enrichment analysis, clustering, and candidate selection.
This article explains how to analyze Oxford Nanopore data for antibody discovery, what each step does, where conventional workflows become difficult, and how to take the analysis from raw long reads to lead antibody candidates in Platforma.
What makes Oxford Nanopore data useful for antibody discovery?
The defining advantage of Oxford Nanopore sequencing is read length.
A single Nanopore read can span an entire antibody variable-domain transcript. This can preserve the native pairing of the heavy and light chains (VH–VL) from the same molecule, something that is difficult to recover from conventional short-read sequencing where VH and VL are sequenced as separate fragments.
Full-length reads also reduce the need to assemble antibody sequences from shorter fragments.
This makes ONT particularly useful for workflows where the complete antibody sequence and chain pairing are important, including:
Phage and other display libraries
Hybridoma sequencing
Single B-cell and paired BCR repertoires
In vivo antibody discovery from immunized animals
ONT is also increasingly practical for antibody workflows as sequencing accuracy improves. R10.4.1 chemistry (Kit V14) delivers >99% single-read accuracy, reaching 99.75% (Q26) with current Dorado basecalling models [1]. UMI-based consensus pushes this much further: combining UMIs with concatemeric consensus sequencing has been shown to reach Q51.7 on immunoglobulin heavy-chain amplicons, effectively removing sequencing error as a source of false variants [2]. This sequencing technology can capture the antibody molecule in a single read, at an accuracy that supports calling real CDR variation.
Since 2025, Oxford Nanopore has also supported 10x's Chromium GEM-X 3' and 5' chemistries, so you can run long reads on 10x single-cell libraries to recover full-length isoforms [3], keeping single-cell partitioning while stopping short reads from throwing away the full-length paired sequence.
How deep do you need to sequence ONT for antibody discovery?
Depth is the most consequential decision in an ONT antibody campaign. It’s made before any analysis happens, and it’s where we most often see money spent for nothing.
The key point: consensus accuracy is driven by reads per molecule, not by total reads. Once you have enough reads behind each UMI or cell barcode to build a confident consensus, further depth is redundant. In practice a few dozen good reads per chain is enough for a confident ONT consensus.
So size the run from the molecules, not from a headline number:
reads needed ≈ cells recovered × chains per cell × 30–50 reads per chain ÷ usable-read fraction
For a 10,000-cell run at two chains per cell and ~50% usable reads, that’s roughly 1.5–2M reads. For 100,000 cells, roughly 15–20M. Run the same arithmetic on a campaign you’ve already sequenced: divide total reads by (cells recovered × chains per cell), then compare against those few dozen. We routinely see it land 10× higher — thousands of reads per chain per cell where dozens would have done.
What does an ONT antibody sequencing analysis pipeline involve?
A complete Oxford Nanopore antibody analysis workflow typically looks like this:
Group reads by UMI and generate consensus sequences to reduce sequencing and PCR errors.
Annotate antibody sequences, including V, D, and J genes and CDRs.
Identify VH–VL pairs from full-length reads.
Assign clonotypes and collapse redundant sequences.
Cluster to understand diversity.
Analyze enrichment (display) or clonal expansion + SHM (in vivo).
Prioritize candidates for downstream structural analysis, developability assessment, and experimental validation.
The challenge is that these steps are often handled by different tools, pipelines, or custom scripts. Moving data between them can make an otherwise powerful long-read experiment surprisingly difficult to analyze reproducibly.
Step 1 — Generate UMI consensus sequences
Raw Nanopore reads aren’t always accurate enough to use directly. The fix is unique molecular identifiers (UMIs): tag each original molecule before amplification, sequence it many times, group reads by UMI, and build a consensus. Random errors wash out; the true sequence remains.
Two distinct things are happening at this stage, and it’s worth keeping them separate:
Barcode correction recovers molecules whose UMI or cell barcode was misread. It determines how many of your molecules you keep.
Consensus accuracy comes from collapsing many reads of the same molecule into one sequence. It determines how correct each sequence you keep actually is.
Tuning one does not improve the other, and conflating them is a common source of misdiagnosis when a run looks noisier than expected.
This matters for antibody work because a single base error can change a CDR or invent a false variant. For display-library selections, UMI consensus is also what lets you separate genuine diversity and enrichment from sequencing artifacts. ONT with dual UMIs has been used to track diversity and enrichment before and after each panning round, recovering rare binders even when the dominant clone exceeded 99% abundance [4].
Step 2 — Annotate antibody sequences
Determine what each consensus sequence is: V, D, and J gene segments, CDR and framework regions, and sequence-level features. Annotation turns nucleotides into interpretable antibodies, but on its own it doesn’t tell you whether a sequence is enriched, unique, redundant, or worth testing.
Step 3 — Identify VH–VL pairs
This is where long reads pay off. When a read spans the paired chains, you know which heavy chain belongs to which light chain: that’s the difference between knowing two sequences and knowing the actual antibody. It’s what makes ONT valuable for scFv/Fab libraries, hybridomas, single B cells, and paired BCR repertoires.
Step 4 — Assign clonotypes
Collapse related sequences into clonotypes while correcting residual PCR/sequencing error. Read counts alone mislead: one clone can appear thousands of times through amplification, and many near-identical reads may be the same underlying clone. Clonotypes are the meaningful unit for everything downstream.
Step 5 — Cluster to understand diversity
If the top 20 sequences are near-identical, testing all 20 tells you little. Clustering groups related antibodies so you can build a panel from distinct families rather than 20 versions of the same binder, cutting redundancy and covering more of the response.
Step 6 — Analyze enrichment or clonal expansion
For display experiments, the question isn’t “which antibodies are present?” but “which are becoming more abundant across selection rounds?” A clone that’s abundant in one sample isn’t necessarily interesting; a cluster that consistently expands round over round is much stronger evidence of real selection. Enrichment analysis then separates genuinely enriched families from merely abundant ones, (and with a negative control) antigen-specific binders from sticky, non-specific ones.
For in vivo campaigns, there’s nothing to enrich across. The signal comes from biology instead: clonal expansion plus somatic hypermutation (SHM), the fingerprint of affinity maturation.
Step 7 — Prioritize candidates
The goal is a shortlist worth validating. Once you have clonotypes, clusters, and either enrichment or SHM/expansion evidence, you screen the survivors for developability liabilities, select the strongest candidate from each cluster so the panel stays diverse.
Structure-based analysis can then be applied to a pre-selected subset of candidates to assess structural diversity and developability. [5]
Parameters reference
For MiXCR, the relevant presets are:
ONT, no molecular barcodes: generic-ont
ONT with UMIs: generic-ont-with-umi
PacBio equivalents: generic-pacbio, generic-pacbio-with-umiFor scFv constructs, MiXCR locates the chains using the linker sequence, so the exact linker nucleotide sequence for your library is required. The example below, ggtggaggtggctctggcggtggcggatcg, encodes (G₄S)₂ — 30 nt. Many scFv libraries instead use (G₄S)₃ (45 nt), and variants are numerous, so substitute your own construct’s linker: an approximate sequence is a leading cause of low alignment rates.
Full-length long-read scFv (heavy–linker–light):
Construct building order: Heavy-linker-light
Linker: ggtggaggtggctctggcggtggcggatcg
Heavy chain tag pattern: ^(R1:*)ggtggaggtggctctggcggtggcggatcg*
Light chain tag pattern: ^*ggtggaggtggctctggcggtggcggatcg(R1:*)
Heavy assembling feature: FR1 - FR4
Light assembling feature: FR1 - FR4Adding a 12 nt UMI — prepend the tag to both patterns; the same tag must appear in both for pairing to work:
Heavy chain tag pattern: ^(UMI:N{12})(R1:*)<linker>*
Light chain tag pattern: ^(UMI:N{12})*<linker>(R1:*)Where ONT is still the wrong choice
If you can’t run UMIs or per-cell consensus, PacBio HiFi remains the safer long-read choice. ONT’s accuracy advantage in antibody work comes substantially from consensus. Without it you’re relying on single-read accuracy, where the margin is thinner.
Reads per molecule is the binding constraint, not total reads. If your protocol can’t deliver enough reads behind each UMI or cell, no amount of total depth will fix it — see the depth section. This is the failure mode that masquerades as “ONT isn’t accurate enough.”
The error profile is indel-biased. R10.4.1 improved homopolymer accuracy substantially, but it remains worth inspecting CDR3 length distributions for indel artifacts rather than assuming a substitution-style error model.
Long reads don’t create pairing that isn’t in the construct. If your library doesn’t physically link VH and VL in one molecule, read length won’t recover it.
Where analysis workflows get difficult
A real workflow often spans basecalling, UMI processing, consensus, annotation, clonotyping, custom scripts, enrichment, clustering, and selection across several tools and file formats.
Every handoff between tools is a place for the analysis to drift: reprocessed files, forgotten parameters, lost metadata. By the time you reach candidate selection, the decision that matters most is the hardest to reproduce.
Tools exist for parts of the workflow — NAb-seq for hybridomas and single B cells [6], Seq2scFv for annotating and quantifying full-length scFvs from long reads across panning rounds [7], UMI-based consensus methods for high-accuracy amplicon sequencing [2], and MiXCR for annotation and clonotyping. Stitching them together is where reproducibility breaks down.
Among no-code platforms, Geneious Biologics’ Antibody Annotator accepts PacBio and Nanopore long reads too, though in long-read mode it turns off antibody numbering and records only amino-acid-level variants (not nucleotide-level) relative to the reference.
Analyzing ONT antibody data end-to-end in Platforma
Platforma is the only software that turns your ONT data into a decision matrix — evidence-based lead selection from your entire NGS dataset.
The same MiXCR engine many labs already use for immune-repertoire sequencing powers both in-vivo and in-vitro antibody workflows. It takes full-length long reads from Oxford Nanopore and PacBio directly, with native UMI support for deduplication and consensus error correction [7]. From there, paired clonotypes flow into study-level QC, clustering, enrichment analysis or clonal expansion, structure-based refinement, and candidate prioritization, which ranks candidates on a composite In Vivo Score for immunized-animal campaigns, or on enrichment quality across panning rounds for display campaigns, while enforcing sequence diversity across the final panel.
Ready to turn your Nanopore data into a lead panel of candidates? Bring your ONT antibody data into Platforma and go from raw long reads to a developable lead panel in one reproducible pipeline.
References
“Nanopore sequencing accuracy.” Oxford Nanopore Technologies. https://nanoporetech.com/platform/accuracy
“R2C2 + UMI: Combining concatemeric and unique molecular identifier–based consensus sequencing enables ultra-accurate sequencing of amplicons on Oxford Nanopore Technologies sequencers.” PNAS Nexus 3(9):pgae336 (2024). https://academic.oup.com/pnasnexus/article/3/9/pgae336/7737794
"Oxford Nanopore expands compatibility with 10x Genomics to unlock deeper insights in single-cell transcriptomics." Oxford Nanopore Technologies. https://nanoporetech.com/news/oxford-nanopore-expands-compatibility-with-10x-genomics-to-unlock-deeper-insights-in-single-cell-transcriptomics
“Deep mining of antibody phage-display selections using Oxford Nanopore Technologies and Dual Unique Molecular Identifiers.” New Biotechnology 80:56–68 (2024). https://doi.org/10.1016/j.nbt.2024.02.001
"Structure-based antibody lead selection: how to find the candidates sequence hides” https://blog.platforma.bio/p/structure-based-antibody-lead-selection
“NAb-seq: an accurate, rapid and cost-effective method for antibody long-read sequencing in hybridoma cell lines and single B cells.” bioRxiv (2022). https://www.biorxiv.org/content/10.1101/2022.03.25.485728v1.full
“Seq2scFv: a toolkit for the comprehensive analysis of display libraries from long-read sequencing platforms.” mAbs 16:2408344 (2024).
“Annotating scFv libraries.” Platforma documentation. https://docs.platforma.bio/guides/antibody-discovery/scFv-clonotyping/


