Structure-based antibody lead selection: how to find the candidates sequence hides
How structure prediction, clustering, and developability analysis turn millions of NGS signals into a smaller, more diverse, more developable lead panel.
When researchers sort through millions of NGS hits using only 1D sequence information, they fail to capture a vital piece of the puzzle. A sequence-only approach can hide functional redundancy, overlook promising candidates, and allow false-positive liabilities to slip through.
The future of antibody discovery requires moving beyond sequence-based selection and instead, prioritizing candidates based on their true functional properties and biological potential.
Why is Structural Prediction Critical for Antibody Lead Candidate Selection?
Relying solely on 1D sequence data creates a blind spot in antibody lead selection. Sequence-based approaches can make diverse candidates appear redundant, and also classify safe antibodies as potential liabilities based on risks that may not exist in the folded structure.
By predicting antibody 3D structures, researchers can group candidates based on their actual physical properties, identify functionally redundant binders, and evaluate developability based on whether potential liabilities are exposed on the surface. This enables more informed lead selection and improves the odds of successful wet-lab validation.
The Illusion of Sequence Diversity
In an antibody discovery program targeting ovarian cancer, we analyzed a lead panel of 300 nanobodies that appeared highly diverse based on sequence alone. The candidates were broadly distributed across the repertoire (Figure 1), and a traditional sequence-based clustering approach would have suggested that all 300 represented distinct candidates.
Then we predicted the 3D structure of each nanobody and clustered them by shape. The 300 “diverse” sequences collapsed into just 6 distinct structural clusters. Two of these structural groups accounted for approximately 95% of the abundance across the entire panel. What appeared to be a diverse panel was actually dominated by a small number of repeated structural solutions.
This has direct consequences for validation strategy. Selecting candidates based only on sequence diversity can lead teams to spend valuable synthesis and characterization resources testing many antibodies that are structurally redundant. With the fully loaded cost of synthesizing, expressing, and characterizing a single antibody candidate reaching approximately $1,000 (excluding time), selecting the wrong candidates can quickly multiply costs across a discovery campaign.
Why sequence diversity can be misleading
The root cause is simple: different sequences can fold into the same shape. Antibody function is largely determined by the three-dimensional arrangement of the complementarity-determining regions (CDRs), the loops that form the antigen-binding surface. Two antibodies with different sequences can adopt similar CDR conformations and, as a result, may recognize overlapping epitopes or exhibit similar functional behavior.
Sequence clustering can’t see this. It groups antibodies based on amino acid similarity rather than the shape of the binding surface. As a result, structurally similar antibodies can appear distant in sequence space, while a panel that looks diverse on paper may contain significant hidden redundancy.
In one study, the Oxford Protein Informatics Group found that clustering antibodies by predicted CDR structure identifies common-epitope binders far more completely than sequence-based clustering.1 Their conclusion: structural data “contains orthogonal functional information to sequence.”
Why sequence-based liability scoring over-flags
Sequence limitations also affect developability assessment.
Many antibody screening workflows identify liabilities by searching for sequence motifs associated with risks such as deamidation, oxidation, glycosylation, or unpaired cysteines. However, the mere presence of a motif does not necessarily mean the antibody will experience a developability issue. This, in fact, depends on whether that residue is exposed and accessible in the folded structure. Without structural context, sequence-based approaches can overestimate risk and eliminate potentially valuable candidates.
Conversely, many important developability risks are inherently structural and cannot be identified from sequence alone. Surface hydrophobic patches, charge asymmetry, unusual surface exposure patterns, and CDR-H3 conformational properties all influence antibody behavior but require a 3D view of the molecule.
These are among the features incorporated into the Therapeutic Antibody Profiler (TAP)2, developed by the Oxford Protein Informatics Group, which established structure-based developability metrics using clinical antibody datasets. Structure provides the missing layer of information needed to distinguish real liabilities from false positives and identify new risks that sequence-only approaches miss.
How to integrate 3D structures into your NGS workflow
To turn millions of sequencing signals into confident lead selection decisions, structural analysis should become part of the standard discovery workflow, rather than being treated as a secondary step.
Instead of taking the place of sequence-based analysis, structural prediction adds an additional layer of information after the initial sequence filter and before the expensive commitment to synthesis. A typical workflow looks like this:
1. Initial sequence filtering. Run your standard initial filtering to collapse millions of clonotypes into a workable lead set. Sequence remains the fastest and most cost-effective first-pass filter.
2. Predict structures. Generate a 3D model for every remaining candidate. Modern deep learning tools such as ImmuneBuilder3 can predict antibody structures in seconds. Each model also includes a per-residue confidence score. This is important because while framework regions are typically predicted with high confidence, the CDR-H3 loop (the most important region for antigen binding) is also the most flexible and difficult to predict. Always consider the confidence score when interpreting individual structural features.
3. Cluster by structure. Group antibodies based on their predicted binding geometry rather than sequence similarity. Structural clustering identifies redundant binders that sequence clustering misses, allowing you to retain representatives of truly distinct structural families. Keep the best representative from each cluster, remove the near-duplicates, and your panel is diverse across distinct binders and distinct epitopes.
4. Scan for liabilities. Use the same models to assess developability from the folded surface: report only the liability motifs that are genuinely exposed, and add the 3D-only metrics sequence can’t produce: hydrophobic surface patches, charge distribution, and other developability metrics described by the Therapeutic Antibody Profiler (TAP). Even though it’s not ground truth, it provides a much more informative assessment than sequence motifs alone.
5. Construct the panel. Build the refined lead panel with all the clustering and structural information included — selecting on enrichment quality, real structural diversity, and surface developability together, then advancing to synthesis and validation.
Most discovery teams only prioritize abundance because it’s the easiest signal to measure, but that doesn’t tell you whether two antibodies represent different biological solutions or whether they’ll behave well as therapeutics.
Better lead selection comes from combining every available signal at the point of decision. Sequence tells you what mutations occurred. Structure tells you what those mutations produced.
Structure-based lead selection in Platforma
None of the workflow above requires exporting data into separate tools or building custom analysis pipelines.
In Platforma, structure prediction, structural clustering, and developability analysis are integrated directly into the antibody discovery workflow. After initial lead selection, Platforma predicts antibody structures with the 3D Structure Prediction block, groups candidates by structural similarity using 3D Structure Clustering, and evaluates surface-based developability with 3D Structure Liabilities, and carries every result directly into final lead selection.
Because every step runs inside the same reproducible pipeline, researchers can trace every decision without moving data between disconnected applications or relying on custom scripts. Scientists can combine sequencing, structure, and developability into a single decision process.
Most antibody software helps teams manage experiments or process sequencing data. Platforma helps teams make discovery decisions.
Ready to see what structural analysis reveals in your own antibody panel?
Try Platforma on your own sequencing data and discover how structure-based lead selection can uncover hidden redundancy, improve developability assessment, and help you build a stronger validation panel.
Spoendlin FC, Abanades B, Raybould MIJ, Wong WK, Georges G, Deane CM. “Improved computational epitope profiling using structural models identifies a broader diversity of antibodies that bind to the same epitope.” Frontiers in Molecular Biosciences 10:1237621 (2023). https://www.frontiersin.org/journals/molecular-biosciences/articles/10.3389/fmolb.2023.1237621/full
Raybould MIJ, Marks C, Krawczyk K, Taddese B, Nowak J, Lewis AP, Bujotzek A, Shi J, Deane CM. “Five computational developability guidelines for therapeutic antibody profiling.” PNAS 116(10):4025–4030 (2019). https://www.pnas.org/doi/10.1073/pnas.1810576116
Abanades B, Wong WK, Boyles F, Georges G, Bujotzek A, Deane CM. "ImmuneBuilder: Deep-Learning models for predicting the structures of immune proteins." Communications Biology 6:575 (2023). https://www.nature.com/articles/s42003-023-04927-7



