Posted on May 5, 2025
There is a limited set of known human germline framework sequences (~150), but CDRs, which dictate antigen recognition, show more variation (1)
There is a limited set of known human germline framework sequences (~150), but CDRs, which dictate antigen recognition, show more variation (1). m unique amino acid sequences from ca. 600 individuals. Despite the great theoretical diversity of antibodies, we find that the majority of sequences coming from such studies can be reliably mapped to an existing structure. Keywords:antibody specificity, B-cell receptor, next-generation sequencing, structural homology, protein, bioinformatics tools == Introduction == Antibodies are proteins that play a key role in recognizing potentially noxious molecules (antigens) in jawed vertebrates. They are produced by B-cells, where they can be secreted or act as a membrane-bound B-cell receptor. In humans, they are composed of two polypeptide chains, referred to as heavy and light. Each of the chains has a variable region that is ca. 110 amino acids long, composed of the framework region and three hypervariable loops referred to as complementarity determining regions (CDRs). There is a limited set of known human germline framework sequences (~150), but CDRs, which dictate antigen recognition, show more variation (1). It is estimated that a typical human is capable of producing more than 1010distinct antibody molecules (26). Thus, in a single individual, there is likely to exist Heparin an antibody capable of recognizing an arbitrary antigen, though perhaps not specifically. Such binding malleability of antibodies has long been a subject of intensive academic and industrial research. Discerning human antibody diversity will help us to understand how our immune system is capable of recognizing such a myriad set of antigens and underpins our ability to exploit them therapeutically (710). Next-generation sequencing of immunoglobulin genes (Ig-seq) facilitates this task as it allows us to obtain a snapshot of the B-cell receptor (antibody) repertoire across different individuals and immune states (1116). The outputs from these Heparin Ig-seq experiments have been characterized by their germline biases and sequence analysis methods (2,10,14,17,18). These studies do not consider the three-dimensional structure of the antibody, but it is this structure that dictates antigen recognition (19,20). In one study, the authors structurally characterized a small portion of their Ig-seq data (ca. 2,000 structural models from ca. 175,000 sequences) but they did not produce TLR4 a structural annotation protocol (21). In this paper, we show it is possible to characterize structurally large percentages of the data and describe Heparin a pipeline for automating the task. As described by Kovaltsuk et al. (20), structural information can give both an overall predicted shape and detail of the binding site (CDRs). Predicting the shape of a sequence can offer sufficient information to link it to an antibody with similar shape and defined antigen specificity (22,23). Enriching Ig-seq datasets with structural information should improve analyses and insights that can be derived from antibody repertoire snapshots. Here, we describe the structural annotation of antibodies (SAAB) algorithm to bridge the sequence-structure gap in antibody repertoire analysis. This protocol, given a FASTA file with potentially millions of antibody sequences, maps the full sequences, frameworks, and CDRs to the high quality antibody structures currently available in the Protein Data Bank (PDB) (24,25). We demonstrate the validity of this approach by testing the protocol on five separate Ig-seq datasets encompassing ca 35 m sequences from ca. 600 individuals. For each dataset, we can associate a majority of frameworks and CDR sequences to an existing antibody structure. This finding recapitulates on a large scale both the structural conservation of the framework and the canonical CDR paradigm (26,27). More generally, however, we demonstrate that it is currently possible to approximate the structures of entire variable region sequences for most of the data. Therefore, despite the theoretically allowed repertoire diversity, currently observed antibody sequence space appears to employ only a conservative set of structural shapes. == Materials and Methods == == Structural Annotation of Antibodies == Our SAAB algorithm accepts amino acid sequences in FASTA-formatted input. The algorithm first Chothia-numbers the sequences (28) and then maps (if possible) the sequences to known antibody structures in the PDB (25). These structures are identified for entire variable region as well as for frameworks and CDRs separately. Details of the steps of the protocol are given below. == Chothia-Number Sequences == Each of the supplied Ig-seq amino acid sequences are Chothia-numbered (28) using ANARCI (29). The numbering provides a consistent frame of reference for antibody sequences. This allows for sequence identity calculation of the entire sequence as well as regions of the antibody separately (frameworks and CDRs). == Antibody Structural Reference == The set of antibody structures accompanying our software and used in this analysis was downloaded from the structural antibody database (SAbDab) on 31st October 2017 (25). Only X-ray structures with resolution better than 3.0 were used. VHH structures were excluded. If both heavy and Heparin light.
Recent Comments