Indeed, we have seen that its performance in distinguishing human from murine sequences is usually higher than the methods relying on just the sequence similarity, due to the fact that this statistical model, on which the MG-score is based, accounts for pair-correlations between residues at different positions
Indeed, we have seen that its performance in distinguishing human from murine sequences is usually higher than the methods relying on just the sequence similarity, due to the fact that this statistical model, on which the MG-score is based, accounts for pair-correlations between residues at different positions. of clinically used antibodies. Finally, we use the humanness score as an optimization function and perform a search in the sequence space, starting from different murine sequences and keeping the CDR regions unchanged. Our results show that our humanness score outperforms other methods in sequence classification, and the optimization protocol is able to generate humanized sequences that are recognized Phenolphthalein as human by standard homology modelling tools. == Introduction == Antibody-based drugs have acquired an increasing importance in the last two decades, both for imaging and for therapeutic uses, especially to treat different types of cancer and autoimmune diseases. However, their development is usually a long and difficult process, prone to fail at different stages. Antibody humanization Phenolphthalein is usually MTC1 a key step in this process, unless the candidate is already obtained from a human library, and is essential in moving from the preclinical to clinical stage. In fact, new antibodies are typically developed in animal models (most often, in mouse); however, the antibodies obtained by this way are usually not tolerated by humans, elicitingin vivoan immune response against the murine antibody. Thus, they need to be humanized, substituting a part of their sequences by the human ones, while preserving their specificity, affinity and stability. Although computational methods are available, nowadays such humanization process is mostly a trial-and-error process, based on CDR-grafting and back mutations1. CDR-grafting implies selecting the Complementarity-Determining Regions (CDRs), responsible for antigen recognition, from the given murine sequence and grafting them into the human Framework Region Phenolphthalein (FR); the latter is usually selected by looking, in the human genome, at the germlines that produce FRs most homologous to the murine ones: the hope is that the combination of such human frameworks with the original murine CDR will result in a molecule that still preserves its stability and activity, but is usually tolerated by the human immune system. However, most of the times this approach is not completely successful at either of its quests, and the researcher is usually left alone in trying further mutations, until an antibody with the selected properties is usually identified. This error-prone process is usually a true bottleneck in the development of new treatments, in a market of increasing global impact. From an algorithmic point of view, CDR-grafting corresponds to a search, in the human germlines, for the sequence with minimal Hamming Distance (i.e highest similarity) to the original murine one. Thus, the similarity to the closest germline represents a humanness score to maximize, for this approach. More in general, any humanization protocol will rely on maximizing some humanness score, whose basic requirement is to be able to distinguish human from mouse sequences with as few errors as possible. Most humanness scores are based on pairwise sequence identity between the sample and a set of reference (most often, germline) human sequences: for instance, the score can correspond to the average similarity2, or the average among the top 20 sequences3, or the highest similarity, over windows of typically 9 residues4,5. Recently, a different approach as been proposed by Seeliger6, introducing a score function that accounts both for local preferences and for pair correlations between residues at different positions. Interestingly, such approach is reasonably capable to distinguish between human and mouse sequences, despite a relevant residual overlap between their distribution; also, the stochastic humanization process under such score function samples regions of low immunogenicity, even though the final basin of attraction of the trajectories presents an intermediate value of immunogenicity (as measured by the Epivax score)7. It is also worth mentioning that such approach breaks with the common logic of considering CDRs as the only antigen binding regions, by treating correlations between any positions (CDR or framework) on the same grounds, which can be a safer option due to the fact that there are relevant antigen binding residues also in framework regions8. However, Seeligers approach uses an ad hoc score function for Phenolphthalein the pairs of residues,, that apparently is Phenolphthalein usually loosely related to mutual information, and may suffer its same problems9, in distinguishing direct and indirect correlations. A more fundamental approach deals with the observed sequences.