From Sequence to Function The structural biology of therapeutic peptides
Dissolve a twenty-residue peptide in water and titrate it. Five of its side chains are histidine — the same imidazole ring, five copies of it — and they release their protons at five different pH values: 4.89, 5.62, 5.94, 6.44 and 6.80. Nothing distinguishes them but where they sit in the chain. At pH 6 some are charged and some are not, and no table of amino-acid properties will tell you which. That measurement is this document's starting point and its argument in miniature. A peptide is a chain of amino acids joined by one repeated bond, and what any part of it does is decided by the rest.
01Twenty letters and one bond
The titration in the lede is a real experiment. A twenty-residue fragment of the salivary protein MUC7 was dissolved at 0.4 mM in 100 mM sodium perchlorate, blanketed with argon at 298 K, and titrated by potentiometry from pH 2 to pH 12; the refinement resolved ten separate deprotonation constants (Szarszoń et al., 2026). Three glutamate carboxylates came off at pH 3.36, 3.58 and 4.21. Five histidine imidazoles came off at 4.89, 5.62, 5.94, 6.44 and 6.80. The N-terminal ammonium group came off at 7.67 and the single lysine side chain at 10.07. A second study by the same approach, on a different MUC7 fragment, found the same thing on a smaller scale: two histidines at 5.83 and 6.64, two lysines at 10.17 and 10.87, a single aspartate at 3.75, the C-terminal carboxyl at 2.72 and the N-terminal ammonium at 7.41 (Ślusarczyk et al., 2026). Chemically identical groups, in one short chain, differing by up to two pH units in when they ionise.
This is not a curiosity but the argument in miniature. Whether a side chain is charged at a given pH is one of the few properties a reader might reasonably expect to be able to look up. It is a property of the imidazole ring, and all five of those rings are the same ring. The answer is nonetheless a property of position in a sequence rather than of the residue's name, and what is true of a proton is true of everything else the chain does.
So it is worth being exact about what a peptide is. Each amino acid carries, on a single central carbon, four things: an amino group, a carboxyl group, a hydrogen, and a side chain that is different in each of the twenty. Two of them join when the amino nitrogen of one attacks the carboxyl carbon of the other, expelling a molecule of water and leaving an amide linkage — the peptide bond. The reaction has a geometric requirement that has been stated precisely from structural and computational analysis of the ribosome's own catalytic centre: the incoming amine must come within 4 Å of the carbonyl carbon it attacks, along a trajectory 76–115° off the plane of that carbonyl (Katoh & Takada, 2026). Everything downstream — every fold, every receptor contact, every proteolytic failure — is a consequence of that one linkage repeated, and of the twenty things hanging off it.
Because the two ends of the linkage are chemically different — one supplies a nitrogen, the other a carbonyl — a chain has a direction. That is a chemical fact, not a drafting convention. The ribosome builds strictly from the amino terminus to the carboxyl terminus, and it builds unevenly: a reconstituted Escherichia coli translation system yielded 0.20 µM for two consecutive D-alanines and 0.08 µM for two consecutive N-methyl-leucines against 0.31 µM for ordinary L-alanine, and spacing the same residues apart recovered much of the loss, so the run rather than the count is the obstacle (Katoh & Takada, 2026). Destruction reads the direction too. Trypsin cuts only on the carboxyl side of lysine and arginine (Szarszoń et al., 2026), and exopeptidases work inward from a free end, which is why closing a chain head-to-tail protects it: of six peptides incubated in concentrated Madin–Darby canine kidney cell lysate at 37 °C, five linear L-forms were destroyed within 2–6 h while the one cyclic L-peptide survived beyond 23 h (Juraszek et al., 2025).
What hangs off the backbone is conventionally sorted into acidic, basic, polar and non-polar, which is a chemist's filing system rather than a functional one. It is more useful to sort the twenty by what they do to a chain. A first group carries charge, conditionally: aspartate and glutamate in the low threes to low fours, histidine in the fives and sixes, lysine around ten — every one of them shifted by its neighbours, as the MUC7 titrations show. A second group repels water and drives the chain to hide it. Here the honest position is that this document's evidence base carries no per-residue hydropathy table and no volume scale at all. The Kyte–Doolittle scale that gave the field its standard numerical treatment of hydropathy along a sequence (Kyte & Doolittle, 1982) appears in the corpus only as aggregate scores for whole sequences — seventeen disordered loops spanning 2.00 to 4.06 on a normalised scale — which predicted neither the stability nor the solubility of the proteins they were grafted into (Ripka et al., 2021).
A third group donates and accepts hydrogen bonds, with a geometric fussiness that only high-resolution structures reveal: in the conotoxin Mu8.1, solved by X-ray crystallography at 1.67 Å, five threonines at the dimer interface each donate a bifurcated hydrogen bond to the backbone carbonyl four residues along, every one locked in the same rotamer with χ1 near −60° (Müller et al., 2023) — five identical residues doing one identical job, the exact opposite of the five histidines. A fourth group breaks the rules the others follow: glycine has no side chain beyond a hydrogen and so no chiral centre at the α-carbon, while proline's side chain loops back onto its own backbone nitrogen, closing a ring that restricts the chain rather than adding to it. A fifth forms covalent bridges, and a peptide that needs them may not be obtainable without help — Mu8.1 had to be expressed alongside a co-expressed disulfide-formation system to fold at all.
One more thing about the twenty: they are not twenty. A cheminformatics screen of the ChEMBL database found 3,845 monomers that could be inserted into a chain by ordinary peptide-bond formation, of which 3,179 are absent from the standard structural dictionary and cannot be represented by current structure-prediction pipelines at all (Han et al., 2026). Nor is this exotic: one incretin lipopeptide solved by X-ray crystallography at 1.59 Å carries six non-proteinogenic residues, including a D-glutamic acid buried mid-chain and an acyl conjugate hung off a lysine (Mitchell et al., 2026), and about 30 % of approved peptide drugs contain at least one D-amino acid, though none is entirely D-configured (Juraszek et al., 2025).
The reader who has followed this far should now expect something specific. Asked what a lysine does in a peptide, the correct first answer is no longer "carries a positive charge" but "depends where it is" — and the second section will show that the same is true of the backbone itself, which turns out to have preferences that shift by tens of percent depending on nothing but the identity of the residue next door.
02Why the bond is stiff, and why that is the whole story
Here is how the stiffness of the peptide bond appears in modern practice. When a structure-prediction pipeline was benchmarked on short test peptides carrying non-canonical residues, the validation rule applied to the output was binary: a predicted structure was scored "good" if the dihedral angle ω about each amide bond satisfied |ω| ≥ 150° or |ω| ≤ 30°, "borderline" if any ω fell between 120° and 150°, and "bad" if any fell between 30° and 120°. Across the test set 85.78 % of predictions landed in the good band (Han et al., 2026). The whole intermediate range — a third of the circle — is treated as an error rather than as a conformation. That is what a partial double bond means, expressed as a piece of software.
The chemical reason is that the amide's electrons are not localised where the formula suggests: part of the nitrogen's lone pair delocalises onto the carbonyl, giving the C–N linkage partial double-bond character and the carbonyl correspondingly less. The corpus assembled here contains one direct experimental handle on that redistribution, and it is indirect — the amide I′ vibration of a primary amide sits about 20 cm⁻¹ above that of a secondary amide when the two are compared in water (Schweitzer-Stenner, 2026). Beyond that, the evidence base gives no C–N bond length for an amide inside a chain and no measurement of the rotational barrier about that bond in energy units. Those figures exist in the physical-chemistry literature; they are not in this document's reading, and they are therefore not printed here.
What follows from the plane is arithmetic. A chain of thirty residues, if every bond rotated freely, would have some ninety degrees of backbone freedom. It has about sixty: two per residue, the rotation φ about the nitrogen-to-α-carbon bond and the rotation ψ about the α-carbon-to-carbonyl bond. A third of the chain's conformational freedom is gone before any side chain has done anything. In 1963 Ramachandran and colleagues showed that most of what remains is gone too, removed by steric collision between atoms the two rotations bring into contact, leaving a small and characteristically shaped permitted minority (Ramachandran et al., 1963); the analysis was extended and formalised two years later (Ramakrishnan & Ramachandran, 1965). The same group later asked whether the planarity assumption underlying the whole construction was exactly true, and argued that it is not (Ramachandran et al., 1968).
Current work on unfolded peptides partitions the permitted region into seven named mesostates with explicit boundaries: polyproline II between φ of −90° and −42° with ψ between 100° and 180°; antiparallel β-strand between −180° and −130° with ψ from 130° to 180°; a transition band between them; parallel β-strand just below; an extended region along the bottom edge; and a turn-like region subdivided into seven sub-regions, four of which sit at positive φ where only glycine goes comfortably (Suresh et al., 2026). These are labels for a single residue's backbone angles, not for secondary structures, which require several consecutive residues to agree. That distinction matters more than it sounds: a residue can sit in the polyproline-II box without any of its neighbours doing so, and calling that "structure" would be a category error.
Two residues redraw the map, in opposite directions. Glycine, lacking a β-carbon, escapes most of the steric collisions that define the boundaries and reaches regions no other residue can. In the 1.59 Å crystal structure of the incretin analogue mentioned earlier, the single Ramachandran outlier in the entire model was Gly30, with density clear enough that no alternative conformation was possible (Mitchell et al., 2026) — an anecdote, but a clean one. Proline goes the other way, its ring fixing φ within a narrow range and removing the backbone amide hydrogen the rest of the chain uses for hydrogen bonding. Here again the honest statement is negative. This evidence base carries no measurement of the cis-to-trans proline population in any peptide, no isomerisation rate, and no quantification of how much of the map glycine actually gains. Those are standard textbook quantities, and they are not in this document's reading.
Now the correction the modern literature insists on, and the first appearance of the idea the whole document rests on. The classical treatment described an unfolded chain as a "random coil": one broad permitted basin within which all conformations were treated as effectively equivalent in energy. Direct measurement says otherwise. Five scalar coupling constants from solution NMR, together with infrared, polarised Raman and vibrational circular dichroism profiles of the amide I′ band, all on short host–guest peptides in water and fitted to a two-dimensional Gaussian model of the map at 2° resolution, resolve that basin into two confined minima separated near φ = −100°: polyproline II on one side, β-strand on the other, with different free energies for different residues (Schweitzer-Stenner, 2026). The unfolded state is not random. It is a statistical coil, carrying measurably less conformational entropy than the random coil it replaces, which survives as a special case rather than the general one.
The neighbour effects are what should change how a reader reads a sequence. Averaged over unlike neighbours in four-residue host peptides in water, alanine's polyproline-II population falls by about 29 %, aspartic acid's rises by about 50 % and phenylalanine's by about 30 %, while serine, leucine and lysine fall only 12–15 %; a serine neighbour destabilises alanine's polyproline-II conformation by roughly 4 kJ/mol (Schweitzer-Stenner, 2026). Like neighbours reinforce instead: the central alanine of a glycine-flanked host sits at 58 % polyproline II, and of three consecutive alanines above 75 % (Suresh et al., 2026). The primary observable follows. Reference three-bond couplings between the amide and α-protons, for single residues flanked by glycine in water, run from 6.1 Hz for alanine to about 7.45 Hz for valine, and within two intrinsically disordered proteins the same residue type varies by 0.7–1.4 Hz on sequence context alone — against an experimental uncertainty that rarely exceeds 0.1 Hz. The spread is ten times the error.
The size of the problem this creates is worth stating. Covering nearest-neighbour effects among all twenty residues on both sides would need 8,000 tripeptide contexts; assuming that upstream and downstream effects simply add would cut that to 400, but additivity fails, so conformational and solvation free energies do not sum along a chain. The largest survey to date covered 361 blocked dipeptides (Schweitzer-Stenner, 2026). Meanwhile every all-atom force field in current use reproduces these ensembles about an order of magnitude worse than the experiment-fitted model, and one did worse than get a number wrong: a spurious 0.2 nm contact between the N-terminal ammonium and a side-chain carboxyl locked 66–78 % of the sampled conformations of an aspartate-rich peptide into a turn, driving the central residue into positive-φ space that the experiment says it does not occupy (Suresh et al., 2026). The simulation was not wrong about the residue. It was wrong about the molecule.
One counterweight keeps this in proportion. Almost all of that local structure disappears from global measurements: denatured and disordered proteins still scale their radius of gyration with chain length as self-avoiding random walks, with the exponent clustering near 0.6 (Schweitzer-Stenner, 2026). A chain can be locally ordered and globally indistinguishable from a random one, which is why the two kinds of measurement must never be merged. The reader should leave this section expecting a backbone to have preferences before it folds, those preferences to be set by the sequence and shifted by the neighbours, and "unstructured region" to name an ensemble with a shape rather than an absence.
03Reading a sequence, and proving it meant something
Everything in the two preceding sections assumes that a peptide has a sequence — one definite order of residues, the same in every copy. Before 1952 that was a hypothesis, and not the favoured one. Proteins were widely suspected of being statistical objects: preparations of consistent composition whose residues were arranged with some regularity but no exact order. Sanger and Tuppy settled it by taking insulin apart into fragments, sequencing the fragments, and reassembling the order of one of its two chains from the overlaps (Sanger & Tuppy, 1952). What that work established was not a fact about insulin. It was a fact about proteins: they have sequences, sequences are definite, and therefore a sequence is the kind of thing that can be read, copied, compared and eventually explained.
The same laboratory then showed that the sequence is not the whole covalent story. Insulin's two chains are held together, and one of them is held to itself, by disulfide bridges between cysteines, and locating those bridges was a separate problem requiring separate chemistry (Sanger et al., 1954). This is the earliest instance of a pattern that recurs throughout this document: a peptide's covalent structure includes connections that a linear reading of the sequence does not show, and a description that omits them is not a description of the molecule. Modern practice records those connections explicitly for exactly this reason.
The demonstration that a sequence means something arrived the following year, in a molecule small enough for the meaning to be visible. Du Vigneaud, Ressler and Trippett determined the amino-acid sequence of oxytocin and proposed a structure for it (du Vigneaud et al., 1953): nine residues, one disulfide, one C-terminal modification. That the same nine residues assembled in glass reproduced the biological activity of the natural material was published in a chemistry journal whose 1953 volumes PubMed does not index, and this document says so rather than attaching an identifier that would imply a verification that did not happen.
Oxytocin is the document's first proof that sequence is meaning, and it needs a partner to make the point. Arginine-vasopressin is also nine residues, also cyclised by a disulfide between the first and sixth, also amidated on the same C-terminal glycine, and identical to oxytocin at seven of nine positions. It differs at position 3, isoleucine against phenylalanine, and at position 8, leucine against arginine. Two substitutions. The two are separate hormones acting at separate receptors, named for separate physiologies — oxytocin for its action on smooth muscle at parturition and in lactation, vasopressin for water retention, which is why it is also called antidiuretic hormone. This document is a structural one and its corpus carries no pharmacology for either; those names are given because they are what the molecules were named for, not as a claim verified here. The structural claim was verified against UniProt entries P01178 and P01185 before the figure below was drawn.
Two substitutions out of nine, on a scaffold that is otherwise identical down to the position of the bridge and the chemistry of the last residue, and the two molecules are different signals. That is the whole thesis of primary structure in one comparison, and it is worth noticing how much the figure carries at once: an ordered sequence, a covalent constraint that is not visible in the sequence, and a one-atom post-translational modification. All three are part of the molecule. None of them is optional.
Knowing a sequence and being able to make one are different problems, and the second took another decade. Merrifield's solution — anchor the first residue to an insoluble bead, add the next in excess, wash away everything that did not react, repeat — converted synthesis from a research project into a procedure that could be run to a schedule. Its founding paper appeared in 1963 in a chemistry journal PubMed does not index that far back, which this document records as a coverage limitation rather than a doubt, citing Merrifield's own later account of the method instead (Merrifield, 1969). Two consequences followed quickly: the method reached the hormones, with an analogue of oxytocin built by solid phase and shown to work (Takashima et al., 1968), and it reached the total synthesis of a chain carrying the enzymatic activity of ribonuclease A (Gutte & Merrifield, 1969).
What that unlocked is the whole practice of structure–activity work, because a chemist could now make the analogue that tests the hypothesis rather than the one that happened to be isolable. Its scale today is measurable: a survey accessed in December 2024 counted 72 synthetic peptide medicines marketed in the United States, European Union and/or Japan, 146 in active clinical development and 24 in phase 3 (Juraszek et al., 2025). Sequencing established that "what is the order?" has an answer; synthesis established that the answer can be changed on purpose, one position at a time. Everything in Parts Four and Five depends on both.
04Learning to see a shape
In 1951, before anyone had seen the three-dimensional structure of any protein, Pauling, Corey and Branson published two hydrogen-bonded helical configurations of the polypeptide chain, derived from the geometry of the amide group and a set of assumptions about where its hydrogen bonds could go (Pauling et al., 1951a). The same series produced the pleated sheet (Pauling & Corey, 1951b) and an application of both configurations to globular proteins (Pauling & Corey, 1951c). The method is worth naming because it is the earliest and cleanest instance of this document's controlling logic running forwards: from the constraints on a bond, to the small set of arrangements those constraints permit, to a prediction about molecules nobody had yet resolved. The prediction was made from the plane described in section 02, and it was right.
Seeing one took another seven years. Kendrew and colleagues obtained a three-dimensional model of myoglobin by X-ray analysis in 1958 (Kendrew et al., 1958), and then did something more consequential than obtaining it: they did it again, better, and attached a number to how much better. The 1960 Fourier synthesis of the same molecule was reported at 2 Å resolution (Kendrew et al., 1960). From that point onwards a structure came with a figure of merit, and the reader of a structural claim acquired an obligation — to ask what that figure was, and what it licenses.
It licenses less than most readers assume, and it is less objective than it looks. On one recent dataset from a therapeutic-class peptide, two standard processing programs applied to the same measured diffraction put the resolution limit in different places: one returned about 1.6 Å from a fitted correlation criterion, the other estimated anywhere from 2.00 Å on a signal-to-noise criterion applied overall to 1.59 Å on a looser criterion applied along one axis. Refining at a series of cutoffs gave better R and R-free values at 2.5 Å than at 1.59 Å, because lower-resolution maps are less noisy — but at 2.5 Å the alternate conformations of one isoleucine side chain vanished from the map. The authors kept the higher cutoff, noting the literature's position that no single statistic reliably fixes where a dataset ends (Mitchell et al., 2026). The number in a deposited structure's header is a decision with a defensible alternative on either side of it.
What better resolution buys is best shown where the same molecule was solved twice. The conotoxin Mu8.1, redetermined at 1.67 Å by X-ray crystallography, showed ordered water molecules that earlier, lower-resolution structures of the identical toxin had not: four refined around a single lysine, which makes saturated hydrogen bonds to two of them and to a threonine hydroxyl, the network continuing through a third water also held by two backbone carbonyls (Müller et al., 2023). That is a chemical explanation for an anomaly, since the same residue is predicted — computationally, not measured — to titrate at 8.16 against 10.5 for a free lysine. Water is what the last half-Ångström buys, and water is very often what explains the chemistry.
At the other end, low resolution constrains what a structure may be asked. In one study of designed antibodies, a 3.0 Å cryo-EM reconstruction let a designed loop be compared to its computational model residue by residue — the loop matched to 0.8 Å root-mean-square deviation — while maps of a different target at 4.6 Å and 5.7 Å permitted only rigid-body docking of the design model, not model building (Bennett et al., 2026). Below that, structures stop being interpretable at all: plate-like crystals of the incretin analogue diffracted to about 10 Å with 30 % completeness, and the unit cell implied roughly 250–500 copies of a 4,920 Da peptide in the asymmetric unit — nothing could be concluded from a dataset that was nonetheless real (Mitchell et al., 2026).
Two further properties of the number are routinely misread. The first is that it is an average over a molecule that is not uniformly ordered. In cryo-EM of the insulin receptor bound to an antagonist peptide, local refinement produced a map at 3.89 Å — worse by the headline figure than the 3.64 Å consensus map — that nonetheless showed more at the site of interest, enough to model an entire B-chain N-terminus rarely visible in such structures because of its flexibility, its five residues joining seven from the A chain to contact thirteen receptor residues across the site-1 interface; in the same study an entire domain on one half of the receptor dimer gave density too low to model at the contour used for the rest of the map and survived every attempt to recover it, which the authors read, correctly, as a positive finding about motion (Vogel et al., 2026). From the other direction, a cryo-EM structure of the GLP-1 receptor with a non-peptide agonist carried a global figure of 3.1–3.2 Å and a local figure of 2.8–3.4 Å at the ligand site, which is what permitted the pocket to be modelled with confidence (Kawai et al., 2020). Quote the local number for the local claim.
The second is that resolution describes only the part of the specimen that is ordered. The incretin analogue diffracted to 1.59 Å, and seven of its C-terminal residues plus the entire fatty-acid conjugate were left out of the model for want of continuous density, with atomic displacement parameters rising from a chain average of 21.8 Ų to 39.2 at one proline and 71.2 at the serine after it (Mitchell et al., 2026). The same lipid chain has gone unresolved in cryo-EM structures of a related acylated peptide bound to two receptors, and solution NMR of related analogues reports the same regions as flexible — three methods converging on an absence, which is evidence about the molecule rather than about any one instrument.
Then there is the lattice itself. Crystallisation does not merely observe a conformation; it can select one. The clearest demonstration here is a β-lactamase crystallised twice: a new trigonal form at 1.8 Å against a previously reported orthorhombic form of the same unliganded enzyme. Overall the two agreed to 0.49 Å on α-carbons, but three loops flanking the substrate cleft deviated up to fourfold more, displacing two phenylalanine rings by 2–3 Å and narrowing the cleft, while a fourth loop held by an internal hydrogen-bond network superimposed to 0.25 Å with no change at all (Saino et al., 2015). Two packings, one protein, two answers about the active site — and, used deliberately, a way of finding out which parts of a molecule are mobile. The same effect has been argued to hide biology: the atomic differences distinguishing two oncogenic mutations at one position of K-Ras4B were reportedly not captured in crystal structures, where crystal contacts favour an alternative stabilisation of one switch region, and were seen instead by NMR and molecular dynamics. The same review supplies the most self-evident case of a real structure that cannot be a working one — and supplies it, fittingly for this section, from models rather than from measurement: in an in-silico inactive state derived from the PI3Kα crystal with its lipid substrate soaked in, the inositol head sits 6–7 Å from the phosphate it must attack, against 1–3 Å in the authors' own modelled active state (Nussinov et al., 2023). Both distances are model outputs, not an observation and a chemical requirement.
It is worth resisting the conclusion this invites. Packing does not always force a conformation; sometimes it accommodates one. In the incretin-analogue lattice, the helical cores stack into square channels with pore cross-sections of roughly 30 × 30 Å, and it is into those channels that the disordered C-terminal residues and the fatty-acid conjugate go — they could be fitted there without atomic clashes, and were deliberately left unmodelled rather than forced into a pose that would have improved the statistics (Mitchell et al., 2026). The lattice hosted the disorder instead of removing it, and the authors offer a reason peptides should differ from proteins here: a protein can bury a flexible or greasy region inside its own fold, whereas a short peptide is solvent-exposed along its whole length, so the lattice has to do the burying.
Every other method trades the same currencies in different proportions, and the honest way to compare them is by what each measures, what it demands, and what it costs.
Read across those rows and a pattern appears that matters more than any individual limit: several of these methods answer questions that look alike and are not. Small-angle scattering cannot see an amino acid, yet fixes a radius of gyration far more precisely than its own 10 Å ceiling, because it is a low-information technique from which a small number of parameters can be extracted very well (Martin et al., 2021). Hydrogen–deuterium exchange needs neither crystals nor high concentration and reports on flexible regions the others cannot resolve at all, yet it is bounded to the peptide rather than the residue and averages over co-existing forms so completely that distinct fibril polymorphs must be physically separated before they can be compared (Meng et al., 2025). Solution NMR, whose reputation rests on large slowly tumbling proteins, has its weakest signal at exactly the size of a macrocyclic peptide drug, and the fix is to slow the molecule artificially with cold viscous solvent (Rüdisser et al., 2023).
That last point generalises into the section's closing claim. Those same authors state that their cyclic peptide's structure in the crystal and in lipophilic solvent differ considerably from its structure in water, and that it is the aqueous conformation that resembles the one it adopts bound to its protein target. A reader given only the crystal structure would have the wrong molecule for the purpose of understanding how it binds — not a false structure, a real one answering a question that was not asked. Resolution decides what a structure can be used for; the method and the conditions decide what question it answers at all.
05Anfinsen's claim, and Levinthal's objection
By the early 1970s two things were established: that a protein has one definite sequence, and that it has a definite three-dimensional structure that can be determined. Anfinsen's contribution was to argue that the second follows from the first — that the information determining a chain's folded structure is contained in its amino-acid sequence, and that the folded state is the one of lowest free energy accessible to the chain under its physiological conditions (Anfinsen, 1973). The thermodynamic hypothesis is the reason the word "sequence" carries the weight it does in every field downstream of protein chemistry. If it is true, then a sequence is a complete specification, and every question about a protein's shape is in principle a question about its sequence plus its solvent.
Levinthal's objection is the counterweight, and it is an objection about time rather than about energy. If the folded state is simply the global minimum, then a chain looking for it by trying conformations would have to search a space that grows multiplicatively with every residue and every rotatable bond. Even at implausibly fast rates of interconversion, the time required to sample that space exhaustively exceeds by an enormous margin the time real chains actually take, which is often less than a second. Something must therefore direct the search. This document states the objection in words and prints no arithmetic for it, because the evidence base assembled here does not carry Levinthal's original calculation — no conformations-per-residue figure, no total search time, no comparison against measured folding times. The numbers are famous, they are reconstructible, and they are not in this reading; they are therefore not invented here.
What the corpus does carry is the modern formal version of the same point. Even a drastically simplified representation of folding — a chain of hydrophobic and polar beads on a lattice — is NP-hard, so the cost of finding the global minimum energy conformation grows exponentially with chain length over a landscape described as rugged (Uttarkar et al., 2026). The practical face of that exponent appears in the same paper's benchmark: a quantum-classical folding pipeline run on 1,224 sequences reached a negative-energy fold for 99.6 % of six-residue peptides and 97.3 % of seven-residue ones, then collapsed to 13.4 % at eight residues, 4.7 % at nine and 1.5 % at ten. The figure benchmarks a search procedure rather than physics, since the implementation used randomised contact potentials and fixed idealised backbone geometry. Read as such, it is still the shape of Levinthal's complaint, in a system where the search is done deliberately and can be counted.
The resolution of the two positions is the one the rest of this document depends on. Both are right, and they are reconciled by giving up the picture of folding as a search over equivalent alternatives. A chain does not sample conformations at random, because its landscape is not flat: it is biased from the first moment by everything sections 01 to 03 described — the frozen amide, the sterically forbidden regions of the map, the sequence- and neighbour-dependent preference of every residue for one basin over another. Folding runs downhill on a surface that is already tilted, and the tilt is what the sequence encodes. Drawn as a landscape, this is the funnel: a broad rim of many high-energy conformations narrowing towards a few low-energy ones.
The funnel is a picture rather than a measurement, and this document treats it as one: the evidence base carries no roughness parameter for a landscape, no relation between folding rate and topology, and no measured downhill folding case. What the corpus does supply is the same logic in population terms — on the energy-landscape account, a protein under physiological conditions largely populates its inactive conformations, with the active state present only as a minor population, and effectors, membranes and mutations work by shifting those populations rather than by pushing a molecule into a new shape (Nussinov et al., 2023). On that view, single structures are snapshots that cannot by themselves explain how function is executed.
Both side valleys are real, and both have been caught in the act. The kinetic trap first. Pulling the 92-residue designed protein Top7 apart in optical tweezers at constant speed, an intermediate state appeared in only about 7 % of unfolding events; holding the same molecule at a fixed force near 15 pN, it appeared in 95 % of them — 121 events out of 128 (Li et al., 2020). The intermediate was always there; at high force it simply lived too briefly for the instrument to see. What it is was settled by elimination and a negative control: it lengthens the chain by 3.9 nm, only the two helices among the candidate substructures could give that length, and the isolated fragment containing one of them unfolded cleanly in two states with no such event. The intermediate was therefore assigned to interactions that do not exist in the folded protein at all. A trap need not be a partly built native structure.
The aggregation exit next. It is a competing outcome rather than a pathological one, and for short peptides concentration often decides. For a cyclic antimicrobial heptapeptide in water, water-exchange cross-peaks in solution NMR bracketed the onset of strong self-association between 0.15 and 0.30 mM, above which every backbone amide was shielded from solvent and below which all were exposed, with amide line widths above 20 Hz indicating high-order oligomers for a molecule near 1,300 Da (Lohan et al., 2025). One sequence, one solvent, two concentrations, two sides of a boundary. And the endpoint of that exit is itself an ensemble rather than a state: hydrogen–deuterium exchange on two separated fibril polymorphs of one 140-residue protein, at 100 % sequence coverage across 28 peptides, localised the difference between them to an eleven-residue stretch (Meng et al., 2025).
One last observation completes the picture, from the same designed protein. Top7 was the first computationally designed globular protein whose experimental structure matched its design at atomic accuracy. Its folding is nonetheless non-cooperative, passes through intermediates, involves non-native interactions and produces fragments stable enough to fold and dimerise alone — where natural globular proteins of that size generally fold in a single cooperative step and their fragments are seldom stable in isolation. The interpretation offered is that the smoothness of a natural folding landscape is itself a product of selection, which a designed sequence never underwent (Li et al., 2020). Getting the structure right and getting the route right are separable achievements.
All of which converges on the sentence this document is built to earn, stated here once and then tested against the biology in every Part that follows. A sequence is not a blueprint. A sequence is a set of constraints on a population of shapes. The constraints are the ones assembled in this Part: one planar bond that removes a third of the chain's freedom, a sterically permitted region that removes most of what is left, per-residue and per-neighbour preferences within that region that bias the rest, and covalent bridges and terminal modifications that pin particular points. What those constraints produce is not a structure but a distribution over structures — one that shifts when the solvent shifts, that a crystal can select from rather than reveal, and that most of the instruments tabulated above report only in average. Part Two takes that distribution seriously, and begins where it must: with the fact that a short peptide in water usually has no single shape at all.
06The four shapes that keep reappearing
A 2026 survey swept more than 230,000 Protein Data Bank entries asking one narrow question: how many contain a run of residues sitting inside the polyproline II window of the Ramachandran map? More than 130,000 contain at least one such run three residues long. Over 4,000 contain two consecutive turns, 238 contain three, and 22 — twenty-two, out of a quarter of a million — contain four or more (López-Sánchez et al., 2026). One turn of the shape is close to universal. Four turns barely exist.
That asymmetry is the shape of the subject. Local structure is not a menu of architectures a chain picks from; it is a small set of hydrogen-bond registers that recur constantly over three or four residues and grow rarer the longer they must persist. The four that recur are best defined by their registers rather than by their pictures, because the register is what an experiment reports.
There are few of them because of a counting problem. Every residue contributes one hydrogen-bond donor, the amide N–H, and one acceptor, the carbonyl C=O, and in water both are satisfied by water. Any conformation that pulls the backbone away from solvent must satisfy them another way, using only the φ/ψ combinations Part One showed to be sterically permitted. Very few periodic arrangements manage both.
The α-helix is defined by an i → i+4 register: the amide N–H of residue i bonds to the carbonyl oxygen of residue i+4, and the chain advances 3.6 residues per turn — which is why older work writes it as the 3.613 helix. Ten residues of the α1 helix of human ACE2 sit in exactly this register across four structures of the receptor bound to the SARS-CoV-2 spike: X-ray at 2.45, 2.68 and 2.50 Å and cryo-EM at 2.90 Å (Ferková et al., 2023). A 1.55 Å crystal structure of the rice protein RLF describes the same register atom by atom rather than asserting it, in an induced eleven-residue helix bonded from the Ser116 α-carbonyl to the Trp120 α-amide (Benson et al., 2024).
Shift the register by one residue and you have the 310 helix, bonding i → i+3, and the two are hard to separate. Their backbone coupling constants are near-identical — 3.9 Hz for the α-helix against 4.2 Hz for 310 — so solution NMR distinguishes them only by a distance: the dαN(i,i+2) contact is 4.4 Å in an α-helix and 3.8 Å in a 310 helix, close enough to give a cross-peak only in the latter. A lactam-bridged nine-residue peptide was assigned α-helical on exactly that basis, by NMR at 600 MHz in aqueous phosphate at pH 6.8 (Ferková et al., 2023). Circular dichroism uses a band ratio instead, R = [θ]222/[θ]208, near 1 for an α-helix and below 0.4 for 310; analogues of the peptaibol trichogin GA IV gave R ≈ 0.7 in 100 mM SDS micelles, read as a mixed 310/α helix and confirmed by NMR in the same micelles showing both diagnostic NOE families (De Zotti et al., 2020).
The β-sheet is the odd one out, because its register leaves the segment. An extended strand makes no hydrogen bonds along its own length; its donors and acceptors point sideways and are satisfied only when another strand lies alongside. A β-strand is half of an interaction, and which half it pairs with is not set by the strand. In a converged simulation of the Aβ(1–42) monomer in explicit water, residues 38–40 appear as a strand in many conformers but pair with residues 33–35 in one and 18–20 in another, and the choice sets the global shape (Sgourakis et al., 2011). One limit belongs here: none of the sources behind Part Two reports the inter-strand spacing or the pleat geometry of a folded β-sheet. The only strand-to-strand distances they measure were measured on amyloid, in Section 09.
The turn is the register that reverses the chain, most often i → i+3 across a four-residue reversal. Turns have explicit φ/ψ boxes rather than one geometry: the convention used in current work on unfolded peptides names a type I/II′ β-turn(i+2) region, an inverse γ-turn region and an asx-turn region, each with numerical bounds (Suresh et al., 2026). The asx turn is the interesting case, because the backbone does not make it — protonated aspartic acid in water is dominated by asx conformations, apparently through a hydrogen bond from its own side-chain carboxylic acid back to the main chain.
Polyproline II solves the hydrogen-bonding problem by refusing it. It is left-handed, extended, completes exactly three residues per turn, and has no internal hydrogen bond at all: donors and acceptors point at the solvent and water satisfies them. It is defined purely by a Ramachandran window — and by whose. The survey above used φ from −95° to −55° and ψ from +125° to +175°, deliberately narrower than other published definitions (López-Sánchez et al., 2026); the mesostate convention used for spectroscopy on short peptides uses φ from −90° to −42° and ψ from 100° to 180° (Suresh et al., 2026). Any percentage of polyproline II is a statement about which box was drawn. This is the conformation older work filed under "random coil".
Why these four is answered at the level of two residues. Backbone preferences in water are side-chain-specific and neighbour-dependent: alanine sits 58% in polyproline II flanked by glycines and above 75% flanked by alanines, valine is β-strand-dominated in GVG and more so in VVV, and like neighbours amplify a residue's intrinsic preference while unlike neighbours damp it (Suresh et al., 2026). The energies are not small — serine and valine destabilise a neighbouring alanine's polyproline II conformation by roughly 4 and 3 kJ/mol (Schweitzer-Stenner, 2026). Helix and sheet nucleation is this effect running cooperatively along a few residues.
The 3.6-residue period has one further consequence: it gives a linear sequence sides. At 3.6 residues per turn each residue sits 100° round from its predecessor, so plotting the sequence on a circle at that spacing — the helical wheel — puts residues far apart in the sequence next to each other in space. A peptide whose polar and non-polar residues alternate near that period winds up with a polar face and a non-polar one, and that property, amphipathicity, belongs to the order of the residues and is invisible in the composition.
The design consequence is measurable. A helical-wheel projection of the antimicrobial peptide temporin-SHf showed only two polar residues, Ser5 and Arg6, on its hydrophilic face. Substituting Ser5 with arginine and appending further arginines dropped the minimum inhibitory concentration against methicillin-resistant Staphylococcus aureus from 50–100 µM to 1.57 µM; reversing the topological order of the sequence, which segregates charged from hydrophobic residues and raises the hydrophobic moment, extended activity to Gram-negative strains from above 100 µM to 6.25 µM (Tang et al., 2023). One limit: hydrophobic moment is invoked qualitatively throughout these sources and never reported as a number with its units. The nearest quantitative descriptor is its one-dimensional cousin, the charge-patterning parameter κ, 0.616 for melittin against 0.179 for histatin 1 (Svensson et al., 2025).
A named secondary structure should now read as a claim of a particular kind: a hydrogen-bond register, holding over a stated number of residues, measured by a named method. It should also read as incomplete. Nothing so far says whether a given peptide in a given tube is in any of these states — and for peptides of the length used as drugs, usually it is not.
07Short peptides mostly have no shape at all
Two nine-residue peptides were cut from the α1 helix of human ACE2 — the helix Section 06 described in a continuous i → i+4 register in four structures of the receptor bound to the SARS-CoV-2 spike. Measured on their own by circular dichroism in phosphate-buffered saline at pH 7.4 and 25 °C, both gave the random-coil signature: a strong negative band near 200 nm and only a small band at 222 nm. Adding trifluoroethanol to 50% raised their helical content to 36.9% and 27.3% (Ferková et al., 2023). Same atoms, same order, same temperature. The helix appeared when the solvent changed.
This is the correction the rest of Part Two depends on. A ten- to thirty-residue peptide in water is usually not folded, and every conformation reported for one is a joint property of the peptide and its surroundings. A helicity of 37% in 50% trifluoroethanol is a true statement about a peptide in trifluoroethanol, not a property of the peptide.
It helps to see what the peptide lost. The ten-residue sequence 10Panx1 forms helix H1 of the first extracellular loop of pannexin-1 in the cryo-EM structure of the heptameric channel; on its own, in phosphate-buffered saline at 100 µM, it gives a random-coil CD spectrum. What held the helix up was context the peptide cannot carry: an N-cap hydrogen bond from the Ser73 hydroxyl to the Gln76 amide, two disulfides to residues far away in the sequence, and an antiparallel three-strand β-sheet packed against it (Lamouroux et al., 2023).
The reverse experiment — hold the peptide fixed and change the surroundings — is where the point becomes vivid. HT2, a ten-residue analogue of the frog antimicrobial peptide temporin-SHf, shows no helical CD signature in water and becomes α-helical in 80 mM SDS micelles or 50% trifluoroethanol. Solution NMR at 700 MHz in deuterated SDS places that helix at residues 3–9. Atomistic molecular dynamics in a 3:1 POPE/POPG bilayer places it at residues 2–7 or 2–8 instead, which the authors attribute to the difference between a curved, negatively charged micelle and a flat mixed bilayer (Tang et al., 2023). Two membrane mimics, one peptide, two answers about which residues are helical. The NMR structure rests on twenty NOEs, none of them long-range — a propensity readout rather than a determined fold, as the paper says.
Nor is it "membrane" that does the folding. Two twenty-residue peptides from the disordered C-terminal tail of the Ku protein were measured by CD in phosphate-buffered saline at pH 7.4 and then in two supported membranes at a peptide-to-lipid ratio of 1:50. Both were random coils in buffer and stayed random coils in a zwitterionic DPPC/cholesterol membrane. In an anionic DPPC/DPPG/cholesterol membrane, the alanine-repeat peptide gained the helical double minimum at 222 and 210 nm, while its proline-repeat sibling, differing only by proline for alanine in a repeating motif, stayed disordered in both (Maity et al., 2023). The variable is the charge of the bilayer, not its presence, and proline vetoes the transition. The peptides that folded were also the only ones in the series to cost cell viability.
Concentration is a third variable, reported even less often. Small-angle X-ray scattering on phosphorylated histatin 1 gives a radius of gyration of 15.8 Å at 0.21 mg/mL and 27.2 Å at 4.77 mg/mL, with rising low-angle intensity indicating self-association (Svensson et al., 2025). The counterweight matters as much: the proline-rich domain of p53 gave CD spectra invariant across 10 and 150 mM sodium fluoride and from 0.2 mg/mL to nearly 9 mg/mL (Berggren et al., 2026). The difference tracks amphipathicity — the property the wheel made visible.
The instrument is the next problem. Circular dichroism is the standard tool for peptide secondary structure, and it is structurally blind to the conformation that dominates short peptides in water. Common deconvolution methods carry no polyproline II basis set, so polyproline II is silently redistributed among the helix, turn and unordered categories; detecting it requires a temperature scan and watching the loss of a maximum near 228 nm, which for the p53 proline-rich domain was linear and reversible on cooling (Berggren et al., 2026). The failure runs the other way too: a strong negative band at 217–218 nm read as helix by a single-wavelength formula was reassigned by full deconvolution to right-twisted antiparallel β-sheet, splitting one peptide into 28.9% helix plus 23.2% sheet (Ferková et al., 2023). In the same study, CD and NMR ranked two stapled peptides in opposite orders.
None of this means an unfolded peptide has no structure. It means it has no single structure. Replica-exchange molecular dynamics on the Aβ(1–42) monomer in explicit water — 52 replicas spanning 270.0 to 601.2 K, 11.7 µs in aggregate — produced an 11,564-conformation ensemble whose representative members include an α-helix at residues 8–12, a 310 helix at 29–33, several distinct β-hairpins and a large coil population, with no dominant structure (Sgourakis et al., 2011). The same paper is honest about the fit: correlation to experimental couplings plateaued around 0.4–0.5, and two independent experimental datasets on the same peptide correlate with each other at only 0.92.
At the other end of the range sits the crystal, and it is not the exception the reader expects. A 42-residue acylated GLP-1/GIP analogue was solved at 1.59 Å. The α-helix runs from Gly4 to Glu28; everything else is randomly coiled; residues Ser33 to Ser39 and the entire fatty-acid conjugate had no continuous electron density and could not be modelled, and isotropic temperature factors rise from a chain average of 21.8 Å2 to 71.2 Å2 at Ser32. The hydrophobic face of that helix is buried not by the peptide's own fold but by its neighbours, which assemble into square motifs producing channels roughly 30 × 30 Å that the disordered residues and the lipid occupy (Mitchell et al., 2026). A crystal structure of a short peptide is a statement about the peptide and about the lattice, in unknown proportion.
The sources disagree about how to read that, and the disagreement is worth stating rather than resolving. One reading is that the environment sets the structure: the same peptide puts its helix on different residues in a micelle and in a bilayer, and has none in water (Tang et al., 2023). The other is that the environment selects from a population the peptide already has. The case for the second is a 2.28 Å crystal structure of an α-methylated ACTR variant bound to NCBD, which buries 1066 Å2 where the solution NMR structure of the same complex buries 1655 Å2, and where a second NMR structure gives the C-terminal helix a third orientation — with no crystal-packing contacts found to explain the difference (Bauer et al., 2020). Induced or selected, both agree on what matters here: no single structure represents a short peptide.
A conformational claim should now read as incomplete unless it names the solvent, the temperature, the concentration and the method, and a percentage of helix without those as a number of unknown provenance. The next section takes the further step: for a large class of sequences, having no fixed structure is not a deficiency but the state in which they work.
08Disorder is a state, not a failure
ACTR is a 71-residue activation domain, disordered in solution, that folds into three helices when it binds the nuclear coactivator-binding domain of CBP. An all-atom ensemble of the free chain, generated by enhanced sampling and reweighted against NMR secondary chemical shifts and paramagnetic relaxation enhancements, decomposes into substates. One accounts for about 2.6% of the population and carries roughly 85% of the backbone contacts ACTR makes in the bound complex, with two of its three binding helices fully formed (Streit et al., 2026). The bound shape is not created by the partner. It is already there, rarely.
The same paper delivers the opposite result in the same breath. Ten unbiased 5 µs simulations started from the experimental bound coordinates with the partner deleted show the tertiary arrangement failing at once: inter-helix orientation and radius of gyration fluctuate widely and the chain collapses into compact states that do not exist in the complex. The helices stay metastable; their arrangement does not. Most tertiary contacts in the complex are ACTR-to-partner. The partner is part of the structure.
Those two facts together are what "intrinsically disordered" means. Such a protein, or a disordered region inside a folded one, has no single dominant structure under physiological conditions — not because folding failed but because the sequence does not encode one. The composition says so: disordered regions are enriched in proline, serine, glutamine, lysine and glycine and depleted in leucine, isoleucine and phenylalanine, which is the physical reason no hydrophobic core forms (Harake et al., 2026).
How much of a proteome is in this state has to be answered carefully, because these sources give at least six different numbers that are not measuring the same thing. Four denominators are in play — proteins, protein-coding genes, residues and proteome — and most figures state neither a length threshold nor a predictor. Neither of the two quoted here states both. Extended disordered regions of thirty residues or more in 92% of human and 91% of Arabidopsis thaliana transcription factors (Salladini et al., 2020) carry the threshold but no predictor: that figure is quoted from an earlier survey, and the IUPred2A, PONDR-FIT and MoRFpred runs the paper does report cover a single protein. Disordered regions in nearly 30% of proteins (Gnanaolivu and Hart, 2026) are borrowed the same way; the AlphaFold2 confidence below 70 and relative solvent accessibility above 0.581 in that paper define disorder for its own variant analysis, not for the 30%. "About a third of the proteome is disordered" is not a single measured fact, and this document does not repeat it as one.
The cleanest demonstration that disorder is a functional state is a set of three domains that are nearly the same sequence and behave nothing alike. The C-terminal transactivation domains of human CITED1, CITED2 and CITED4 are each about fifty residues and 45% identical, with a further 13% conservative substitution. All three give 1H-15N HSQC spectra with amide protons clustered near 8.1 ppm, the disordered signature, and all three gain helix on mixing with the TAZ1 domain of CBP. But only CITED2 carries residual helix free in solution, shown by the α-helical double minimum at 208 and 222 nm by CD and by positive 13Cα secondary shifts across residues 225–235. Surface plasmon resonance at 25 °C gives dissociation constants of 7 ± 2, 23 ± 6 and 220 ± 40 nM for CITED2, CITED1 and CITED4, with CITED1 and CITED2 sharing an association rate, so the difference between them lives entirely in the off-rate (Do et al., 2026). The helix predictor got it wrong: AGADIR returned about 7% helicity for CITED2's helical region and about 7% for CITED4's, which both CD and NMR contradict. A near-identical sequence does not specify a near-identical population of shapes.
Binding has two named limiting cases with different kinetic signatures. In conformational selection the partner captures a conformation the free chain already visits, so the population of that conformation is rate-limiting; in induced fit the chain meets the partner first and orders afterwards, so the rearrangement appears as a separate step whose rate does not depend on how much partner is present. These sources disagree about which dominates, and neither side argues from primary kinetics: one review holds that coupled folding and binding proceeds predominantly by induced fit because free-state folding is too slow (Harake et al., 2026); another holds that binding is selection from the pre-existing ensemble followed by minor optimisation, and calls the alternative incompatible with a physics-based outlook (Nussinov et al., 2023).
The measured systems support neither cleanly. For the NCBD/CID pair, alanine-to-glycine mutations that modulate helix propensity give Brønsted slopes of 0.45 ± 0.04 for CID helix 1 — roughly 45% native helical content already present at the rate-limiting barrier — against a single slope of −0.03 ± 0.02 for helices 2 and 3 treated together, which are effectively unformed. Helix 1 is also the segment transiently populated in the free state. That is selection. Yet mutating the buried salt bridge between NCBD Arg2104 and CID Asp1068 exposes a second stopped-flow phase at 15–20 s−1 that does not change across the whole partner-concentration range — a unimolecular rearrangement after complex formation, the induced-fit signature (Karlsson et al., 2020). The disordered POSH fragment binding the GTPase Rac1b resolves the same way by 15N chemical exchange saturation transfer, into an anchoring step and then a unimolecular folding step at 141 and 49 s−1 (Kjaer et al., 2026). All-atom simulation now suggests the two are stages rather than alternatives: the chain binds from a partially helical sub-ensemble and rearranges inside the complex, and chain length decides how painful that is (Ghosh et al., 2024). One absence: no experiment here reports the classic selection diagnostic, an observed rate that decreases with partner concentration.
Evolution has been run on one of these interactions, and the result is not the tidy one. A roughly 500-million-year-old version of the NCBD/CID pair was reconstructed and its binding transition state mapped residue by residue. The ancient complex has the more ordered transition state — interface φ-values mostly 0.3–0.6 against 0–0.3 for the modern human one — and restrained molecular dynamics confirms that the ancestral transition-state ensemble is more compact and less heterogeneous. The modern complex is about seven times tighter (Kd 0.11 ± 0.01 against 0.81 ± 0.07 µM), and the gain sits almost entirely in dissociation (Karlsson et al., 2020). Evolution made this interaction fuzzier while making it stronger.
The thermodynamic accounting is equally uncooperative. Calorimetry on CBP-NCBD binding ACTR gives an enthalpy of −132.6 kJ/mol against an entropic penalty −TΔS of +89.1 kJ/mol at a dissociation constant of 34 nM — enthalpy-driven association paying heavily for ordering the chain. But complexes of the same hub family run the other way: the plant hub RCD1-RST binds three partners with −TΔS of −27.2, −36.3 and −16.9 kJ/mol, entropy-driven. And high affinity does not require folding at all: the same hub binds one disordered partner at 9 nM with no helix induction detected (Bugge et al., 2021). The review reporting these values cautions against cross-comparison without heat-capacity data; that caution is part of the finding.
Two consequences follow for anyone trying to act on a disordered target. The first is that you cannot aim at a site, because there is not one. Eight conformations of the disordered p53 transactivation subdomain were sampled from simulation, screened for transient pockets and used to dock roughly 20,000 compounds; only molecules ranking in the top 1% against at least three different conformations were pursued. Of 244 tested, ten bound by surface plasmon resonance, the best of them at 13.8 ± 4.8 µM, and a later round of seventy analogues built from that hit reached 8.9 ± 3.7 µM — while ten compounds selected for binding a single conformation very well showed no detectable binding at all (Ruan et al., 2020). The second is that the recognition information is not where the textbook puts it. Screening over 800 human domains against a million-peptide library tiling the human disordered proteome mapped 20,009 motif–domain interactions; of 103 of those pairs retested one by one, 96 bound measurably, at a median dissociation constant of 30 µM; of 222 sequences predicted to bind purely by consensus-motif match, only 44 did (Madhu et al., 2026, preprint).
One caution transfers directly to peptide work. Seven segments of the disordered carboxy tail of neurofilament light were measured both as isolated synthetic peptides and in place within the full-length protein, by small-angle X-ray scattering and time-resolved Förster resonance energy transfer. Every one was more expanded in context than in isolation — the source states that direction for every segment under every condition, but gives the per-segment magnitudes only in its tables and supplementary tables, never as a summary figure — and replacing the native flanking sequence with inert proline-alanine-serine repeats abolished the effect, so it is the specific neighbouring sequence and not tethering that does the work (Koren et al., 2023). A peptide cut out of a precursor is a different molecule with its own ensemble.
Disorder should now read as a measurable state with its own parameters — a population with a width, a composition bias, a partner-dependent narrowing, and an entropy bill that can fall either side of zero — rather than as an experimental failure. The last section of Part Two follows the same ensemble to the exit nobody designs for.
09The same physics, run the wrong way
Teriparatide is parathyroid hormone residues 1–34, approved for osteoporosis and manufactured at scale. Left in 50 mM sodium phosphate with 150 mM NaCl at pH 7.4 and 37 °C until it fibrillates, then pelleted and dried, it gives a wide-angle X-ray scattering pattern with a meridional reflection at 4.75 Å and an equatorial one at 10.8 Å (Sachan et al., 2026). That pair of numbers is the cross-β signature — the architecture of the disease amyloids, produced by a marketed drug in a buffer a formulation scientist would recognise.
Cross-β is a simple arrangement described awkwardly. The chain runs in extended β-strands lying roughly perpendicular to the fibril axis; hydrogen bonds run along that axis, stitching each strand to the ones above and below into a sheet thousands of layers deep; and sheets stack face to face across the axis. The meridional reflection is the strand-to-strand repeat along the fibre, the equatorial reflection the sheet-to-sheet spacing across it.
Three independent methods here put numbers on them, and the numbers differ in a way worth reading correctly. Wide-angle scattering on dried teriparatide fibril pellets gives 4.75 and 10.8 Å (Sachan et al., 2026). Oriented fibre diffraction on dried aligned fibres of a YB-1 protein fragment gives about 4.7 and about 10 Å (Timchenko et al., 2026). A 3.6 Å cryo-EM helical reconstruction of wild-type amylin fibrils reports a layer separation of about 4.9 Å within each protofilament, with the layers tilted about 10° from perpendicular (Gallardo et al., 2020). These are not competing values: diffraction measures a lattice repeat averaged over a dried, aligned bulk, while cryo-EM measures a modelled layer separation in one reconstruction of one polymorph whose layers are not perpendicular to begin with. Name the method in the sentence that reports the structure.
A peptide monograph has to care because this architecture is not the property of a few unlucky sequences. Running the predictor TANGO over 20,162 reviewed human protein sequences under a purely geometric criterion returned 26,715 candidate amyloidogenic hairpin regions across 10,711 proteins — 53% of the human proteome, averaging 2.5 regions per protein (Heid et al., 2024). Prediction on that scale invites scepticism, so the same group tested it, though not on that list: they filtered it to 2,505 candidate regions in 2,098 proteins, then chose eight segments from that shorter set — a selection, not a random draw, and one that deliberately included two proteins already annotated as amyloid. Seven aggregated, four giving fibrils by atomic force microscopy and three binding an engineered β-hairpin-binding protein at 4.7 ± 0.9, 4.7 ± 0.5 and 39 ± 5 µM by calorimetry. Aggregation is a default destination for polypeptide chains; folding, flanking disorder and cellular housekeeping are what normally prevent it.
If aggregation is generic, what does the sequence contribute? The fullest answer available is a deep mutational scan of islet amyloid polypeptide — amylin, the 37-residue hormone from which the drug pramlintide derives. 1,916 variants were measured in one selection assay: 64.6% of mutations reduced nucleation, 15.3% behaved like wild type and 20.1% accelerated it, giving 516 gain-of-function variants where at most four amylin variants had previously been compared in parallel (Badia et al., 2026). The loss-of-function map is what a structural biologist would draw: an unbroken sensitive stretch from residue 15 to 32, with per-position sensitivity tracking burial across eight published fibril structures.
The gain-of-function map is the surprise, and it is the part that matters for anyone engineering a peptide. Sixty-one accelerating substitutions sit in residues 1–12, a region not resolved in most amylin fibril structures. In residues 11–21, 131 of 228 insertions sped nucleation up, and deleting the whole 10–14 stretch nucleated faster than wild type — evidence for a gatekeeping element that stabilises the soluble monomer rather than the fibril. A mutation can make a peptide aggregate faster from a position that touches nothing in the final structure.
That is why prediction sits where it does, and these sources disagree about how well it works. A protein-language-model classifier reached receiver-operating-characteristic areas of 0.91–0.93 on a 222-hexapeptide benchmark, ahead of PASTA at 0.89, CANYA at 0.83, TANGO at 0.71 and WALTZ at 0.61 (Lobo et al., 2026). The deep mutational scan reports the opposite: Zyggregator, TANGO, CamSol, s4pred, AlphaMissense and PopEVE all performed poorly at predicting measured nucleation scores, and gain-of-function mutations in particular were unpredictable (Badia et al., 2026). These are different tasks, and together they mark the honest boundary: predictors find aggregation-prone segments and fail at predicting what a change to a sequence will do. Across nine immunoglobulin light chains, the flagged segments were mostly buried and highly protected in the native fold, so propensity says nothing about whether a segment is ever exposed (Peterle et al., 2026).
The structural consequence of a single substitution can be total. The S20G variant of amylin sits on the solvent-exposed surface of the wild-type fibril and participates in no stabilising contact in any solved structure, yet its fibrils cannot be superposed on the wild-type fibril at all: only one six-residue segment matches, the protofilament interface shifts by two residues, and a new three-protofilament architecture appears in which two different folds of the same sequence coexist in one fibril (Gallardo et al., 2020). Even the handedness of the wild-type fibril is contested, and the dispute comes down to sample handling: left-handed by microscopy of hydrated fibrils, right-handed by microscopy of dried ones, with the monomer folds in agreement.
The core is stabilised by a repetitive hydrogen-bond ladder along the fibril axis and laterally by two named interface classes — dry, interdigitated apolar contacts called steric zippers, and arrays of polar groups called polar zippers (Taylor et al., 2026). The amylin structure carries, in one reconstruction, essentially every motif catalogued in the amyloid literature: a homotypic dry steric zipper at residues 23–25, asparagine ladders stacking along the axis at three positions, aromatic stacking of two phenylalanines and a tyrosine, a polar tube of side chains running the fibril's length, and parallel in-register sheets (Gallardo et al., 2020). One limit: these sources refer to steric zippers repeatedly without enumerating the eight symmetry classes of the classical scheme, and this document therefore does not either.
The kinetics explain why a small problem becomes a total one. For Aβ42 measured by thioflavin-T fluorescence, quiescent, at pH 8.0 and 37 °C, no dataset could be fitted by any model lacking secondary nucleation — new aggregates forming catalytically on the surface of existing fibrils — and the mass generated by that route exceeded the primary route by at least an order of magnitude at 3 µM monomer. The surface catalysis is nonetheless structurally specific: a single valine-to-serine change lengthened the fibril twist and abolished cross-seeding in both directions (Thacker et al., 2020). Secondary nucleation is simultaneously a generic physical process and a faithful template, which is why a trace of the wrong species in a vessel is a different problem from a trace of dust.
That brings the physics to where it costs money. Native mass spectrometry of liraglutide at 1 mg/mL in 20 mM ammonium acetate at pH 6.7 and 37 °C, quiescent, detected oligomers of two to eight monomers within ten minutes and thirteen to sixteen by four and a half hours. Single-ion charge detection then resolved species of roughly 100–250 kDa corresponding to 25 to 62 monomers, in four discrete clusters, where the prior consensus from scattering and chromatographic methods had been six to fourteen. Raising the pH from 6.7 to 8.1 halved the mid-range oligomers within ten minutes and made the highest species undetectable (Kuo et al., 2025). The lipid conjugation that gives liraglutide its duration of action by promoting self-association is the same feature that produces the aggregates.
A solubility limit can determine a product's entire presentation. Pramlintide precipitates above pH 5.5 and therefore cannot be co-formulated with insulin; it is supplied and administered as a separate product (Bousch et al., 2026). Its parent hormone cannot be used as a drug at all, because it aggregates both in solution and at the injection site. And "non-fibrillating", the property pramlintide was designed for, is a kinetic claim over a stated window: atomic force microscopy at 0.5 mM in water at pH 7.4, quiescent, showed only spherical oligomers of 85 ± 17 nm at 24 hours — and fibrils at seven days, or at 24 hours with equimolar zinc present (Dudek et al., 2022). Other work holds that zinc inhibits amylin fibrillisation; the same paper concludes the effect is concentration-dependent and can go either way. The sign of a metal ion's effect is not a fixed property.
The intuitions are unreliable in both directions. Parathyroid hormone's C-terminal disordered region contributes nothing to the fibril core, and removing it makes everything worse: the critical concentration for fibrillation falls from 70.3 ± 9.5 µM for the full-length hormone to 8.83 ± 1.5 µM for the 1–34 fragment, the lag phase shortens from 40–80 hours to 5–40 hours, and full-length fibrils were more than 90% dissolved in 1 M urea while the truncated fibrils released no significant extra monomer up to 5 M (Sachan et al., 2026). Truncating a hormone to its active pharmacophore made it more prone to stable fibrils. The same paper supplies the countervailing point: parathyroid hormone fibrils release monomer at 0.46 ± 0.01 h−1, and that reversibility is what lets the hormone use fibrils as a storage form. Meanwhile sorbitol, a standard stabiliser, shortened the fibrillation half-time of a model protein from 3.54 ± 0.29 to 1.91 ± 0.27 hours at 750 mM while leaving the elongation rate unchanged — acting on nucleation, in the direction no formulator wants (Rahimi et al., 2026). That was κ-casein, not a drug product; the transferable part is the mechanism.
Whether this applies to peptide drugs as a class is answered here only by prediction, and the prediction is not reassuring: on axes of aggregation and phase-separation propensity, insulin, glucagon, prolactin and calcitonin all fall in the high-aggregation quadrant, and insulin's own amyloid structure has been solved (Lobo et al., 2026). What these sources cannot supply is worth naming: no accelerated-stability dataset for a marketed peptide, no subvisible particle counts, and no systematic study of the stresses that actually cause aggregation in manufacture and shipping — air–liquid interfaces, silicone oil, pumping, freeze–thaw. Almost every kinetic experiment quoted above was quiescent or shaken at one fixed speed.
Part Two set out to replace one picture with a distribution, and this is the form the replacement takes. A sequence specifies a population of shapes. Which member you observe depends on the solvent, the temperature, the concentration, the partner and the instrument. Some members bind, some aggregate, and the ones that aggregate can teach the rest to do the same. Part Three asks what forces set the shape of that population in the first place, and what each of them costs.
10Five forces and one accountant
Part Two ended with a warning about solvents, and the warning has a consequence that is easy to miss. Magainin‑2, a 23‑residue antibacterial peptide, gives a circular dichroism spectrum in 10 mM phosphate at pH 7 with a strong negative band near 203 nm and a shoulder at 225 nm — a coil. Titrate dodecylphosphocholine into the same cuvette past its critical micelle concentration, to 1.4 and then 5.7 mM, and the spectrum acquires minima at 208 and 222 nm and a maximum at 195 nm — a helix (Sancho‑Vaello et al., 2025). Nothing about the molecule changed. What changed is that a surface appeared which was willing to pay for the helix.
That is the subject of this Part. A conformation is populated when something repays the cost of adopting it, and a complex forms when something repays the cost of forming it. Five kinds of payment are available to a polypeptide — hydrogen bonds, salt bridges, van der Waals contact, the hydrophobic effect and aromatic interactions — and one accountant decides whether the transaction goes through. The accountant is the second law, and it charges for two things structural pictures do not show: the water that had to be moved, and the freedom the chain gave up.
Say at the outset what this document cannot give you. The corpus assembled for it contains no measured free energy for a single hydrogen bond, buried or exposed; none for a single salt bridge; no coefficient in calories per square Ångström for burying non‑polar surface; and no double‑mutant cycle with a coupling energy. Four staples of the textbook account, all absent. What the corpus carries instead is better than a table of averages: several cases where the value a chemist would have predicted from the picture turned out to be wrong in sign.
Hydrogen bonds and salt bridges, which are counted rather than weighed
Both are geometric objects in this literature, not energetic ones. A hydrogen bond is a donor–acceptor separation under 3.5 Å with a donor–hydrogen–acceptor angle above 135° (Hicks et al., 2021), or whatever the PISA program assigns from a set of coordinates (Ishii et al., 2021); a salt bridge is likewise an assignment. Both are reported as counts. The nearest thing to an energy is a software parameter: a rigidity analysis of the complement protein C5 admitted hydrogen bonds at calculated cut‑offs of −0.5 to −2.0 kcal/mol in turn, and the answer changed unevenly with the setting — one domain held 36–61 % of its Cα atoms in rigid clusters even at the strictest cut‑off, while another was rigid at −0.5 kcal/mol and below 10 % by −1.0 kcal/mol (Zhivnov et al., 2026). The network is graded, much of it is worth under a kilocalorie per mole even by the program's own reckoning, and which bonds you count decides which parts of a molecule look rigid.
Counting them does not rank them. Of four antibody fragments raised against the same trimethylated‑lysine peptide, two made 17 hydrogen bonds and 4 salt bridges to it each, and two made 7 hydrogen bonds and no salt bridges at all; by surface plasmon resonance the affinities ran 3.0 × 102, 14, 0.90 and 54 nM, so the two binders with less than half the polar contacts are the second‑ and third‑tightest of the set, and the weakest of the four is one of the high‑contact pair (Ishii et al., 2021).
What the corpus supplies in place of an energy is a demonstration that a charge is never simply positive. In the disordered N‑terminal region of the mycobacterial protein ChiZ, every residue with a membrane‑contact probability above 0.25 in simulation was an arginine, with Arg37 highest at 60 %. But the acidic residues Asp11, Asp20 and Glu28 all lie in the N‑terminal half, and the arginines near them bond those carboxylates rather than the lipid: Arg25, flanked by Asp20 and Glu28, contacts the membrane less than either neighbour. The sequence is using intramolecular salt bridges as negative design. The same study fixes how sharply electrostatic membrane binding switches on — no significant loss of amide crosspeaks against a bilayer 20 % acidic, and essentially every crosspeak broadened beyond detection at 70 % acidic (Hicks et al., 2021). The recognition is of a surface charge density, and it has a threshold.
Van der Waals contact, and the myth of the well‑packed core
The one place the corpus puts a measured number on a non‑covalent force is packing, and the result argues against the received picture. In a de novo designed protein, ten leucine and isoleucine residues in the hydrophobic core were changed to valine — the deletion of one methylene group each, nothing more. Melting temperature, followed by circular dichroism at 222 nm and fitted to the Gibbs–Helmholtz equation, fell from 129.6 °C to 106.1 °C; the change in unfolding free energy estimated from that fit was −5.5 kcal/mol, against −5.8 kcal/mol calculated independently from the non‑polar surface buried on folding (Koga et al., 2020). Measurement and surface‑area calculation agree to within half a kilocalorie, which is the kind of agreement that gets quoted.
Everything else undercuts the quotation. The gutted variant still folded to the same topology, with a backbone Cα RMSD of 1.4 Å by NMR and a well‑dispersed spectrum rather than the smear of a molten globule, despite visible cavities and a packing score of 6.0 against −0.45 for the parent, where lower is better packed — and it melted at 106 °C. Changing nothing but the length of the loops, meanwhile, abolished folding altogether. Nor is every buried residue worth the same: of the ten substitutions made singly, one cost 6.3 °C of melting temperature while four changed it by less than one degree. What distinguished them was contact topology — residues touching both adjacent and distant structural elements carried the stability, those touching only their neighbours carried none that could be measured, and Rosetta's predictions tracked neither at single‑residue resolution. Burial is not contribution.
The hydrophobic effect and aromatic contact
The hydrophobic effect is bookkeeping about water. Indirectly, and by secondary report, moving a side chain from water into a membrane interior costs roughly 3–5 kcal/mol for uncharged polar groups such as threonine and glutamine against more than 14 kcal/mol for formally charged ones, the penalty on ionisable groups being partly relieved by pKa shifts of +2 to +5 units for glutamate and −4 to −5 for lysine in the low‑dielectric interior (Sun et al., 2026, citing earlier scales). The ratio — roughly threefold between burying a polar group and burying a charge — is the part worth carrying forward. Directly, and by calorimetry, the effect is measured as an entropy, and that measurement belongs to the accountant.
Aromatic interactions are visible everywhere in this corpus and priced nowhere. The four antibody fragments above all recognise their trimethylated lysine by building a cage of aromatic rings around the methyl groups, and the composition of the cage differs every time — two tyrosines and a tryptophan, two phenylalanines and a tyrosine, three tyrosines, three tryptophans — so what is being satisfied is a property of aromatic rings in general (Ishii et al., 2021). Aromatic contact also holds molecules to each other: the 1.05 Å crystal structure of magainin‑2 shows two antiparallel helices meeting across a 510 Å2 interface built from five phenylalanine–phenylalanine contacts, all eight lysines on the opposite face (Sancho‑Vaello et al., 2025). And a cation sitting over a ring face is directly observable in molecules too small for anything else to be responsible: in four‑residue peptides, upfield ring‑current shifts on basic side‑chain nuclei appeared for three of four, and whether they appeared depended on the register (Mitchell et al., 2022).
The accountant
Now the arithmetic that makes the rest of this document behave sensibly. Four cationic‑aromatic tetrapeptides were titrated calorimetrically into large unilamellar vesicles of 80:20 phosphatidylcholine and cardiolipin (Mitchell et al., 2022). All four bound with statistically indistinguishable affinity, dissociation constants of 27.5 to 39.5 µM and free energies of −26.2 to −25.9 kJ/mol — a respectable binding event, roughly −6 kcal/mol. Isothermal titration calorimetry separates it into its parts, and the parts are startling: the binding enthalpies were −5.1, −4.5, −3.2 and −7.1 kJ/mol, while the entropic term TΔS ran from +19.0 to +22.1 kJ/mol. The free energy is almost entirely entropic, with a small enthalpic top‑up: roughly −5 kJ/mol of enthalpy against about +20 kJ/mol of TΔS.
Now watch what happens when a chemist tries to improve it. The four peptides differ in the polar groups on their aromatic side chains — two indole NH groups, one phenol OH, one phenol OH, none — and the enthalpy tracks that series faithfully, spanning a 2.2‑fold range. The free energy does not move at all: it stays within 0.3 kJ/mol across the set, and the affinities are not significantly different. Every kilojoule of enthalpy bought by adding a hydrogen‑bond donor was handed straight back in entropy. This is enthalpy–entropy compensation, measured rather than asserted, and it is why a designer who optimises the enthalpic term alone gets nothing for the work.
The compensation is not the whole ledger. There is a second, separable cost: the flexibility the chain surrenders. That has been priced too, in a designed repeat protein into whose loop seventeen different disordered sequences were grafted, up to 58 residues long (Ripka et al., 2021). By guanidinium denaturation followed by tryptophan fluorescence, fitted to a two‑state model with a shared m‑value, the free energy of unfolding fell from −6.6 kcal/mol for the parent to −4.71 kcal/mol for the longest graft — about 1.9 kcal/mol to tether 58 disordered residues into a fold. Stability correlated with loop length at a Spearman coefficient of −0.98 and not at all with charge; two loops of identical length but very different composition gave indistinguishable stabilities, −5.56 against −5.42 kcal/mol. The price of a tether is set by its length and is nearly blind to what it is made of, and it is sublinear — six inserted residues cost about 14 °C of melting temperature, 48 and 58 only 20 °C. In the same seventeen proteins solubility behaved as a wholly separate property, correlating with net charge per residue at −0.86 and not with length at all: a preview of the argument in Part Five that "stability" is at least four unrelated things.
The payoff: a good‑looking interaction can be worth nothing, or less
If the reader takes only one thing from this section, it should be this. A pair of chimeric peptides was built to bind an RNA G‑quadruplex, and each was simulated for 500 ns to identify which residues spent the most time hydrogen‑bonded to the target (Mou et al., 2026). In the first chimera the three highest‑occupancy residues were a tryptophan, a lysine and an arginine; replacing any one with alanine barely changed the measured dissociation constant, and only replacing all three at once cost affinity. In the second — the same two modules fused in the opposite order — the three highest‑occupancy residues were a histidine and two arginines that stayed associated with the RNA throughout. Substituting any of them, singly or all together, made the peptide bind better; the authors attribute the effect to local conformational strain rather than to the contacts themselves. Occupancy is not contribution.
The same lesson arrives from a different direction. The standard explanation for why some cyclic peptides cross membranes is the intramolecular hydrogen bond: a chameleonic bond that folds back and shields an amide NH from water, lowering the cost of leaving the aqueous phase. Across an experimental permeability dataset for cyclic peptides of 3 to 15 residues, with Boltzmann‑weighted bond counts computed from a conformational search, the correlation with measured permeability was −0.24 against Caco‑2 and −0.16 against PAMPA (Sun et al., 2026). Negative: the more permeable peptides had numerically fewer of the bonds supposed to make them permeable. The corpus does hold the opposing case — peptidomimetics built from pyrrolinone rings whose repeating intramolecular hydrogen bonds were confirmed by infrared spectroscopy and by the temperature dependence of the NH chemical shifts, and which behaved as more cell‑permeable than the analogous peptides (Smith et al., 2011). The two need not conflict: one is a single designed scaffold in which the bonds also enforce an extended conformation, the other a population statistic over a diverse chemical space. The defensible statement is that intramolecular hydrogen bonding can lower desolvation cost in particular designs and is not what distinguishes permeable macrocycles from impermeable ones on average.
The third case is the most practically damaging. Of the four antibody fragments discussed above, the two that made 17 hydrogen bonds and 4 salt bridges made several of them to the free carboxylate at the C‑terminus of the synthetic peptide they had been raised against. Amidate that terminus — a change of one atom — and their binding response nearly disappears, while the two fragments that had ignored it are unaffected. In practice, the tightest binder of the panel at 0.90 nM gave only a faint band against the full‑length protein on a Western blot, while a fragment sixty times weaker gave a strong specific one (Ishii et al., 2021). A structurally visible, chemically reasonable, energetically productive set of interactions turned out to be an artefact of the assay construct. In a real protein that carboxylate does not exist.
11Amphipathicity, and a molecule with two faces
An antimicrobial peptide has a problem that looks impossible. It must destroy a bacterial membrane while leaving the host's own membranes alone, using a molecule of twelve to thirty residues that carries no enzyme, no receptor and no recognition domain. The solution is a consequence of Section 06: a helix advances 100° per residue, so a residue and the one four places later end up on nearly the same side of the cylinder. Segregate the charged residues into positions that fall on one side and the greasy ones into positions that fall on the other, and the wound‑up chain has two faces. Neither face exists in the sequence. Both are made by the winding.
This is the cleanest case in the whole document of a property that belongs to sequence order rather than to composition, and it can be summarised in a single scalar. The mean hydrophobic moment is the vector sum of side‑chain hydrophobicities taken around the helical wheel; it is large when the two classes segregate and small when they are interleaved. In a thirteen‑member series of analogues of a thrombin‑derived peptide, the computed moment was raised from 0.610 in the parent to 0.877 in the most extreme analogue by replacing phenylalanine and glutamine with lysine and arginine and moving a tryptophan, and the geometric‑mean minimum inhibitory concentration fell from 57.0 to 14.25 µM (Jahan et al., 2025).
The therapeutic window, however, turns over. The analogue with the highest moment was haemolytic from about 40 µM, giving a therapeutic index of 2.81 — worse than the parent's 3.68. Its immediate neighbour in the series, whose computed moment differs by 0.005, reached 10 % haemolysis only at 135 µM and a therapeutic index of 9.47; a third analogue at 182 µM and 9.03. A descriptor that cannot resolve a 48‑fold difference in haemolytic concentration is not measuring the thing that decides selectivity. In the same series, swapping tryptophan for isoleucine — a change of less than 0.02 in the moment — raised the minimum inhibitory concentration against E. coli in 2.5 mM calcium chloride from 32 to 128 µM. Divalent‑cation screening of the electrostatic approach step is simply not in the descriptor.
That the moment is nonetheless a real physical quantity, and not merely a fitted number, is established elsewhere and by loss of function. A cryo‑EM structure of a Plasmodium invasion complex at the host membrane resolved seven strongly amphipathic helices with hydrophobic moments of 0.3–0.6 on the authors' normalisation, their leucine‑ and phenylalanine‑rich faces buried in the detergent belt and their basic faces at the headgroup interface. Isolated peptides of those helices perforated liposomes, at a frequency rising with amphipathicity; mutants designed specifically to lower the moment produced no membrane deformation and no perforation at all (Haile et al., 2026). That normalisation and the one quoted above are different scales and cannot be compared directly.
How the two‑faced molecule tells a bacterium from a host cell is a question about charge density, and the corpus answers it with the experiment already met in Section 10. A disordered, arginine‑rich sequence showed no detectable engagement with pure phosphatidylcholine liposomes, with 4:1 DOPC:DOPE, or with 4:1 POPC:POPG — 20 % acidic lipid. At 70 % acidic lipid essentially every amide crosspeak in its NMR spectrum broadened beyond detection (Hicks et al., 2021). Bacterial membranes carry anionic phospholipid on their outer face; mammalian plasma membranes largely do not. The discrimination is a threshold in surface charge, and the peptide does not need to recognise anything.
What happens after contact is a second, separable question, and the same study answers it unexpectedly: the peptide bound the anionic bilayer tightly and did not fold. Solid‑state NMR experiments reporting on mobile sites gave abundant crosspeaks at the same undispersed chemical shifts as the free protein, while the complementary experiment, which reports on rigid sites, gave only a handful — consistent with a single four‑residue stretch ordering in part of the population. Simulation found no gain in secondary structure on binding at all. A peptide can pay for membrane binding electrostatically and decline to fold, which is worth holding against the magainin case where the same driving force does produce a helix.
What the corpus does not contain
The textbook demonstration of this section is the scrambled control: take an amphipathic antimicrobial peptide, reshuffle its residues at constant composition, and watch the activity disappear. That experiment is not in this corpus, and this document will not pretend otherwise. Three composition‑preserving experiments were found, and none of them abolishes anything.
The first is a scrambled module. A module derived from a helicase was fused to a selected quadruplex‑binding peptide, and the fusion bound its RNA target with dissociation constants of 285 ± 20 nM and 99 ± 6 nM in the two orders, against 1.52 µM and 2.23 µM for the modules alone. Replacing one module with a scrambled version of identical composition gave 447 ± 37 nM and 203 ± 14 nM — a loss of only about 1.6‑ to 2‑fold, and still three to five times better than either module by itself for the first of those and seven to eleven times for the second (Mou et al., 2026). Order mattered. Bulk composition and charge carried most of the binding.
The second is a circular permutation. A short "tadpole‑like" peptide with a hydrophobic tail and a helical head was rearranged head‑to ‑tail — an operation that preserves composition exactly — and the rearranged peptide had minimum inhibitory concentrations of 3.13–6.25 µM against methicillin‑resistant S. aureus, where the natural parent's minimum inhibitory concentrations against the same organism are 50–100 µM (Tang et al., 2023). Order sharpened an activity the parent already had, by about sixteenfold, rather than destroying it, and the comparison against the natural parent is not itself composition‑controlled.
The third is the chirality series, which belongs to Section 13 and is the strongest of the three. What the reader should take from all of this is that the sequence‑order claim is well supported in direction and poorly supported in magnitude. Composition does a great deal of the work in these molecules. Order does the rest, and the rest is what decides selectivity.
12The covalent shortcuts
Everything so far has been reversible. A hydrogen bond breaks and reforms thousands of times a second; an amphipathic helix exists only while a membrane is there to hold it. There is one exception in the standard amino‑acid alphabet: two cysteines can be oxidised into a covalent sulfur–sulfur bond, and the shape that bond enforces stops being a matter of probability. This is the only tool in the natural set that converts a conformational preference into a constraint.
Nature's most developed use of it is the cystine knot, in which three disulfides are arranged so that the third threads the macrocycle formed by the other two together with the backbone between them. Add head‑to‑tail cyclisation of the backbone and the result is the cyclic cystine knot of the plant cyclotides: about thirty residues, six cysteines, three disulfides, no free termini, and a fold conserved across eight plant families (Cândido et al., 2026). The six loops between the cysteines are hypervariable, which is exactly what makes the scaffold useful to engineers: loops 1 and 4 are pinned by the knot itself, but the remaining four accept foreign sequence.
The best matched pair in the corpus makes the case in one comparison. A linear antibacterial peptide in human serum at 37 °C had a half‑life under one hour and had lost its main chromatographic peak entirely by four hours. The identical sequence grafted into loop 6 of a trypsin‑inhibitor cyclotide had a half‑life over thirty hours, with only about a 30 % reduction of that peak — an improvement the authors put at more than 30‑fold (Mourenza et al., 2026). The graft also bound its bacterial target about four times more tightly, at 33.4 ± 6.4 nM against 131.2 ± 30.2 nM by fluorescence polarisation, and acquired something the linear peptide never had: it entered human lung epithelial cells and killed bacteria hiding inside them, restoring half of host viability at 15.6 ± 3.2 µM where the linear peptide failed at 100 µM.
It also cost. Against the primary target strain the minimum inhibitory concentration rose roughly threefold, from 6.3 to about 17–19 µM, and the graft lost the broad activity of the parent against methicillin‑sensitive S. aureus and E. coli entirely. The linear peptide had permeabilised bacterial membranes strongly; the graft did not measurably permeabilise them at all. That is the trade in its clearest form. Constraint suppressed the promiscuous membrane mechanism and left the specific one, and whether that is a gain depends on which mechanism you wanted.
The bond is context‑dependent, not sacred
Four results in this corpus should cure anyone of the habit of treating a native disulfide as load‑bearing by default.
The first is the sharpest. In an eleven‑residue conopeptide closed by a single disulfide, replacing both cysteines with alanine raised the inhibition of the target receptor, from about 35 % to about 50 % at 1 µmol/L, which the authors attribute to the extra flexibility letting the peptide adapt to its site. Then the peptide was optimised at a different position, to about 70 % inhibition — and in that background the identical cysteine substitutions dropped inhibition to 26–47 %. In a further truncated analogue at 78 %, they dropped it to 12–33 % (Zhang et al., 2026). The same bond was dispensable in one sequence background and essential in another.
The second: in full‑length bovine lactoferricin, the native disulfide joining the peptide's two ends is essential for stabilising the β‑sheet in aqueous solution and yet has no significant functional role, because the oxidised and reduced forms have similar antimicrobial activity (Nguyen et al., 2010) — structural and functional necessity are different properties of the same bond. The third: a specific non‑native disulfide isomer of the conotoxin KIIIA, a "misfolded" product of the kind an oxidative folding reaction is designed to avoid, exceeds the natural peptide in potency against the sodium channel NaV1.7 (Mao et al., 2026, reporting earlier work). The wrong answer to the folding problem was the better molecule.
The fourth is a warning about the closing chemistry. Oxytocin's nine‑ residue ring is closed by one disulfide. In a human colonic model built from pooled faecal slurry under anaerobic conditions at 37 °C, native oxytocin retained 6.0 ± 1.36 % of its starting concentration after 1.5 h, and a linear analogue with both cysteines replaced by alanine was gone within the first thirty minutes. But rebuilding the same ring with a non‑reducible bis‑thioether staple was worse than the disulfide, not better: that variant was completely degraded in under an hour, which the authors attribute to the larger and more flexible ring the linker creates (Taherali et al., 2023). A ring is not a ring. Its span and geometry matter as much as whether it can be reduced.
The same study gives the disulfide's most useful positive result. Adding three D‑amino acids to the disulfide‑cyclised peptide raised survival at 1.5 h from about 6 % to 85.1 ± 0.21 %, and only that combination gave measurable permeation of excised rat colonic tissue in an Ussing chamber — an apparent permeability of 5.70 × 10−5 cm/s at one hour, against exactly zero for native oxytocin, which was undetectable in tissue or receiver chamber at either time point. Position within the ring mattered more than count: a single D‑residue outside the ring left 22.3 % intact, while two inside the ring left 3.70 % — less than the unmodified peptide.
The folding cost, which scales badly
Making these molecules is where the architecture charges for itself, and the corpus supplies three points on one scale from three laboratories. One disulfide closes by dissolving the crude peptide in aqueous acetonitrile, adjusting to pH 8–10 and stirring in air for 48 h (Zhang et al., 2026); or with iodine in 50 % aqueous acetonitrile, for a reaction time that source never states — a deliberately non‑selective oxidation that is safe only because the sequence contains exactly two cysteines (Dahal et al., 2023). Three interlocking disulfides in an inhibitor cystine knot required four days at 25 °C in deoxygenated buffer containing 5 mM reduced and 0.5 mM oxidised glutathione, with efficiency varying unpredictably between mutants differing at a single position elsewhere in the sequence (Wu et al., 2026). Two cysteines admit one pairing; six admit several, and the corpus never says how many — the combinatorial arithmetic of oxidative folding appears nowhere in this reading, nor any measured isomer distribution or folding yield.
Biology solves the problem with machinery that expression hosts do not have. Recombinant conotoxins made without the cone snail's own disulfide‑isomerase environment come out roughly thirtyfold weaker, or entirely inactive (Mao et al., 2026). Forming five disulfides on one 89‑residue conotoxin in E. coli required co‑expression of an engineered disulfide‑forming system and still yielded 2 mg per litre of culture, though the product did give a 1.67 Å crystal structure resolving all five bonds — two within one helical sub‑domain, one crossing between domains, one linking two further helices, and one tethering the free C‑terminus back onto the core (Müller et al., 2023). Prediction is often said to inherit the same difficulty, but the evidence base read for this document does not carry it: neither 2026 source consulted here reports that AlphaFold‑class methods degrade specifically on disulfide pairing, and one of them mentions AlphaFold only favourably (Ogundele et al., 2026; Zhang et al., 2026).
The bond that is meant to break
The most elegant result in this theme inverts the whole framing. A cyclic amphipathic peptide designed to carry small interfering RNA into cells is closed by a disulfide between two terminal cysteines — and it is closed that way because the cytosol is a reducing compartment. The ring is cut by cytosolic glutathione, proteases finish off the linearised peptide, and the cargo is released. The prototype gave over 90 % knockdown of its target transcript; the identical ring rebuilt through a non‑reducible thioether gave about 45 % in the same assay, and negligible knockdown in a second cell line (Jagrosse et al., 2023). The disulfide is not armour. It is a lock keyed to a compartment: robust in the oxidising extracellular space, designed to fail inside the cell.
The same series carries a caution that recurs throughout this document: the tightest binder of the cargo, at 0.87 ± 0.05 µM, gave only about 50 % knockdown, while two much weaker binders at roughly 14 and 27 µM gave about 86 % and 85 %. Holding on tightly is not the objective if the point of the molecule is eventually to let go.
13Chirality, and the mirror that does not work
Nineteen of the twenty proteinogenic amino acids are chiral, and life builds with one hand of them. Why it chose that hand is not a question this corpus can answer — nothing in the reading addresses the origin of biological homochirality, and nothing addresses spontaneous racemisation either. What the corpus answers, and answers unusually well, is the operational question: what happens when you break the convention, and what that tells you about what a target is actually reading.
Begin with the uncontroversial part. Proteases are enzymes and their active sites are chiral, so a D‑residue at a scissile bond is not a substrate. Substituting two residues with their D‑enantiomers at a trypsin site abolished cleavage there, confirmed by the disappearance of the corresponding fragment ion from the mass spectrum (Szarszoń et al., 2026). Enantiomer pairs of six structured peptides in concentrated cell lysate behaved the same way: every D‑peptide was entirely stable for 24 h while every L‑peptide was destroyed within 2–6 h — with one exception, a head‑to‑tail cyclic peptide whose L‑form was also stable, because cyclisation removes the free termini exopeptidases require. That study also found D‑peptides essentially non‑immunogenic in mice where their L‑enantiomers were not, and reported that around 30 % of approved peptide medicines contain a D‑residue while none is entirely D (Juraszek et al., 2025).
The reason none is entirely D leads to the second point: blocking a named protease site is not the same as surviving a body. A thirteen‑residue peptide had a half‑life of 23 ± 5 min in human plasma at 37 °C. Its two‑D‑residue analogue — the one whose trypsin site had been successfully blocked — managed 16 ± 2 min, slightly worse. Only the fully D‑configured version survived, losing about a quarter of the material over two hours (Ślusarczyk et al., 2026). Plasma carries many peptidases and they find the L‑residues you left behind.
The strongest single finding in this document
Now the interesting part. Retro‑inversion is the standard trick: write the sequence backwards and make every residue D. The justification, repeated in introductions for four decades, is that reversing the chain direction and inverting every centre cancel out, so the side chains end up in the same positions in space while the backbone becomes unrecognisable to proteases. A retro‑inverso peptide is supposed to be a conformational mirror image with a protease‑proof spine.
It is not. A study of an antimycobacterial peptide made all four members of the factorial — parent, retro, inverso and retro‑inverso — rather than the usual two, and put them through circular dichroism in pH 7.4 phosphate buffer at 25 °C (Glossop et al., 2025). The parent spectrum carries helical minima at 206 and 219 nm, a β‑sheet minimum at 212 nm, and an exciton band at 228 nm assigned to a tryptophan zipper. The all‑D inverso analogue gives an exactly mirrored spectrum, band for band, including that 228 nm feature — which is what an enantiomer must do. Both backbone‑reversed forms lose it entirely. Pure enantiomerisation preserves the tryptophan‑mediated structure; backbone reversal destroys it. Solution scattering agreed: the parent had a radius of gyration of 8.2 Å and a maximum dimension of 22.2 Å, the retro‑inverso analogue 10.2 and 30.6 Å. It is a different molecule, not a mirror of the first.
An independent study on a ten‑residue peptide reaches the same conclusion by a different route: molecular‑dynamics simulation placed the parent's helix in the N‑terminal half and the retro‑inverso analogue's helix in the C‑terminal half (Tang et al., 2023). Solution NMR in SDS micelles gives the reverse placement for the parent, its helix sitting near the C‑terminus, a divergence the authors attribute to the curved micelle surface. In the simulations the helix moved to the other end, which is what a reversed sequence requires and what a mirror image forbids.
The same factorial explains something the pairwise experiment never could. Against M. smegmatis the four analogues gave minimum inhibitory concentrations of 20 µM for the parent, 12.6 µM for retro, 20 µM for inverso and 1.3 µM for retro‑inverso — a roughly fifteenfold gain that neither single operation produced. Selectivity against a Gram‑negative comparator rose from 1–2 to 32, and the ratio of human‑cell toxicity to antimycobacterial potency from 5.3 to 368. Now the crucial control: against proteinase K, the parent and the retro analogue were both near‑completely digested in an hour while the inverso and retro‑inverso analogues were both highly resistant. The pure D‑enantiomer was every bit as protease‑proof as the retro‑inverso and gained nothing in potency. Resistance and potency came from different changes. Chirality inversion bought the stability; backbone reversal alone bought the activity. No pairwise comparison could have taken those apart.
It is not a rule. The same paper ran five further native/retro‑inverso pairs: three improved, by as much as sixteenfold in one case, one showed no improvement at all — attributed to a formal charge of +8 making a detergent‑like mechanism insensitive to conformation — and one pair was inactive in both forms at every concentration tested. Elsewhere in the corpus the strategy is described as buying very high protease resistance at the routine cost of affinity, and as being limited in scope to short linear and hairpin binders (Juraszek et al., 2025; Lucana et al., 2024). Retro‑inversion is a hypothesis to be tested per target, not a design rule.
Where chirality does not matter, and where a single centre does
The corpus also contains the clean negative control. A lipid bilayer is achiral, so a peptide whose target is bulk lipid should not care which hand it is, and it does not: all‑D versions of two thrombin‑derived analogues kept minimum inhibitory concentrations of 8–16 µM and permeabilised over 90 % of bacterial cells at twice that concentration where the parent managed 58 % (Jahan et al., 2025), and two cathelicidin enantiomer pairs killed with overlapping confidence intervals (Blower et al., 2015). A ten‑residue peptide with a defined head‑and‑tail architecture behaves differently, its straight all‑D mirror image losing activity the retro‑inverso form kept (Tang et al., 2023) — a long regular helix presents a near‑equivalent amphipathic face in either handedness where a short shaped peptide does not. Both results agree the target is not a chiral receptor.
At the other extreme, a single centre can move a property that composition cannot explain. In a cyclic heptapeptide library, replacing one arginine with D‑arginine roughly halved haemolysis, raising the 50 % haemolytic concentration to 165 µg/mL, while minimum inhibitory concentrations across the whole library stayed within 1.5–6.2 µg/mL — selectivity bought for nothing. Making the same peptides entirely D did not help: the all‑D analogues were as haemolytic as their all‑L parents, and showed the highest backbone flexibility of any variant in simulation (Lohan et al., 2025). More D is not better D.
One last observation closes the physical picture. Potentiometric titration of a thirteen‑residue peptide and of its complete enantiomer gave side‑ chain pKa values that are essentially the same — 3.75 against 3.94 for an aspartate, 5.83 and 6.64 against 6.08 and 6.62 for two histidines, 7.41 against 7.47 for the N‑terminal ammonium (Ślusarczyk et al., 2026). This is not a discovery; it is a requirement. Enantiomers have identical thermodynamics in an achiral medium, and water is achiral. Any measured difference between a peptide and its mirror image must come from an interaction with something chiral — a protease, a receptor, another chain — and never from the molecule alone.
Modified residues of exactly these kinds are not exotic. A 1.59 Å crystal structure of an acylated dual‑incretin analogue, a molecule of the class now dominating peptide therapeutics, resolves six non‑proteinogenic residues in a 39‑residue chain: two 2‑aminoisobutyric acids, a fluorinated α‑methylphenylalanine, an α‑methylleucine, an ornithine and a single D‑glutamate buried mid‑chain (Mitchell et al., 2026), most of them visible directly in the difference density. Single‑site D‑substitution and α,α‑dialkylation are routine chemistry, and the reason is the argument of this Part: each pays a small, specific amount of conformational entropy in advance for a specific gain.
What the reader should now be able to predict is modest but real. Given a sequence and a target, the question is no longer "what shape is it?" but "what is each candidate shape being paid, and who is paying?" A contact that looks good in a structure has told you about one term in a sum whose other terms are larger. A constraint that removes flexibility has bought something and charged for it. And a mirror image is only a mirror image if you did not also reverse the chain. Part Four takes this apparatus to the interface, where two residues out of thirty turn out to be carrying the transaction.
14Two residues out of thirty
A fourteen-residue recognition peptide, crosslinked into a bicycle, inhibits the interaction between the histone chaperone RbAp48 and its partner MTA1 with an IC50 of 12.3 ± 2.0 nM in a competitive fluorescence-polarisation assay. Replace each of nine of its positions in turn with alanine and measure again, and the nine substitutions do not behave like nine equal parts of one number. Five of them barely register. One of them costs a factor of 351 (Hart et al., 2021).
| Alanine substitution | IC50 (nM) | Fold loss |
|---|---|---|
| Thr677 | 12.4 ± 3.2 | 1.0 |
| Pro687 | 15.6 ± 2.2 | 1.3 |
| Lys686 | 19.8 ± 3.6 | 1.6 |
| Arg683 | 24.6 ± 3.9 | 2.0 |
| Tyr685 | 38.3 ± 1.9 | 3.1 |
| Arg679 | 77.4 ± 9.6 | 6.3 |
| Pro684 | 131.3 ± 15.3 | 10.7 |
| Lys678 | 262.9 ± 42.7 | 21 |
| Arg682 | 4319 ± 974 | 351 |
Every one of those residues is present in the bound peptide. All nine were inside a molecule that binds; none of them was added by accident. Yet five of them can be deleted, one at a time, for less than a factor of two — and one of them, an arginine whose guanidinium group makes an extended hydrogen-bond network with the receiving surface, carries more of the binding energy than the other eight together. Scrambling the same nine residues into a different order pushes the IC50 above 10,000 nM, which establishes that the sequence and not the composition is what is being read. This is the observation the rest of Part Four is built on: the residues that touch and the residues that matter are different sets, and nothing in a picture of the complex tells you which is which.
The method that produces such a table is alanine scanning, and its logic is subtractive. Alanine has a single methyl group where other residues carry rings, charges or hydroxyls, so replacing a residue with alanine truncates its side chain at the β-carbon while leaving the backbone geometry as nearly unchanged as any substitution can. The change in binding energy is then read as the contribution of the deleted part. That reading has two well-known soft spots, and the corpus assembled for this document illustrates both. Truncating a side chain does not simply remove a contact; it also leaves a cavity, which the solvent and the partner must then deal with. And alanine scanning assumes that what a residue contributes can be measured with the rest of the molecule held fixed — an assumption that fails, measurably, in more than one of the cases below.
The shape of the RbAp48 result is not idiosyncratic. A twelve-residue macrocycle, Z1, selected by phage display from a library of more than 109 sequences, binds the oncogenic RhoA G17V variant with a KD of 136 nM measured by biolayer interferometry. Its alanine scan divides the same way but at a different ratio: six positions — W3, F4, F5, W6, E8 and D12 — each cost more than ninefold, with the worst penalties falling in the contiguous aromatic block W3-F4-F5-W6, while the remaining positions cost only two- to fourfold (Abraham et al., 2026). The linear counterpart of the same sequence showed no detectable binding at all, which is the entropic argument of Part Three restated as a measurement. Note what the two scans have in common and where they differ. Both are strongly non-uniform. But the RbAp48 peptide has a single dominant residue and the RhoA macrocycle has a contiguous patch of six, so "hot spot" describes a distribution rather than a fixed number of residues, and the distribution is a property of the particular interface.
Until recently that claim rested on a scatter of individual scans. It is now measured exhaustively. Deep mutational scanning of five human PDZ domains against seven peptide ligands, read out by two selections in yeast and fitted to a three-state thermodynamic model that separates folding energy from binding energy, produced 21,802 free-energy measurements — 9,064 for folding and 12,738 for binding (Martí-Aranda & Lehner, 2026). Defining a hot spot as an interface residue whose substitutions exceed the interface median of 0.75 kcal/mol, the median interface contains six hot spots, which is 41.67 % of its contacting residues, with a range across the seven interfaces of 29.41 % to 61.54 %. Fifteen interface positions are never hot spots in any of the seven interactions, and eight of those contact the ligand in more than one domain. The energetically important residues are also the buried ones: median relative solvent-accessible surface area 0.14 for conserved hot spots, 0.33 for non-conserved hot spots, and 0.58 for interface residues that are not hot spots at all. That is the O-ring picture with numbers on it — a small buried centre that carries the energy, ringed by contacts that are real, visible in the structure, and worth nothing.
Which brings the argument to what a pharmacophore is. In its loosest use the word names a list of chemical groups — a positive charge here, an aromatic ring there, a hydrogen-bond donor at the end. Taken that way it predicts badly. Screening 222 sequences chosen purely because they matched the consensus motif of one of nine bait domains, only 44 of them — 20 % — bound at all (Madhu et al., 2026). Four fifths of the sequences carrying the right groups in the right order were not ligands. Taken seriously, a pharmacophore is not a list but a spatial arrangement with tolerances, and the tolerances are narrow enough to be measured. From an identical peptide sequence, changing only the geometry of the crosslinker that holds it — ortho-, meta- or para-xylene, or a meta-pyridine — moved the IC50 from 77.4 nM to 18.7 nM, 15.1 nM and 34.9 nM, a 5.1-fold spread from the substitution pattern on a single ring. Moving that same crosslinker to a different pair of anchoring positions in the same sequence inverted the ranking and cost far more: 321.7, 388.0, 2342 and 146.8 nM respectively (Hart et al., 2021). Nothing about the chemical inventory changed. The geometry did.
The same paper contains a small, useful correction of the kind this document is built to display. The crystal structure of the monocyclic precursor bound to RbAp48 (PDB 6ZRC) placed the xylene linker beside a receptor phenylalanine, and the authors proposed a π-stacking interaction. Rather than leave the proposal standing, they tested it: five linkers bearing electron-withdrawing and electron-donating substituents, which should have moved a π-stacking interaction in opposite directions, all lost between 1.5- and 2.5-fold. The contact is real; the mechanism attributed to it was wrong, and the correction came from chemistry rather than from a better structure.
Mutations do not add up
The deepest problem with reading an interface residue by residue is that the residues are not independent, and one recent structure supplies the arithmetic. Cryo-EM of the insulin receptor ectodomain bound to the bivalent antagonist Ins-AC-S2 revealed a contact patch not previously described, in which the ligand's B-chain N-terminus and part of its A-chain engage thirteen receptor residues across the fibronectin-type-III and insert domains (Vogel et al., 2026). Six substitutions across that patch were then tested one at a time in a functional antagonism assay, against 43 nM human insulin, reading phosphorylated AKT at Ser473 in cells overexpressing the human B isoform of the receptor. Wild-type Ins-AC-S2 gave an IC50 of 5.71 ± 0.73 nM. Each single mutant cost between 1.4- and 3.5-fold. All six together cost 9.5-fold. Multiply those six ratios — 2.00, 1.96, 2.47, 1.44, 3.45 and 2.66, an arithmetic this document performs and the authors never state — and independence predicts roughly 130-fold. The measured combination is fourteen times cheaper than that.
And then the same combined mutant was measured a second way. By isothermal titration calorimetry against the purified ectodomain, wild-type Ins-AC-S2 bound with a Kd of 0.5 nM and the six-fold mutant with 18.0 nM — a 36-fold loss of affinity against a 9.5-fold loss of potency in the cellular assay. Both numbers are correct. They are answers to different questions, and a reader who conflates them will mis-rank two molecules. That distinction runs through this whole theme. A bicyclic K-Ras binder bound the inactive GDP-loaded protein with a KD of 0.37 ± 0.09 μM and the active nucleotide-analogue-loaded form at 0.86 ± 0.28 μM, disrupted the Ras–Raf interaction with an IC50 of 3.4 ± 0.7 μM, and killed cells with an LD50 near 8 μM in one line and 17 μM in another (Trinh et al., 2016). The gap between those three numbers is not noise. It is conformational selectivity plus target abundance: cells hold 0.4–20 μM of the inactive form against 0–50 nM of the active one, so the inhibitor is sequestered by the state it binds best, and the authors say plainly they would have screened against the active state in hindsight.
There is a second failure mode of the method, and it is the mirror image of the first. Where a partner is engaged by several weak sites at once, mutating any one of them reports nothing. The ALYREF RRM domain binds six consensus motifs in the N-terminal region of POLDIP3, five of them detectably, with individual peptide affinities in the 80–150 μM range; the full-length proteins co-immunoprecipitate robustly; mutating a single site does not affect that co-immunoprecipitation, and only truncating the whole N-terminal region abolishes it (Madhu et al., 2026). An alanine scan on such a system would conclude that no residue matters. The same survey also reframes how weak most real recognition is: of 103 interactions retested by fluorescence-polarisation competition across 29 domains, 96 showed measurable binding spanning nanomolar to millimolar, with a median Kd of 30 μM.
The claim that got taken back
For about two decades the protein DLK1 was described as a non-canonical ligand of Notch, on the strength of two-hybrid interaction data and the downregulation of Notch target genes when DLK1 was expressed. It does not bind Notch. Surface plasmon resonance between the full extracellular domains of DLK1 and NOTCH1 detected no interaction even above 10 μM, while the canonical ligand DLL4 bound and fitted a 1:1 model in the same experiment; flow cytometry showed DLL4 but not DLK1 staining NOTCH1-overexpressing cells; immobilised DLK1 activated no Notch reporter, and 3 μM soluble DLK1 did not inhibit DLL4-driven activation (Antfolk et al., 2025). The interesting part is not the negative result but the explanation attached to it. The original two-hybrid assays were run in the reducing environment of the cytosol and nucleus, where the disulfide bonds that hold a Notch receptor's extracellular repeats together cannot form — so the reported interaction was with a conformation the receptor never adopts where it works. Sequence comparison agrees: DLK1 lacks the two domains Notch engagement requires, and the three bulky hydrophobic interface residues of the canonical ligand JAG1 are replaced in DLK1 by proline, serine and glycine.
The same paper then finds the real receptor and solves the complex. DLK1 EGF5-6 bound to the ACVR2B extracellular domain, by X-ray crystallography at 2.7 Å, PDB 9D20, burying 783.5 Å2 — slightly more than the 649.6–775.2 Å2 range of the receptor's canonical cystine-knot ligands, which it reaches with the major β-sheet of an EGF-like domain instead of a knot's finger hairpins. Alanine substitution here is almost binary: ACVR2B W78A and F101A each abolish binding outright, while R56A costs about sixfold, and on the ligand side a single charge reversal, R193D, abolishes binding both to recombinant receptor and at the cell surface. Two domains of a six-domain extracellular region carry the whole affinity — the EGF5-6 fragment binds at KD 1.0 μM against 1.4 μM for the full ectodomain — which is worth setting beside the opposite result from the RbAp48 series, where truncating a 26-residue peptide to the fourteen residues that looked critical in the crystal structure cost 196-fold, because two of three residues forming a hydrophobic cluster went with them. Whether a minimal fragment keeps its potency depends on whether the discarded material was contributing energy or merely present, and a structure does not distinguish those two states.
What the reader can now do with an interface picture is smaller and more precise than before. A contact map tells you which residues are within reach of each other; it does not tell you which ones pay, how much, or whether the payment survives when a neighbour is also removed. The energy is unevenly distributed, usually buried, sometimes concentrated in one residue and sometimes spread across a contiguous patch, and it is measurable only by mutation and only in the assay whose question you actually care about. When a paper reports that a residue is "critical", the first question is which measurement it was critical in.
15Message and address
Take a peptide agonist, remove eight residues from its amino terminus, and put the truncated molecule back on cells carrying its receptor. Exendin(9-39), the N-terminally truncated form of the lizard peptide exendin-4, binds the human GLP-1 receptor on fixed HEK293T cells with an equilibrium dissociation constant of 3.93 nM, calculated from an association rate of 1.64 × 106 M−1s−1 and a dissociation rate of 2.09 × 10−3 s−1 by surface plasmon resonance microscopy, fitted to a 1:1 model (A Thirumurthy et al., 2026). It binds tightly. It does not activate. And the histogram of dissociation constants across 600 regions of interest on the chip is a single Gaussian — one binding mode, not two.
Every full agonist measured in the same experiment gave two. Native GLP-1 resolved into modes at 149 nM and 1.14 nM; exendin-4 at 92.9 nM and 3.06 nM; liraglutide at 29.3 nM and 1.22 nM. In each case the two modes shared similar association rates and differed roughly twenty- to fiftyfold in dissociation rate. The authors read the fast-off mode as the peptide held by one arm and the slow-off mode as the second arm engaging, since an on-rate should not change once the ligand is already receptor-bound. That reading is an interpretation of avidity kinetics rather than a direct structural observation, and it should be labelled as such. But its content is the two-domain model of class B peptide receptors, measured on intact cells: the truncated antagonist engages one domain, the agonists engage two.
The model itself is stated the same way by several independent sources. The peptide's C-terminal half docks into a groove on the receptor's extracellular domain — a globular α-β-βα fold held by three conserved disulfide pairs — at typically micromolar affinity; the peptide's N-terminal half then inserts into the transmembrane helical bundle, rearranges it, and enables G protein coupling; and the two steps together give nanomolar binding (Babin et al., 2025; Liu et al., 2025). Because the activating element is the N-terminus, cutting from that end characteristically produces a competitive antagonist — a molecule that still finds the receptor and no longer switches it on. The framework was originally established by solution NMR on a corticotropin-releasing factor receptor extracellular domain in complex with the antagonist astressin (PDB 2JND), and the extracellular-domain fold was first seen in atomic detail in an X-ray structure of the parathyroid hormone 1 receptor ectodomain bound to its hormone at 1.95 Å, PDB 3C4M, where the groove cradles the hormone's amphipathic helix; the first transmembrane-domain crystal structure of the family, the glucagon receptor at 3.4 Å, PDB 4L6R, came five years later.
The clearest quantification of the two steps comes from a single receptor measured as a ladder. Against the calcitonin-receptor-like receptor with RAMP1 in COS-7 membranes, by nanoBRET equilibrium competition with Cheng–Prusoff correction, the C-terminal fragment CGRP(27-37) — which can reach the extracellular domain and nothing else — bound with a Ki of 1 μM. The longer truncation CGRP(8-37), which retains the transmembrane-engaging segment but has lost the activating residues, bound at 96.7 ± 2.4 nM in the G-protein-uncoupled state and 92.1 ± 1.2 nM coupled. Full CGRP(1-37) bound at 74 nM uncoupled and 3 nM with G protein present (Babin et al., 2025). These are binding affinities, not functional potencies, and the fact that only the full agonist shows a G-protein shift is the expected behaviour of an agonist rather than an additional finding.
Three further measurements in the same study put the ladder on firmer ground. Apparent melting temperature of the detergent-solubilised receptor rose step-by-step with what the peptide could reach: 38.2 °C ligand-free, 39.8 °C with the (27-37) fragment, 43.6 °C with (8-37), and 43.9 °C with the full peptide — an apparent midpoint from a sigmoidal fit to native-PAGE densitometry, not a calorimetric value. Functionally, a C-terminally truncated adrenomedullin fragment comprising only the transmembrane-binding portion produced detectable cAMP agonism at 30 μM, and the (27-37) peptides failed to antagonise it while the (8-37) peptides significantly diminished it. And kinetically, the address alone lets go almost at once: wild-type CGRP(27-37) has a dissociation rate of 12.2 min−1, a binding half-life of 3.6 seconds, whereas an affinity-matured (8-37) probe dissociates biphasically with half-lives of 13 and 52 minutes and a slow-phase residence time of 76 minutes. Holding on is what the second contact buys.
The structural side of the model is now unusually well populated. Cryo-EM structures of the family, reported with resolutions, include the GLP-1 receptor with GLP-1 at 2.10 Å (PDB 6X18) and 4.10 Å (5VAI), the GIP receptor with GIP at 3.24 Å (7RA3) and 2.98 Å (7DTY), and the glucagon receptor with glucagon at 3.70 Å (6LMK); the dual agonist tirzepatide has been solved on both the GLP-1 receptor (7FIM) and the GIP receptor (7FIY) at 3.40 Å (Sangwung et al., 2024). Across these, the peptide's N-terminus is inserted in the helical bundle and its C-terminus is held by the extracellular domain, which is what the model predicts. What the corpus cannot supply is the next level of detail: the review that carries the resolutions reports the conservation of the binding pocket as percentages — contacting residues of GLP-1R, GIPR and GCGR share a minimum of 40 % identity and 60 % similarity, the pairwise pocket figures running 47, 42 and 44 % identity, with no whole-receptor identity given anywhere to set them against — and does not enumerate a single named contact. No source in this evidence base gives a residue-by-residue contact list for a peptide bound to a class B receptor, and none is invented here.
One further experiment localises the two ends independently of any structure. Fluorophores were attached to secretin at positions −1, 13, 22 and 29 and their environments probed by anisotropy and iodide quenching against the active receptor structure PDB 6WZG. The C-terminal probe reported an extracellular-domain environment, the N-terminal probe reported deep burial in the helical bundle, and positions 13 and 22 sat at the junction. Only the N-terminal probe cost the peptide both binding affinity and cAMP potency; the C-terminal probe bound and signalled normally (Harikumar et al., 2024). Decorating the address is tolerated. Decorating the message is not.
Order alone
If a peptide's parts have different jobs, the strongest possible demonstration would be two molecules containing the identical parts, making the identical contacts, with opposite pharmacology. That molecule pair exists. S961 and S597 are bivalent insulin-receptor ligands built from the same two binding segments joined by a linker; S961 carries the site-1 segment first and is an antagonist, S597 carries the site-2 segment first and is an agonist. Cryo-EM of the receptor bound to S961 at 3.68 Å, reconstructed from 378,182 particles, shows each segment occupying the same receptor surface in the same orientation as it does in isolation and in the agonist complex (Vogel et al., 2026). The difference is geometric. With site 1 membrane-proximal, the receptor is held in an inverted-V with its fibronectin stalks about 137 Å apart — against about 112 Å in the unliganded ectodomain — and cannot close; with the segments arranged parallel to the membrane, the receptor rises into the active geometry. Isolated, the site-1 segment is itself an agonist with an affinity of 11–40 nM and the site-2 segment a weak antagonist; assembled in one order they antagonise with reported half-maximal inhibitory concentrations of 2.4 nM at one receptor isoform and 7.4 nM at the other, and in the other order they activate.
The identity of a single terminal residue can do something similar to signalling character rather than to direction. Tirzepatide behaves as a G-protein-biased agonist at the GLP-1 receptor, with reduced β-arrestin recruitment; mutating its N-terminal tyrosine to the histidine that native GLP-1 carries increased activity and restored some β-arrestin recruitment (Sangwung et al., 2024). The first residue of the message is not merely required for activation; it is part of what kind of activation occurs.
What changes for the reader is the meaning of a truncation experiment. A peptide that has been shortened from one end and lost activity has not simply "lost potency"; which end was cut determines whether what remains is a weak version of the original or a different pharmacological object. And the two questions a truncation raises — does it still bind, and does it still switch the receptor on — require two different instruments, which is why the affinity ladder above is quoted entirely in inhibition constants and the activation evidence entirely in cAMP.
16What conservation tells you, and what it does not
Somewhere between 450 and 500 million years ago, in the deuterostome lineage, a disordered binding domain called NCBD acquired a disordered partner called CID, and the two have been binding each other ever since. Their interface contains a salt bridge — NCBD Arg2104 to CID Asp1068 — that has survived the whole of that interval. Ancestral sequence reconstruction makes it possible to build the Cambrian-like versions of both proteins and measure them alongside the modern human pair. Breaking the salt bridge slows association about twentyfold in the human complex and less than twofold in the ancestral one. Molecular dynamics restrained by the measured φ-values shows it is populated in the native state and formed in neither transition-state ensemble. One of its two residues, mutated away, increases both the association and the dissociation rate, giving a negative φ-value near −0.4 — meaning the conserved aspartate makes an unfavourable interaction at the rate-limiting step (Karlsson et al., 2020).
Half a billion years of conservation, and at the step the residue pair appears designed for it does nothing, or worse than nothing. This is not a paradox. It is what conservation actually measures. An alignment records that substitutions at a position have been removed by selection; it does not record why, and the reason need have nothing to do with the function under study. A residue can be invariant because it is required for folding, for expression, for a second interaction, for avoiding aggregation, or because the position sits in a region with a low substitution rate for reasons no one has identified. Conservation is evidence about constraint. Importance to a particular function is a separate claim that requires a separate experiment.
Two peptide families make the point from the opposite direction, by conserving the part that gets thrown away. Cathelicidins take their name from a conserved N-terminal cathelin domain with high homology to a cathepsin inhibitor; the antimicrobial peptide itself is the C-terminal region, and it is hypervariable. In buffalo against cattle, the exon and intron regions encoding the propiece are highly conserved and one mature peptide is identical between the species — while the mature region of another family member is 94 residues in buffalo against 43 in cattle (Brahma et al., 2015). Within buffalo alone, a single cathelicidin gene carries twelve distinct mature-peptide variants across five breeds, with individual animals carrying four to six of them, so there is no single sequence to align in the first place. Those variants are not equivalent: against one Pseudomonas aeruginosa reference strain, minimum inhibitory concentrations across the series run from 0.2 μM to more than 200 μM, and the variants differ measurably in helical content by circular dichroism as well as in potency. The source prints no sequences and attributes the spread to tryptophan and arginine content rather than to any one substitution. Conotoxins do the same thing: gene superfamilies are defined by the conserved signal peptide, which is cleaved off during maturation, while the mature toxin is hypervariable, and fewer than 3 % of an estimated millions of distinct sequences have been characterised structurally or functionally at all (Mao et al., 2026). In both families, the conserved region is the discarded one and the functional region is the fastest-moving.
Conservation also fails in the other direction, by being uninformative exactly where a drug designer needs information. The human dopamine D2 and D3 receptors share 78 % sequence identity across their transmembrane regions and 100 % identity at the orthosteric site — every residue where the natural transmitter binds is the same. A bitopic agonist nonetheless achieves fiftyfold selectivity for D3, and cryo-EM at 3.05 and 3.09 Å shows why: its secondary pharmacophore reaches past the conserved pocket to a groove at the extracellular vestibule, where D3 reads 93-GGV-95 and D2 reads 98-GE-99. Deleting the extra glycine abolishes activation entirely (Arroyo-Urea et al., 2023). Selectivity lives where conservation stops. The same study separates two things that alanine scanning usually blends: substitutions at the conserved orthosteric site reduced potency by more than a hundredfold in half-maximal effective concentration, while substitutions at the divergent secondary site mostly reduced maximal efficacy instead — one of them by only about threefold in potency while markedly reducing efficacy. The authors also note, unusually, that this residue reduces efficacy even for a ligand that never reaches it, so part of its effect is on the receptor's intrinsic function rather than on ligand contact, and they decline to quantify its binding contribution for that reason.
Where conservation is used as a predictor, its performance can be measured, and in disordered regions it collapses. Benchmarked on 2,104 ClinVar-labelled missense variants across 290 genes, all falling in regions predicted disordered, the unsupervised predictor AlphaMissense reached a receiver-operating characteristic area under the curve of 0.907 and a precision-recall area of 0.733; a protein language model reached 0.847 and 0.606; and EVE — the other unsupervised model, built explicitly on sequence conservation — reached 0.715 and a precision-recall area of 0.500 against a random baseline of 0.14 (Gnanaolivu & Hart, 2026). Published figures for the same unsupervised predictor put it near 0.94 in ordered regions against 0.85 in disordered ones. Of the 15,999 ClinVar missense variants located in predicted disordered regions, 85.7 % remain of uncertain significance. Where a residue has no fold to hold it in place, evolutionary conservation stops being a usable proxy for functional importance — which is a strong hint that most of what conservation scores were tracking all along was structural constraint.
The consensus motif, conservation's most familiar output, is a similarly weak object. Across the highest-confidence tier of human SUMOylation sites, 2,485 of 8,639 — about 29 % — match the textbook core motif, and adherence falls to 21 % and 18 % in the successively lower-confidence tiers. The strongest single positional preference in the whole dataset is a 3.91-fold enrichment of glutamate at one position, measured against non-modified lysines observed in the same mass-spectrometry peptide pool (Al-Momani et al., 2026). The same study also supplies a small model of intellectual honesty about alignments: it found one strongly enriched non-canonical motif matching 294 sites on 224 proteins, noticed that all of them were zinc-finger proteins whose peptides mapped ambiguously across homologous domains, and concluded that its own apparent conservation signal was probably a peptide-to-protein mapping artefact.
None of this makes conservation useless, and the evidence base carries two cases where it earns its keep. The first is predictive and was confirmed: aligning DLK1 against the canonical Notch ligand showed the two domains required for Notch engagement to be missing and three bulky hydrophobic interface residues replaced by much smaller ones, and direct biophysics then found no binding (Antfolk et al., 2025). The second inverts the usual reading altogether. Mature growth factors of the TGF-β family are so conserved that clinical-stage antibodies raised against one cross-react with several others, and one receptor-based ligand trap was terminated early for toxicity attributed to that overlapping specificity. The way around it was to target the part of the same molecule with the most sequence diversity — the prodomain that is removed on activation. The resulting antibody binds a conformational epitope across three strands of the prodomain arm, burying 1,650 Å2, solved by X-ray crystallography at 2.79 Å, and its selectivity is attributed to precisely the divergence at that epitope (Dagbay et al., 2020). Conservation predicted cross-reactivity; variability predicted selectivity; both predictions were used deliberately.
The strongest positive result in the evidence base for the more ambitious use of alignments — covariation between positions, which promises to find functionally coupled residues without any structure at all — is a retrospective comparison. A previously published PDZ "sector", a set of twenty coevolving positions defined from sequence alone, was cross-tabulated against the 21,802 new energy measurements. Fourteen of the twenty sector residues turn out to be experimentally measured energetic hot spots, an odds ratio of 27.90 at p = 4.21 × 10−9 (Martí-Aranda & Lehner, 2026). That is a real, statistically strong correspondence between a sequence-only prediction and a direct measurement, and it is the best such result available here. It is also partial: six of the twenty sector positions are not hot spots, and eight hot spots are not sector positions.
An alignment, then, is a hypothesis generator with a known and quantifiable failure rate. It tells you where to look and how confident to be that something is under selection; it does not tell you what for. The operational consequence for reading a peptide paper is small and sharp: a sentence of the form "this residue is highly conserved and therefore critical" contains one observation and one inference, and the inference is the part that needs the experiment.
17Knowing versus modelling
The same enzyme, in the same unliganded state, crystallised twice. Overall the two structures agree: the average deviation between corresponding α-carbons is 0.49 Å. But three loops flanking the substrate cleft deviate by up to four times that, displacing two phenylalanine rings by 2–3 Å and narrowing the cleft, so that one packing shows an open active site and the other a closed one. Meanwhile a fourth loop, held by a hydrogen-bond network, superimposes at 0.25 Å with no conformational change at all — and even the ordered water molecules split, some sitting more than 1 Å from their positions in the other form while two conserved ones do not move (Saino et al., 2015). Both structures are real. At least one of them is not the population in solution. The authors' own stated motivation is the useful part: comparing packings of the same protein is how you find out which loops are mobile.
That is the honest starting point for a methods reckoning, because it shows both what crystallography does and how to catch it doing it. A crystal reports a conformation the molecule can adopt, selected in part by the lattice. For a folded protein this is usually a small correction. For a peptide it is not, because a short chain is solvent-exposed along its whole length and has no tertiary structure of its own to bury anything in; whatever burial happens has to be done by the lattice. Two documented consequences follow, and they point in opposite directions. A crystal structure of the phosphoinositide 3-kinase α catalytic subunit with its substrate soaked in places the inositol head roughly 6–7 Å from the phosphate it must attack, when the chemistry requires 1–3 Å; the structure is a real conformation and cannot be the catalytic one (Nussinov et al., 2023). And the atomic difference between two oncogenic mutations of the same K-Ras residue, which accounts for their different clinical behaviour, was invisible in crystal structures because lattice contacts stabilised an alternative loop conformation; it was seen by NMR and simulation instead.
The counter-case is equally instructive. A 42-residue acylated lipopeptide analogue of the incretin hormones crystallised in a lattice whose helical cores stack into square channels roughly 30 × 30 Å in cross section, and the seven disordered C-terminal residues and the entire fatty-diacid conjugate sit in those channels — unmodelled, but demonstrably able to occupy the void without clashing (Mitchell et al., 2026). Here the lattice hosted the disorder rather than removing it. The two positions are not contradictory: a lattice can hold a mobile loop in one of its accessible conformations while leaving a genuinely high-entropy segment unresolved, and the same practical test settles both — does the feature reproduce across independent crystal forms, and does an orthogonal solution method see it?
One more thing that same dataset makes visible is that the resolution figure printed in a structure's header is a decision. Two standard processing programs placed the diffraction limit of one dataset anywhere between 2.00 and 1.59 Å depending on which statistic was used, and paired refinement across the range produced nominally better agreement statistics at a conservative 2.5 Å cutoff — at which point an alternate conformation of one isoleucine disappeared from the map. The authors kept the higher-resolution cutoff and accepted the worse-looking numbers, and separately accepted an elevated free R-factor rather than fit unexplained density near the acylation site. Global quality metrics can penalise restraint and reward overfitting.
Cryo-electron microscopy, and what averaging costs
Single-particle cryo-EM builds a map by averaging over many copies of the molecule, and the copies are not identical. For the insulin receptor bound to the antagonist S961, template picking found 7,329,122 particles across 21,434 movies from five grids; iterative classification reduced that to 745,509 for a 3.78 Å reconstruction and then to 378,182 for the deposited 3.68 Å map. About one particle in twenty reached the final map, and discarding particles is what improved the resolution (Vogel et al., 2026). What the discarded heterogeneity was doing is sometimes recoverable and sometimes not. In the same map, one half of the receptor dimer showed no density at all for an entire fibronectin domain at a contour level appropriate for the rest of the structure — and focused refinement on each half separately still left one much better ordered, showing the motion originates within that half rather than between them. Three-dimensional variability analysis then split the particle set into two subsets at essentially identical global resolution, 3.99 and 3.97 Å, differing in whether the ligand's linker was occupied. Resolution and occupancy are independent quantities, and a missing domain in a cryo-EM map is a positive statement about dynamics, not an absence of the domain.
It follows that the headline resolution can mislead about what a map shows. For a second complex from the same study, the consensus map reached 3.64 Å, and a focused, symmetry-expanded local refinement reached only 3.89 Å — yet the worse map resolved a whole binding interface the better one did not, including a ligand N-terminal segment normally invisible in insulin-receptor structures, making contacts with thirteen receptor residues. In the other case, a GLP-1 receptor complex reported at 3.1–3.2 Å globally had local resolution of 2.8–3.4 Å at the ligand site — a band the source calls comparable to the global figure, straddling it rather than beating it. What licensed modelling the ligand was the well-defined density, not a better local number (Kawai et al., 2020). And the construct can set the ceiling more firmly than the processing: the same ligand gave 6.44 Å with the full-length receptor and 3.64 Å with the ectodomain alone, with combining the datasets giving no improvement. Finally, a small caution about what a structure's conditions are: the GLP-1 receptor complex above carried a positive allosteric modulator added during sample preparation to stabilise the complex further, its density was visible but its orientation could not be determined, so it was left out of the deposited model — and it was not present in the experiments showing the ligand works.
Nuclear magnetic resonance, and ensembles that need not exist
NMR reports a time and population average, which for a flexible molecule is not the same thing as a structure. At peptide scale the technique also starts from a disadvantage: the nuclear Overhauser effect, its primary distance observable, passes through zero for molecules tumbling at the rate a roughly 1,000-dalton macrocycle tumbles, so the measurement has to be rescued by deliberately slowing the molecule down — cyclosporin A was measured in a 70:30 chloroform/hexadecane mixture at 274 K to reach a rotational correlation time of 0.47 ns (Rüdisser et al., 2023).
The deeper problem appears when the restraints disagree with each other. For a second macrocycle in the same study, seven distance restraints involving a tryptophan indole were violated by more than 0.2 Å in over 80 % of a 100-structure bundle calculated using all the data. The response was to reclassify those restraints as conformationally averaged, recalculate without them, cluster the result into two states — "indole-in" and "indole-out", defined by their side-chain torsions — and split the averaged restraints between them. The populations of those two states were then set to 1/N, that is, assumed equal, because the paper states plainly that estimating them accurately would require more precise restraints than exist. Since the observable scales as the inverse sixth power of distance, short distances dominate the average disproportionately, and a sparsely populated compact conformation can dominate a measurement without being common. This is the mechanism by which a published NMR ensemble can correspond to no populated state: an averaged restraint is satisfied by a structure that need not exist.
Precision is not accuracy here either. On the same molecule, upgrading from 108 semi-quantitative distance restraints to 163 restraints including exact ones shrank the structural bundle from 0.95 Å to 0.10 Å backbone root-mean-square deviation — but only ten of the 163 restraints were long-range, so a very tight bundle was being built almost entirely from local geometry, and the authors note that a previously published, even tighter structure of the same molecule was described by its own authors as over-restrained. In the same dataset, the standard threshold criterion for detecting an intramolecular hydrogen bond from the temperature dependence of a chemical shift returned a false negative on a proton known from crystallography to be hydrogen bonded — a conformational average masquerading, under a threshold, as a different structure.
Small-angle X-ray scattering, the method of choice for disordered chains, is subject to the same category of limit stated more starkly by its own practitioners: the majority of scattering curves from disordered regions can be fitted with fewer than five conformations, and the molecular details of those conformers — the contacts producing their size and shape, and the correlations between those contacts — cannot be extracted from the data at all. The only reliably recoverable quantity may be the ensemble-average level of compaction, which simple analytic models already give directly (Martin et al., 2021). The conformers drawn in a published ensemble figure are a compatibility statement, not a census.
Prediction, dated
This subsection carries a date on purpose, because this axis moves faster than any other in the document and every capability claim below expires. The rule that governs its grammar is the one this document uses throughout: a prediction is not a measurement and a docking pose is not a structure. Predictions are reported here as predictions, with the predictor named, and the only sentences below that use the language of measurement are the ones reporting experiments.
Start with chirality, because it is the failure that cannot be waved away. Scored against 33 cyclic-peptide X-ray crystal structures, AlphaFold3 assigned the configuration of D-amino acid residues correctly in 50.3 % of cases, against 75.9 % for one competing model and 100 % for two others (Cao et al., 2025). A second study, on a different test set of 50 non-canonical residues drawn from the standard chemical dictionary, put the same model's chirality accuracy near 64 % and characterised the error precisely: it systematically renders D residues as L, even when the input features carry the correct handedness (Han et al., 2026). That second study rules out the obvious excuse by rebuilding every input through two independent conformer generators and getting the same result, with 0.903 Å all-atom deviation between the two input paths. The two figures come from different test sets with different denominators and should be reported as a range, not a single number — but they agree on direction and severity, and Part Three has already established what a chirality error means: a D-residue predicted as L is a different molecule. In the same 2025 study, the authors report that their own improved model classified every peptide bond in the test set as trans and got no cis bond right — in N-methylated macrocycles, where cis amides are the defining structural feature.
The rest of the picture for peptides specifically is a set of numbers worth holding together. Without an explicit encoding of the ring closure, AlphaFold predicts cyclic peptides as linear chains, at 5.736 Å mean α-carbon deviation from the crystal structures — not a near miss but the wrong topology; adding a cyclic position-offset matrix brings that to 2.291 Å, and fine-tuning on Rosetta-generated structures to 1.368 Å, with 14 of 33 predictions below 1 Å (Cao et al., 2025). Force-field relaxation of the predictions improves local geometry and clash scores without changing global accuracy at all — 1.374 Å both relaxed and unrelaxed — so a model that passes geometry validation says nothing about whether the fold is right. And acceptance is not validity: fed 3,179 non-canonical residue monomers absent from the standard dictionary, AlphaFold3 ran without error on 100 % of them and produced structures satisfying all five geometric criteria for 72.44 %, with good peptide-bond torsions in 85.78 % when embedded in a test peptide (Han et al., 2026). For the broader task of predicting a protein–peptide complex from sequence, the reviews available here put success near 50 %, with the best single method at 53 % — but neither restates the target count or the deviation threshold that defines success, so the figure is quoted here only with that caveat attached (Chang et al., 2024).
Design is where prediction is tested hardest, and the two halves of the result belong in one sentence. De novo designed antibodies were confirmed by cryo-EM for the first time in 2026: a 3.6 Å reconstruction matched its design model at 0.9 Å over the fold, with the six designed loops agreeing at 0.2, 0.2, 0.3, 0.4, 0.7 and 1.1 Å and agreement extending to side-chain rotamers (Bennett et al., 2026). The success rate of the campaigns that produced them was 0 % to 2 % per target, with 9,000 designs screened per target by yeast display. First-round affinities ran from 78 nM to 5.5 μM by surface plasmon resonance, and roughly two orders of magnitude of laboratory affinity maturation were applied on top. One design bound its intended epitope but through a predominantly framework-mediated interface rather than the designed one, established by cryo-EM at 3.9 Å, and its authors classified it a design failure. Binding is not evidence that the designed mechanism was right.
Docking and energetics carry their own boundaries. One review, and only one source in this evidence base, states the limit: classical docking degrades sharply beyond about five residues, and the reason is that the peptide's rotatable bonds create a conformational space the sampling cannot cover — with the further wrinkle that the ability to generate a near-native pose exceeds the ability to rank it first (Chang et al., 2024). On the energetic side, the sharpest available comparison is internal to a single designed-peptide study: for one complex, an end-point method returned a binding free energy of −32.8 ± 0.1 kcal/mol while a rigorous alchemical calculation on the same complex returned −10.2 ± 2.4 kcal/mol; for a second, −22.3 ± 0.1 against −6.6 ± 3.5 kcal/mol (Thakkar et al., 2023). The tight error bar on the fast method is a measure of how reproducibly it produces its answer, not of how close that answer is. The rigorous calculation predicted 15 μM for the second peptide and biolayer interferometry measured 31 ± 9 μM — agreement that must never be quoted without the accompanying statement, made by the same authors, that their ±3.5 kcal/mol uncertainty admits any dissociation constant between 0.04 and 5,200 μM. Alchemical free-energy methods in routine pharmaceutical use carry errors of 0.5–1 kcal/mol, which at room temperature is roughly a fivefold error in the constant they are predicting.
Molecular dynamics is a sampling method, and its outputs are samples. Three replicate simulations of one peptide–protein complex, started from the same structure at the same temperature, stayed bound for approximately 200, 2,300 and 3,300 nanoseconds — a sixteenfold spread across three trajectories of an identical system, which makes any residence time quoted from a single run uninterpretable (Thakkar et al., 2023). The gap between what simulation reaches and what it is asked about is arithmetic: conventional all-atom simulation is limited to tens of microseconds, and the small protein GB1 folds experimentally in 10 milliseconds (Bačić Toplek et al., 2024). And the force fields themselves are the largest uncontrolled variable. Benchmarked residue by residue against experimental coupling constants in short unfolded peptides, a model fitted directly to the spectroscopy achieved a reduced chi-squared of 4.04 ± 3.34, while three widely used force fields achieved 32.54 ± 26.43, 46.60 ± 34.72 and 31.02 ± 36.76 (Suresh et al., 2026). In one five-residue peptide that experiment shows to be unfolded, two of those force fields folded it — 78 % and 66 % of conformations collapsing into a turn held by a contact between the N-terminal ammonium and an aspartate side chain — while the third kept it extended, in agreement with the data. The peptide had to be excluded from the study's own summary because the result contradicted experiment. None of the three reproduced the effect of neighbouring residues on local structure, which is precisely the variable that distinguishes one therapeutic analogue from another.
A final example shows how much the reported subset matters. Two peptides of identical amino-acid composition differing only in sequence order were measured by fluorine NMR titration: the one with dispersed residues bound the target more tightly (156.8 ± 5.80 μM) than the one with clustered leucines and more helical character (223.1 ± 13.50 μM). Eight independent one-microsecond simulations per peptide gave mean interaction energies that overlapped within error and ranked the two peptides the wrong way round; analysing only the single best-docked pose flipped the ranking to agree with the measurement (Tino et al., 2026). Two defensible analyses of the same trajectories give opposite answers, and only one of them matches the experiment.
The reckoning does not end in scepticism, and it should not. Every method above is doing something no other method can, and several of the results in Parts One through Four exist only because two of them disagreed and someone went looking for the reason. The correction is narrower than a general warning. It is that the grammar of a structural claim carries most of its content: solved at 2.7 Å by X-ray crystallography is a different kind of sentence from predicted with high confidence, which is different again from observed in a one-microsecond trajectory, and a reader who flattens the three has lost the ability to tell which claims can be relied on. Two facts from this section make the point without any interpretation. Three replicate simulations of one complex, started identically, disagreed by a factor of sixteen. And a flagship structure predictor, in 2025 and 2026, got molecular handedness wrong somewhere between a third and a half of the time on the exact chemistry that therapeutic peptides are made of.
Part Four began with the observation that two residues out of thirty can carry an interface, and it ends with the reason that observation was so hard to establish: it is not visible in any structure, at any resolution, by any method. It had to be measured by taking the molecule apart one residue at a time, in the assay whose question was actually being asked, and the answer changed depending on which question that was. What remains is the other half of the peptide's life. A molecule that binds is not yet a molecule that survives, and Part Five starts where this one stops — with the clock that begins running the moment the peptide enters a body that has enzymes.
18The clock starts at injection
Take a 21‑residue peptide with no fold — random coil by circular dichroism in 20 mM sodium phosphate at pH 6.9 — and put a single chymotrypsin‑preferred site in the middle of it. Incubate it in buffered saline with 50 nM purified bovine α‑chymotrypsin. The enzyme cuts exactly one bond, between Tyr11 and Lys12, with a half‑life of 8 min. Now change that one tyrosine to alanine. Under identical conditions the enzyme does not visibly touch the peptide in five days (Werner et al., 2016).
That is the whole of this section stated as a single experiment. A protease does not attack a peptide; it reads a short stretch of sequence and cuts a bond inside it. Which bonds a peptide offers, and to which enzymes, is written into the order of its residues in exactly the way its receptor affinity is. Part Four spent its length showing that two residues out of thirty carry most of the binding. The symmetrical and less comfortable fact is that one or two residues out of thirty also carry most of the destruction, and they are usually not the same residues.
The claim generalises past a single model substrate. Curating experimentally verified cleavage sites from a specificity database for 29 proteases implicated in peptide‑drug degradation and fine‑tuning a protein language model to classify each bond gave F1 scores up to 0.94 for the best‑behaved enzymes — 0.9354 for caspase‑6, 0.7924 for thrombin — from sequence alone, with no structure supplied (Cifuentes et al., 2025). The performance is enzyme‑dependent and sometimes poor: MMP‑3 reached only 0.3077. But the same paper carries a statistic that matters more than any of the scores. Across the 29 datasets, the ratio of non‑cleaved to cleaved bonds ran from 36 to 592, with a median of 188. For any given protease, the overwhelming majority of the bonds in a sequence are not substrates. Susceptibility is sparse, positional and legible — which is what makes it an engineering target rather than a fact of life.
The cleanest rule in the field, and the two drugs built on it
Dipeptidyl peptidase‑4 removes the first two residues from a peptide's N‑terminus, and it does so only when proline or alanine occupies the penultimate position. That is the entire specification. Inspecting the three mouse interferon‑inducible CXCR3 ligands, only CXCL10 has a proline at position 2 — and CXCL10 was the one converted by DPP‑4 into a form that still binds CXCR3 but no longer signals chemotaxis (Matsumoto et al., 2023). Cleavage inactivated the message without removing the molecule, which is worth holding onto: destruction and disappearance are not the same event.
Two marketed peptides exist because that rule can be read off the first two letters of a sequence and then edited. Teduglutide is glucagon‑like peptide‑2 with alanine at position 2 replaced by glycine, which deletes the DPP‑4 site; its terminal elimination half‑life is reported as 2 to 3 h after subcutaneous administration in humans (Zhang et al., 2025). Albiglutide carries the same alanine‑to‑glycine substitution in each of two copies of a modified GLP‑1, fused to human albumin, and has an average plasma half‑life of five days in humans (Nilsen et al., 2025). Native GLP‑1's circulating half‑life is given as one to two minutes, attributed jointly to DPP‑4 in blood and renal clearance (Jang et al., 2026; Nilsen et al., 2025).
The temptation is to multiply those numbers into a fold‑change, and it must be resisted. No document read for this monograph reports a with‑and‑without DPP‑4 pair measured in one experiment with matrix and species stated. The one‑to‑two‑minute figure, the two‑to‑three‑hour figure and the five‑day figure come from three different papers and are three different quantities; albiglutide's five days is driven mainly by the albumin fusion, not by the position‑2 edit. Nor does the corpus contain a Michaelis constant, a turnover number or a specificity constant for DPP‑4 acting on anything.
Mapping the cuts in real molecules
Where a peptide is actually attacked has been mapped in several cases by mass spectrometry of the fragments. The decapeptide 10Panx1, residues Trp74 to Tyr83 of human pannexin‑1, is cut in human plasma at 37 °C at two bonds: Trp74–Arg75 and Ala78–Phe79 (Lamouroux et al., 2023). Leu‑enkephalin's known weak point is the Tyr1–Gly2 bond, hydrolysed by aminopeptidase N with a half‑life of 16 min in Dulbecco's phosphate‑buffered saline at 37 °C with the purified porcine enzyme (Byerly‑Duke et al., 2026). The salivary mucin fragment MUC7(84–96), FPNPHQPPKHPDK, is cut by trypsin at the Lys–His bond, identified by the HPDK fragment at m/z 496.253 (Ślusarczyk et al., 2026).
What blocking a cut is worth, in numbers with their assays attached
The fold‑changes available from backbone chemistry are large, and they are position‑specific to an almost absurd degree. In the chymotrypsin model system, replacing a backbone amide's hydrogen with an amino group — N‑amination — at the P1 position took the half‑life from 2.0 min to 300 min, and at the adjacent P1′ position to 202 min, while the same modification three residues away moved it from 2.0 to 2.5 min (Anwar et al., 2026). Against proteinase K, a deliberately non‑specific protease, a linear α‑peptide had a half‑life of 0.27 min; the same sequence with a hydrocarbon staple, 1.6 min; and the stapled sequence carrying three additional cyclic β‑amino‑acid residues, 150 min (Checco et al., 2015). Four dispersed α‑to‑β3 replacements in a MUC1 glycopeptide moved its proteinase‑K half‑life from 1.0 min to 124 min (Gibadullin et al., 2025). And a triazole staple raised 10Panx1's half‑life in human plasma at 37 °C from 2.27 ± 0.11 to 66.13 ± 0.52 min (Lamouroux et al., 2023).
Notice how heterogeneous that list is. One number is a buffer incubation with a single purified enzyme; two are incubations with an aggressive generic protease; one is human plasma. Only the last describes anything resembling the environment a drug meets. They are all real, and they are not on a common scale.
The result that spoils the story
Here is the finding that should be read before any of the fold‑changes above. The MUC7 fragment's trypsin site was mapped, and then two of its residues — exactly the two flanking the scissile bond — were inverted to their D‑enantiomers. The mapping worked perfectly: after 24 h with trypsin the characteristic HPDK fragment was simply absent. The site was gone. In pooled human plasma at 37 °C, the native peptide had a half‑life of 23 ± 5 min and the two‑D‑residue analogue 16 ± 2 min — no improvement, if anything slightly worse. Only the wholly D‑configured peptide survived, losing about 25 % over 120 min where the others had lost 94 % (Ślusarczyk et al., 2026).
The lesson is not that cleavage‑site mapping is useless. It is that plasma is not one protease. A peptide's lifetime in blood is set by the aggregate of every site it presents to every peptidase present, and the site you can see in a purified digest is merely the one that a chosen enzyme happened to find first. Blocking it removes one term from a sum with many terms. The same arithmetic explains a companion observation from a different laboratory: in 25 % human serum at 37 °C, capping the C‑terminus of two cationic hexapeptides as an amide — the standard countermeasure against carboxypeptidases — produced degradation profiles indistinguishable from the free‑acid parents, and mass spectrometry of the intermediates recovered fragments missing two or three C‑terminal residues but never one, meaning an endoprotease had cut first and the carboxypeptidase was never the entry point at all (Nguyen et al., 2010). Reviews continue to state that C‑terminal amidation confers carboxypeptidase resistance (Voronko et al., 2025; Mao et al., 2026). The rule is not wrong. Its predictive value in a particular matrix, for a particular sequence, is close to zero.
Three things called half‑life
One paper in this reading separates the senses more cleanly than any other. A series of cyclic thrombin inhibitors was profiled on three independent axes. Most withstood 8 h in simulated gastric fluid with porcine pepsin at pH 1.2 and simulated intestinal fluid with porcine pancreatin at pH 6.8 at 37 °C — proteolytically excellent. The same molecules gave parallel‑artificial‑membrane logPapp values from −5.6 ± 0.1 to −5.2 ± 0.1, against warfarin at −5.3 ± 0.1 — permeable. And the most potent of them was cleared by rat liver microsomes at 260 ± 20 μL per minute per milligram of protein, which the authors judged too fast to dose at all. The degradation products carried masses 16 and 32 units above the parent: the molecule was being oxidised on its thioether, and that is not proteolysis (Merz et al., 2024).
Species belongs in the same warning. The valine–citrulline linker, one of the protease-cleavable peptide linkers carried by eight of the fifteen approved antibody–drug conjugates, is stable in human serum and cleaved in mouse plasma by the carboxylesterase Ces1C, confirmed with purified enzyme and a knockout‑mouse control — which is why preclinical work on these molecules has to be run in knockout animals. Putting an acidic residue three positions from the scissile bond fixed it; putting a lysine there made the linker more labile than the parent (Balamkundu et al., 2023).
The kidney, which does not read sequence at all
Proteolysis is only one exit. The other is filtration, and it is indifferent to chemistry: it sorts by size. The threshold most often quoted for glomerular filtration of proteins is approximately 60 to 70 kDa, with the podocyte slit diaphragm identified as the principal size‑selective barrier, the glomerular basement membrane as a secondary filter and the fenestrated endothelium contributing both size and charge selectivity (Heaps et al., 2026); the same figure is restated independently in a review of albumin‑binding engineering (Argyle et al., 2026). The threshold also appears in length units: about 3 nm hydrodynamic radius, proposed as a design target for generative models (Heaps et al., 2026), and about 5.5 nm in a nanomedicine design rationale (Yu Q. et al., 2026).
All three of those numbers must be flagged. Every one of them is a review‑level or design‑rationale restatement carrying a citation, not an original measurement; two of the three come from the same paper; no document read here reports a sieving curve or a measured Stokes radius; and the two length figures use different quantities, one a radius and one a size whose definition is not stated. Present the threshold as the accepted figure, which it is. Do not present it as a result.
What the threshold buys, when crossed, is large. Free insulin‑like growth factors have circulating half‑lives under 10 min; more than 99 % of them circulate in complexes with binding proteins and, for two of those, an acid‑labile subunit, producing a 150 kDa ternary assembly far above the filtration cut‑off and a half‑life of roughly 12 to 16 h (Heaps et al., 2026). Two glomerular filtration barriers are also not the whole renal story: filtered protein that crosses is normally reclaimed in the proximal tubule by megalin–cubilin‑mediated endocytosis, and impaired megalin expression or variants in the cubilin gene raise susceptibility to albuminuria (Rroji et al., 2026); the route is saturable and can be competed with free ligand (Yu Q. et al., 2026).
Neprilysin, the zinc metallopeptidase most often named alongside renal clearance in discussions of circulating peptide hormones, appears in this reading in outline only. It has more than 50 reported substrates, is anchored ectodomain‑out on the plasma membrane, can be shed into the circulation, and prefers to hydrolyse after a valine; the preference was tested by mutating a conserved C‑terminal valine in phospholamban to alanine, which abolished cleavage (Cunningham et al., 2026, preprint). But no paper read here gives a turnover number, a Michaelis constant, or a half‑life change attributable to neprilysin for any peptide, and the one substrate worked out in detail is an intracellular membrane micropeptide rather than a circulating drug. Two further enzymes that any account of peptide clearance would be expected to name are simply absent. A regex sweep across every retrieved record and full text in the gap‑fill retrieval returned zero hits for insulin‑degrading enzyme. Angiotensin‑converting enzyme appears only as a drug target, never as a degrader of therapeutic or endogenous peptides. Nothing in this section should be read as covering them.
What the reader can now predict is this. Given a sequence, the first two residues tell you whether one named enzyme will truncate it; the distribution of lysines, arginines, aromatics and large aliphatics tells you roughly how many other sites it offers; and its mass tells you whether the kidney will remove it regardless. What none of that tells you is a number, because a number requires an assay, and the assay is half the claim.
19What the body does to a sequence after it is made
Human amylin is a 37‑residue hormone whose C‑terminus is an amide rather than the free carboxylate every ribosome produces. Replace the amide with the carboxylate — exchange one nitrogen for one oxygen, at one end of a 37‑residue chain — and receptor activation falls 58‑fold at the human Amylin1 receptor and 20‑fold at Amylin3, but only 2.6‑fold at the calcitonin receptor (Yang et al., 2024, reporting values from a 2018 primary study). By comparison, mutating the C‑terminal tyrosine's entire side chain to phenylalanine or alanine, with the amide left intact, caused only modest reductions, and a proline substitution with the amide intact slightly increased activation.
Read that pair of facts carefully, because they are the argument of this section in miniature. The whole side chain of the last residue matters less than one atom of its terminus. And because the loss is severe at two receptors and mild at a third, de‑amidation does not merely weaken amylin — it changes which receptor amylin prefers. A modification is not a decoration on a sequence. It is an instruction the receptor reads, and this one carries selectivity information.
The same paper measured what that atom is worth for a different property. Non‑amidated amylin still forms amyloid: in thioflavin‑T kinetics at 25 μM in phosphate‑buffered saline at pH 7.4, the time to half‑maximal signal rose slightly more than fourfold at 25 °C and close to sixfold at 37 °C, and the fibrils that formed gave superimposable infrared spectra with amide‑I maxima at 1622 to 1624 cm‑1. Toxicity to a rat insulinoma β‑cell line moved from an EC50 of 62.7 to 65.5 μM for the amidated hormone to 94.4 to 105.6 μM for the free acid — less than twofold (Yang et al., 2024). Fifty‑eight‑fold for signalling; four to sixfold for aggregation kinetics; under twofold for cell death. Function is far more amide‑sensitive than physics is.
Evolution appears to agree. An alignment of 479 unique annotated amylin sequences across jawed vertebrates found the C‑terminal glycine–basic–basic extension — the signal that specifies amidation — strictly conserved apart from three camel species, with the donor glycine itself conserved in every sequence, and the amidating enzyme present in every species in the alignment. The N‑terminal processing signal is less conserved than the C‑terminal one (Yang et al., 2024).
The enzyme that writes the amide
Amidation in animals is performed by one bifunctional enzyme working in two steps. A copper‑dependent peptidylglycine α‑hydroxylating monooxygenase domain hydroxylates the α‑carbon of the C‑terminal glycine; a zinc‑ and calcium‑dependent peptidyl‑α‑amidating lyase domain then cleaves the resulting carbinolamide into the peptide amide plus glyoxylate (Blackburn, 2025). The amide nitrogen is the glycine's nitrogen. The rest of the glycine leaves. One source in this evidence base states, without qualification, that this is the only enzyme in humans capable of the reaction (Blackburn, 2025).
Its chemistry is unsettled in an interesting way. The monooxygenase domain carries two copper centres, one coordinated by three histidines and the other by two histidines and a methionine, separated in the resting oxidised enzyme by roughly 11 to 14 Å across a solvent‑filled cleft. The textbook mechanism has an electron travel that distance. An alternative reading, argued from structures of this enzyme and its sister dopamine β‑monooxygenase showing fully open and closed conformations with the coppers 4 to 5 Å apart, and from X‑ray absorption spectroscopy on a selenium‑containing substrate analogue giving a selenium–copper occupancy near 1.8 — that is, a bridge between two coppers — is that the domain closes instead (Welch et al., 2025). The computed cost of the open‑to‑closed motion is about 2 kcal/mol.
Amidation is also not run to completion in the body. Adrenomedullin circulates predominantly as its inactive glycine‑extended precursor, at ratios reported between 5.6 to 1 and 2 to 1 over the bioactive amide in healthy people (Ilina et al., 2025). A pool of one‑atom‑short, inert hormone persists in blood alongside the active material.
The other terminus has its own cap. N‑terminal pyroglutamate is installed enzymatically by glutaminyl cyclase, and searching human and mouse plasma for peptides carrying both caps re‑discovered gonadotropin‑releasing hormone and gastrin and turned up twelve further capped peptides mapping to precursor proteins with signalling annotation, six of them detected in both species with complete sequence conservation. One of them, a capped fragment of the tachykinin precursor, agonised the human NK1 receptor in a β‑arrestin recruitment assay with an EC50 of 0.7 nM against 1.7 nM for substance P, and was roughly twofold more potent than the sequence‑identical peptide whose N‑terminus had not been cyclised (Wiggenhorn et al., 2023). The cap is worth measurable potency by itself.
Sulfation, and one atom of a different kind
Tyrosine O‑sulfation is written in mammals by two Golgi enzymes, TPST1 of 370 residues and TPST2 of 377, which transfer a sulfonyl group from a nucleotide donor to the phenolic oxygen of a tyrosine and prefer acidic sequence context within about five positions either side (Jin et al., 2026; D'Antona et al., 2024). The human TPST2 catalytic domain has been solved with the spent nucleotide and sodium at 1.75 Å and with manganese at 2.00 Å, the metal sites confirmed by anomalous scattering at the manganese absorption edge; the metal orders the entrance to the active site rather than rearranging it (Jin et al., 2026).
What one sulfate is worth was measured on a therapeutic antibody. An anti‑interleukin‑4 human IgG1 expressed in HEK293 cells resolved into a main species and an acidic species carrying an extra 79.9572 Da. By surface plasmon resonance the sulfated species bound with a dissociation constant of 32 pM against 72 pM for the unmodified one, the difference carried by a faster association rate; in a cell‑based signalling neutralisation assay its IC50 was 0.055 nM against 0.196 nM. Chlorate treatment suppressed the acidic species and co‑transfecting both sulfotransferases enriched it to 92 % of the total, and the modified residue was localised to the first complementarity‑determining region of the light chain (D'Antona et al., 2024). One sulfate on one tyrosine, inside a binding loop, worth roughly twofold in affinity and 3.6‑fold in potency. In the tick anticoagulant madanin, sulfation of two tyrosines is reported to enhance the interaction with thrombin about 1000‑fold, and a 1.55 Å structure of the tick sulfotransferase with a substrate peptide shows why the two sites are modified in a fixed order: a singly‑sulfated peptide binds the enzyme about two‑ to 2.7‑fold more tightly than its unsulfated counterpart, so the first sulfate recruits the second (Yoshimura et al., 2024).
One gene, two molecules
The clearest demonstration that a mature peptide is not its gene comes from the human cathelicidin. A single four‑exon gene on chromosome 3 encodes a 170‑residue precursor from which a 37‑residue peptide of net charge +6 is cut. Mass spectrometry shows that the peptide made by macrophages is largely unmodified while the peptide made by neutrophils is extensively modified at its N‑terminus, and the formylated and acetylated forms lose the ability to induce autophagy that the unmodified form has — which is why the neutrophil‑derived peptide cannot do it (Voronko et al., 2025). One gene, one sequence, two cells, two molecules, two behaviours. A third modification of the same peptide, citrullination by peptidylarginine deiminases, disrupts its helix and abolishes its antimicrobial activity outright — while the N‑terminal acetyl and formyl groups leave both intact.
Not every modification does what it is credited with. Adding a single N‑acetylgalactosamine to a threonine inside the MUC1 tandem repeat raised the affinity of an anti‑MUC1 antibody measured by surface plasmon resonance from 6.50 ± 0.73 μM to 1.30 ± 0.51 μM, a fivefold gain. In a proteinase‑K digest, the glycosylated and non‑glycosylated peptides had half‑lives of 1.1 and 1.5 min — indistinguishable (Gibadullin et al., 2025). Glycosylation is routinely credited with both recognition and stability. Here it delivered one and none of the other.
And a modification can be invisible to the method that would most like to see it. Capping a research peptide's C‑terminus as an amide abolished the binding of two of four antibodies raised against the free‑acid peptide, because those two grip the carboxylate directly — which also explains why they can never recognise the intact protein, where that carboxylate does not exist (Ishii et al., 2021). Conversely, in a 1.59 Å crystal structure of an acylated dual‑incretin analogue, neither the amidated C‑terminal residues nor the entire fatty‑diacid conjugate could be modelled at all: the isotropic displacement parameters climb from an average of 21.8 Å2 over the ordered chain to 71.2 Å2 at the last residue with density (Mitchell et al., 2026). Absence of density is not absence of the group. It is the ensemble, showing up as a hole in the map.
Two final cases close the section. Attaching a sixteen‑carbon fatty acid to lysine 26 of GLP‑1 — the modification that defines liraglutide — is reported to take the half‑life from one to two minutes to more than ten hours by borrowing albumin, a protein that itself survives about twenty days in blood (Jang et al., 2026). That single design move underwrites an entire drug class, and it is the bridge into Section 21. And the lantibiotic epidermin, a finished 21‑residue antibacterial peptide, is assembled by an ordinary ribosome from ordinary amino acids and then converted enzymatically into a molecule containing dehydroalanine, 2‑aminoisobutyric acid, two meso‑lanthionines, a 3‑methyllanthionine and a C‑terminal aminovinyl ring (Chevrollier et al., 2026). No codon spells any of those. The gene and the molecule are not the same chemical object, and for a great many biologically active peptides they never were.
20Buying rigidity
The linear decapeptide 10Panx1 is a random coil in phosphate‑buffered saline by circular dichroism and has a half‑life in human plasma at 37 °C of 2.27 ± 0.11 min. Tether two of its side chains with a triazole staple across an i, i+4 spacing and the helicity of the best analogues rises to 41 and 56 % and the half‑life to 66.13 ± 0.52 and 62.42 ± 2.51 min — a thirtyfold gain in the same assay. Mass spectrometry of the metabolites shows precisely why: the two bonds cut in the linear parent are enclosed by the macrocycle in the stapled analogues, and the bonds that are now cut are the ones immediately outside the tether (Lamouroux et al., 2023). Constraint protects what it encloses and nothing else.
The functional gain from that same thirtyfold stability gain was about twofold. At 100 μM in a mouse melanoma line, the linear peptide reduced hypo‑osmotic‑shock‑induced ATP release by 19 % and the two best stapled analogues by 39 and 33 %. Both the parent and the staples then did something a data sheet would not advertise: across 400 to 6.25 μM they inhibited more at lower concentrations, an inverted concentration–response that the authors could not explain mechanistically and that complicates every potency comparison in the paper.
Two matched series, with the numbers on both sides
| Matched series | Linear parent | One constraint | Two constraints |
|---|---|---|---|
| 10Panx1 — half-life, human plasma, 37 °C | 2.27 ± 0.11 min | 66.13 ± 0.52 min | ~20 % intact at 24 h |
| 10Panx1 — helicity by CD in PBS | random coil | 41–56 % | 44 % |
| 10Panx1 — ATP-release inhibition at 100 μM | 19 % | 33–39 % | 48 %, but insoluble without a tail |
| RbAp48 — IC50, competitive fluorescence polarisation | 2,621 ± 786 nM | 47.7 ± 12.5 nM | 12.3 ± 2.0 nM |
| RbAp48 — half-life in breast-cancer cell lysate | 16.0 min | 11.9 min | 94.3 min |
| RbAp48 — ITC binding free energy | −11.31 ± 0.06 kcal/mol | no binding curve obtained | −11.08 ± 0.44 kcal/mol |
The RbAp48 series is the most complete matched set in this reading, because one laboratory measured affinity, calorimetry and proteolysis on the same molecules. Truncating a fragment of a chromatin‑remodelling partner to fourteen residues cost most of its potency, giving an IC50 of 2,621 ± 786 nM. A single side‑chain‑to‑side‑chain macrocycle restored it to 47.7 ± 12.5 nM — a 55‑fold gain that the authors explicitly attribute to reducing the entropic penalty of binding rather than to adding contacts. Bicyclisation took it to 12.3 ± 2.0 nM ('t Hart et al., 2021).
The entropy argument, quantified — and what it actually says
This is where the section stops being a success story. Isothermal titration calorimetry on the same RbAp48 molecules gives the whole ledger. The long linear peptide bound with a free energy of −11.31 ± 0.06 kcal/mol, composed of an enthalpy of −11.03 ± 0.13 and an entropic term of −0.28 ± 0.19. The bicyclic peptide bound with a free energy of −11.08 ± 0.44, composed of an enthalpy of −7.10 ± 0.02 and an entropic term of −3.98 ± 0.42 ('t Hart et al., 2021). Constraint moved roughly 3.9 kcal/mol out of enthalpy and roughly 3.7 kcal/mol into entropy. The total is the same to within the error bars. And circular dichroism showed the linear, monocyclic and bicyclic peptides were all random coil in solution, so whatever the entropic gain is, it is not helix pre‑formation.
An entirely independent series says the same thing. Stapling an all‑D antagonist of the p53–Mdm2 interaction in six different ways raised helicity from about 20 % to between 24.7 and 46.7 %, and swung the binding enthalpy from −15.8 to −8.25 kcal/mol — while the free energy stayed at about −11 kcal/mol across the whole series (Kannan et al., 2020). Two laboratories, two targets, two chemistries, one answer: pre‑organisation reliably changes where the binding energy comes from and unreliably changes how much of it there is.
That is not an argument against constraint. The 55‑fold and 213‑fold IC50 gains in the RbAp48 series are real, and the thirtyfold plasma stability gain in 10Panx1 is real. It is an argument against the causal story usually told about constraint, in which pre‑paying conformational entropy adds binding energy. Measured honestly, and in the only two places in this reading where it has been measured at all, it does not.
The costs, which papers report quietly
Every published series of constrained analogues contains failures, and they are more instructive than the leads. In the stapled Mdm2 series, one staple position — the one that produced the most helical single‑stapled analogue in the set, at 46.7 % — lost 180‑fold of affinity, from 41.4 nM for the linear parent to 7,540 nM. The reason was traceable: the staple's α‑methyl group sits where a backbone hydrogen bond to a specific glutamine of the target forms, and that bond is lost. The crystal structure of a doubly stapled member of the same series shows both staples fully solvent‑exposed, contacting the protein nowhere — pure conformational tax (Kannan et al., 2020).
Constraint can also make a peptide worse at surviving. In the chymotrypsin scan, two β3 residues placed six α‑residues apart gave a peptide less stable than any single‑substituted analogue and more susceptible than the wholly natural parent (Werner et al., 2016). In the RbAp48 series, the monocyclic peptide had a cell‑lysate half‑life of 11.9 min against 16.0 min for its linear parent ('t Hart et al., 2021). Cyclising an allosteric activator peptide bought an eightyfold potency gain and about 30 % greater resistance to proteolysis in mouse plasma — an unusually lopsided ledger, and it is the potency number that gets quoted (Tokodai et al., 2025, preprint).
Sometimes the gain is real and simply does not reach the endpoint. Ten constrained peptides grafting an antibody's third heavy‑chain loop onto small brominated templates — eight single macrocycles and two bicycles — gave one member that bound group 1 influenza haemagglutinins 8‑ to 45‑fold better than the linear peptide, and, in the authors' own words, with no improvement in neutralisation, which stayed at EC50 values from 59 to over 100 μM. What eventually rescued the series was not more constraint but two non‑canonical side chains, which improved neutralisation 156‑ to 190‑fold. The same paper asserts that its constrained peptides are more protease‑resistant, and did not measure it (Kadam et al., 2026).
And sometimes the reported activity is not the activity at all. In the stapled Mdm2 series, three analogues activated a p53 reporter in a human colorectal line with EC50 values of 13.7 to 30.3 μM — and two of them released lactate dehydrogenase at 18.5 to 24 μM, essentially the same concentrations. A stapled peptide that did not bind Mdm2 at all, at 7,540 nM, still lysed cells with an EC50 of 17.7 μM and scored positive in a p53‑independent counterscreen. Of the whole series, one compound survived all three controls (Kannan et al., 2020). Any stapled‑peptide cellular result reported without a lysis control and a counterscreen is uninterpretable.
What it costs to make one
Synthetic accessibility is not a footnote to this field; it decides which constraint chemistries anyone tries. Commercial synthetic peptide procurement runs at roughly US$10 per amino acid for 90 %‑pure material, described in one methods paper as cost‑prohibitive for academic screening of many stapling architectures; the recombinant workaround the same group developed costs about £20 for the oligonucleotide pair encoding a 30‑residue peptide and yields around 10 nmol from a 10 mL culture — but cannot install a single non‑natural residue, which removes D‑amino acids, N‑methylation, α,α‑dialkylation and hydrocarbon staples, that is, most of the toolkit (Pantelejevs et al., 2023).
Yields are the other tax. Macrocyclisation to a hexameric pyrrolinone ring ran at about 12 % (Smith et al., 2011). A recently developed tetrazine macrocyclisation gives 71 to 97 % isolated yields on unprotected peptides — but the reagent window is one methylene wide: the chloromethyl compound degraded and gave 10 % of the desired product, the chloroethyl compound gave 81 %, and the chloropropyl compound gave none at all (Liu et al., 2026). And in the 10Panx1 series, one designed staple could not be made cleanly at all: it gave a persistent dimer on resin and in solution and was never tested (Lamouroux et al., 2023). Failures of that kind do not appear as data points anywhere.
Oral bioavailability, the prize most often invoked for macrocyclisation, has been achieved in this reading exactly five times, in three species, and never in a human. A de novo thrombin inhibitor reached 4.5 ± 2.5 % and then 18.4 ± 1.9 % in rats after two rounds of fixing oxidative metabolism — each round deliberately trading affinity for exposure, sixfold and then a further 1.3-fold — against a non‑binding designed reference hexapeptide at 27 ± 4 % in the same hands, and against the same group's earlier phage‑derived nine‑residue macrocycle at 0.2 % in mice (Merz et al., 2024). A monopyrrolinone scaffold reached about 13 % oral bioavailability in two dogs (Smith et al., 2011). Across twenty optimised analogues in the thrombin campaign, no single molecule performed adequately in potency, gut stability, permeability and metabolic stability at once. The authors say so themselves.
Two disagreements should be recorded before leaving the section, because both undercut a standard teaching. Whether macrocyclisation improves permeability depends on which macrocyclisation: end‑to‑end backbone cyclisation raised passive diffusion in all eleven matched linear–cyclic pairs of a depsipeptide series (Thorpe et al., 2025), whereas side‑chain‑to‑side‑chain stapling of small peptide–steroid conjugates gave negligible enhancement in a cited comparison — and the distinction, that one removes the chain termini and their polarity while the other leaves the backbone donors intact, is itself the finding. And the textbook mechanism for macrocycle permeability may be wrong: across 1,273 cyclic peptides with measured cell‑monolayer or artificial‑membrane permeability, the most permeable ones had modestly fewer intramolecular hydrogen bonds, with a weak negative correlation after conformational search in a membrane‑mimicking solvent, and the authors argue for global reduction in polarity rather than transannular shielding (Sun et al., 2026). Cyclosporin's conformational shape‑shifting is taught as the paradigm; at dataset scale it does not generalise.
One last figure of a different kind. Of 8,751 catalogued cyclic peptides, 62.5 % derive from natural sources. No stapled peptide has been approved; the flagship stapled p53–Mdm2 antagonist remains in phase 1b/2 more than a decade on; and the one macrocycle in this reading to have completed phase 3 came from messenger‑RNA display, not from structure‑based stapling (Salvi et al., 2026). After thirty years of deliberate constraint engineering, most clinically used cyclic peptides are still natural products or their derivatives.
21Buying time
Stability and duration are different problems, and the second one is not solved by solving the first. A peptide can be made completely protease‑proof — the all‑D peptides of Section 20 were more than 90 % intact after 4 h in human plasma (Kannan et al., 2020) — and still disappear from circulation in minutes, because the kidney does not need a protease. Below roughly 60 to 70 kDa the glomerulus filters, and almost every therapeutic peptide is far below it. Duration therefore has to be bought by adding size or by borrowing it.
Borrowing is the strategy with a marketed proof. Serum albumin is 66.5 kDa, circulates at 35 to 50 g/L, is filtered to the extent of less than 0.001 %, and has a half‑life of about 19 to 21 days, maintained by pH‑dependent recycling through the neonatal Fc receptor in acidified endosomes (Heaps et al., 2026; Argyle et al., 2026). Fatty acids from ten to eighteen carbons bind seven sites on it. A cargo that binds albumin inherits both of albumin's escapes — from the filter and from the lysosome — and that is the whole mechanism behind liraglutide's C16 chain on lysine 26 and its reported move from one to two minutes to more than ten hours (Jang et al., 2026). Semaglutide's design is on the record as a separate paper (Lau et al., 2015); no pharmacokinetic data from it was read for this monograph. Adding size directly works too: fusing an albumin‑binding domain to a 10 kDa scaffold protein raised its half‑life in mice from 40 min to 60 h (Argyle et al., 2026), and unstructured polypeptide extensions have taken a GLP‑1 analogue from about 2 h to more than 100 h in rodents and a fused leptin from about 0.5 h to about 20 h (Heaps et al., 2026).
That is as far as the evidence read for this document goes, and it is not far enough for the section this outline originally asked for. There is no primary data here on PEGylation: no matched pair with a half‑life gain and a receptor‑potency loss, no molecular‑weight series, nothing but a passing list of half‑life‑extension modalities in a single review. There is no primary data on fatty‑acid acylation either — liraglutide and semaglutide are named as clinical validation, and the specific cost that any albumin‑binding strategy incurs, that a fraction of the drug is sequestered at any moment and therefore not available to its receptor, is nowhere quantified. Fc fusion appears in this reading as a name and nothing else. A quantitative treatment of half‑life extension cannot be written from this corpus, and this section will not pretend otherwise.
What can be said is where the two engineering problems collide. Everything that buys duration adds mass, and mass is made of residues. Linkers longer than about twenty residues are noted to raise molecular weight, aggregation propensity and the number of proteolytic cleavage sites (Argyle et al., 2026) — and the observation is not theoretical. The doubly stapled 10Panx1 analogue needed an appended solubilising tail to be usable at all; after 24 h in human plasma at 37 °C the bicyclic core was still intact and the only metabolite detected was the tail, clipped off (Lamouroux et al., 2023). Constraint protects what it encloses. Everything bolted on outside the ring to buy time, solubility or targeting is, by construction, outside the protection.
22What the immune system reads
Begin with a claim that is repeated so often it has stopped being checked. At least six of the documents read for this monograph state that peptide therapeutics have low, or lower, immunogenicity than antibodies and protein biologics. Not one of them supports the statement with an incidence figure (Al Khzem et al., 2026; Jang et al., 2026; Ding et al., 2025; Nguyen et al., 2026; Wekalao et al., 2026). It is a genre convention: a sentence that appears in introductions, cites other introductions, and has no denominator anywhere behind it.
The one clinical number available cuts the other way. Tirzepatide, a dual incretin‑receptor agonist built from endogenous human sequence, induced anti‑drug antibodies in 51.1 % of patients in phase 3 trials (Jang et al., 2026). That figure must be handled with care in both directions. It is reported second‑hand, and the source names no assay — not the screening format, not the confirmatory step, not the neutralising‑antibody assay, not the drug‑tolerance limit, not the sampling schedule, not the denominator. A 51.1 % detection rate says the assay was sensitive. It does not say the antibodies mattered, and the source reports neither a neutralising fraction nor any effect on exposure or efficacy. It cannot be set beside an incidence figure from another drug measured by another assay. The honest summary is that anti‑drug antibodies against a marketed peptide are common enough to be detected in half of a trial population, and that this corpus cannot say what follows from that. A systematic review of twelve studies on modified antimicrobial peptides developed as anticancer agents makes the general point unimprovably: none of the twelve measured antibody induction, complement activation or cytokine release at all (Wekalao et al., 2026).
Short peptides are not too small to be seen
The mechanistic defence of the low‑immunogenicity claim is that a short linear peptide is close to the worst case for B‑cell recognition: most B‑cell epitopes are conformational, built from segments discontinuous in sequence, and the minority that are linear span roughly 8 to 15 residues and yield lower‑affinity antibodies because the interface is smaller and the chain is flexible — an entropic penalty (Parkkinen et al., 2026). That defence is correct as far as it goes, and the same paper shows exactly how far.
Native mass spectrometry of a human IgE antibody fragment isolated from a peanut‑allergic child against free short peptides gave a dissociation constant of approximately 400 μM for a pentapeptide; the same pentapeptide with its proline hydroxylated bound at approximately 140 μM. The entire difference is one oxygen atom. A tetrapeptide from a different allergen bound weakly at approximately 1300 μM (Parkkinen et al., 2026). Micromolar affinity is roughly a thousandfold weaker than a matured antibody–protein interaction, and binding is not the same as immunising. But five residues is enough to be seen specifically by a human antibody, and that is not what "too small to be immunogenic" would predict.
The structure is more striking than the affinities. A 3.2 Å crystal structure of the same antibody fragment with its whole allergen showed electron density for a 26‑residue flexible loop and for nothing else: the protein's own helical, four‑disulfide core was disordered and contributed no contact at all. That loop carries three copies of a six‑residue motif, and each copy captured one antibody fragment, cross‑linking three antibodies on a single molecule at spacings of 1.5 and 2.3 nm, with individual interfaces of 820, 460 and 510 Å2 (Parkkinen et al., 2026). Because cross‑linking needs two motifs, the shortest peptide that could trigger the effector event is calculated at 13 residues, and a 15‑mer containing both motifs was reported to cause measurable mediator release with human serum. The unit of immunological consequence is not the epitope. It is the repeat. And a class I major‑histocompatibility ligand is 8 to 15 residues long (Arosa et al., 2026) — which is to say, exactly the size of a therapeutic peptide. Length does not exempt a peptide from immune recognition. It determines which arm of the system could see it.
The claim that is asserted everywhere and measured nowhere
The standard account of unwanted immunogenicity in biologics puts aggregates at the centre: aggregated and misfolded species formed during synthesis, formulation or storage are read as danger signals, activate antigen‑presenting cells and raise immunogenicity (Ding et al., 2025). Regulatory framing says the same in hedged language — aggregation is a critical quality attribute because of its "potential impact" on immunogenicity (Fedorko et al., 2026), and leachable silicone oil nucleates particles that "have the potential to increase immunogenicity" (Tripathi et al., 2026).
Not one document read for this monograph measures it. There is no dose–response. There is no correlation between aggregate content or subvisible particle counts and anti‑drug‑antibody titre. There is no sequence‑matched monomer‑versus‑aggregate immunisation in a therapeutic setting. A targeted re‑retrieval designed specifically to close this gap did not close it. The nearest measurement in the whole reading is a vaccine study, which is the opposite problem: an influenza receptor‑binding domain of 27.4 kDa was driven by pH into three defined states — predominantly monomeric at pH 4.7, oligomeric at pH 6.0 and large aggregates at pH 7.4, characterised by light scattering, circular dichroism and tryptophan fluorescence — and immunised without adjuvant into two strains of mice. The oligomeric form gave the highest IgG titres; monomer and aggregate both gave weaker antibody responses; and neutralisation potency ranked the three states in a different order again (Tu et al., 2026). Aggregation did not maximise the antibody response. If a single experiment cannot support the general claim, it certainly cannot support this general claim.
This is worth stating flatly rather than softening. The relationship between aggregate content and immunogenicity is the most confidently repeated causal claim in peptide and protein developability, and the evidence base assembled for this document contains no measurement of it. That is a statement about this corpus, not a statement that the relationship is false. But a reader who has been told the claim ten times is entitled to know that they have been told it ten times and shown it none.
Two conflicts, reported as conflicts
What this document cannot tell you about developability
The chemical liabilities are named at residue level and no further. Oxidation at methionine, cysteine and histidine; deamidation at asparagine; cyclisation to diketopiperazine (Jang et al., 2026; Ding et al., 2025). No document read here names the flanking‑sequence motifs that actually govern those reactions, so this monograph does not print them. There is no accelerated stability study anywhere in the reading — no temperature, no duration, no percentage degraded. There is not one solubility limit in milligrams per millilitre for any therapeutic peptide, with buffer and pH; the qualitative statement that solubility collapses near the isoelectric point is present, the number is not. No conventional formulation excipient is named in any document read, although "careful excipient selection" is recommended. And for the prediction tools on which epitope‑engineering rests — the MHC‑binding and T‑cell‑epitope predictors used without comment in one of the papers read (Puerta‑González et al., 2026) — not a single sensitivity, specificity, area under the curve or positive predictive value is reported anywhere in this corpus, and no prediction in it is checked against an experimental T‑cell assay.
Two smaller things are worth carrying forward. Immunogenicity may be a property of the delivery context and the molecule together rather than of the sequence alone: cyclotides used as inert scaffolds are described as weakly immunogenic or silent, and the same scaffolds become immunogenic when deliberately targeted to antigen‑presenting cells (Cândido et al., 2026). And not every immune hazard is an antibody: strongly cationic sequences, the property that makes cell‑penetrating tags work, are separately associated with mast‑cell degranulation and histamine release (Nguyen et al., 2026). A sequence property that Section 11 treated as a mechanism of selectivity is, in a different assay, a toxicity.
23Where the ensemble picture stops helping
This document has spent five Parts arguing one sentence: a sequence is not a blueprint, it is a set of constraints on a population of shapes. The sentence has earned its keep. It explains why a peptide that looks helical in a crystal behaves like a coil in water, why disorder is a functional state rather than a failed experiment, why binding free energy is a small difference between large numbers, why two residues out of thirty carry an interface, and why a protease and a receptor are reading the same object for different reasons. It is now obliged to say what it cannot do.
It predicts that constraint should work. It does not say which bond to make.
The ensemble view licenses a general expectation: narrow the population in advance and you pay conformational entropy before binding rather than during it. Section 20 shows what happens when someone measures the result. In the only two calorimetric datasets in this reading, constraint moved binding energy from enthalpy into entropy and left the total unchanged to within error — −11.31 against −11.08 kcal/mol in one series ('t Hart et al., 2021), and a flat −11 kcal/mol across an entire stapled series whose enthalpy swung by 7.5 kcal/mol (Kannan et al., 2020). The theory is not refuted by that; the promotional version of it is.
Worse for the practitioner, the theory has nothing to say about position, and position is everything. Cyclising the same peptide across an alternative residue pair, four positions away, gave IC50 values between 126 and 2,342 nM instead of 47.7 nM, and the whole strategy failed to transfer to a nearly identical homologous binding site in the same protein ('t Hart et al., 2021). Adding or removing a single glycine from the N‑terminus of a constrained binder moved its dissociation constant from 427 to 53 to 3,000 nM (Liu et al., 2026). One staple position out of six cost 180‑fold of affinity, and it was the most helical one in the series (Kannan et al., 2020). An ensemble argument tells you that narrowing a population is worth something. It does not tell you where to put the tether, and in this reading that question has only ever been answered by making all of them.
It explains why prediction is hard. It does not fix it.
Part Four ended on the limits of prediction, and Part Five makes them concrete. When de novo design of antibody binders was validated experimentally for the first time, the reported success rates against each target ran from 0 to 2 % — nine thousand designs screened per target to find a handful of binders (Bennett et al., 2026). The same paper's cryo‑electron‑microscopy confirmation matched the designed loops to 0.2 to 1.1 Å. Both facts are true simultaneously: the field can now specify a protein interface atom by atom, and cannot predict which of its specifications will bind.
For peptides specifically the position is worse, because the tools were not built for the chemistry. Structure predictors are reported to lose accuracy on exactly the cyclic peptides carrying N‑methylated and D‑residues that this Part is about; immunogenicity predictors cannot accept a D‑residue at all (Jang et al., 2026); and in one instructive case a predicted antibody model placed a 17‑residue loop in the wrong conformation, with the consequence that the peptide‑binding cavity did not exist in the prediction (Parkkinen et al., 2026). An ensemble picture explains why all of this should be expected — a flexible chain has no single answer to predict, and the populated states depend on solvent, partner and concentration. Explanation is not remedy.
It is silent about the things that actually kill programmes
Nothing in the physics of populations of shapes speaks to whether a molecule can be made, at what yield, at what cost, in what buffer, at what concentration, or how a regulator will read the resulting particle counts. Those are the constraints this reading shows biting hardest. A ring closure at 12 % (Smith et al., 2011); a reagent series in which one methylene decides between 81 % yield and none (Liu et al., 2026); a designed staple that could only ever be isolated as a dimer and was therefore never tested (Lamouroux et al., 2023); roughly US$10 per residue, enough to determine which architectures an academic group can afford to compare (Pantelejevs et al., 2023); not one measured solubility limit in the entire reading; not one named excipient; not one accelerated stability study; and no measurement anywhere of the aggregation–immunogenicity relationship that the same literature treats as settled. After decades of deliberate design, 62.5 % of catalogued cyclic peptides are still natural products or their derivatives, and no stapled peptide has been approved (Salvi et al., 2026).
The position this document ends in
An ensemble view of a peptide is not a design engine. It is a discipline for reading evidence. It tells you that a reported conformation is a claim about a solvent as much as about a molecule; that a half‑life without a species and a matrix is not a fact; that a crystal structure of a flexible chain is a selection from a population and the selection was made partly by the crystal; that an affinity gain and a stability gain are different transactions and neither implies the other; and that any constraint which buys one property has charged for it somewhere, in a currency the paper may not have measured. Every one of those is a question you can put to a result, and each of them has caught something in this Part: a rule that failed in the matrix it was written for, a staple that made a peptide worse, an entropy argument that came out flat, a claim about aggregates that no one has tested, a field‑wide assertion about immunogenicity with a single number behind it and that number unattributed to an assay.
What it will not do is tell a chemist which bond to make on Monday. The sentence at the head of this document says that a sequence constrains a population of shapes. It does not say which member of that population a receptor will select, at what cost, or whether the molecule that does it can be synthesised, dissolved, stored and given to anybody. Those remain empirical questions, answered one compound at a time, at roughly ten dollars a residue. The value of the picture is that it tells you which of your questions are well posed. That is less than the field usually promises and rather more than it usually delivers, and it is where an honest account has to stop.
24Glossary
Terms are defined as this document uses them. Where a definition is conventional rather than natural — as with the peptide/protein boundary — that is said. Where a word is used in two incompatible senses in the literature, both are given and the one used here is named.
| Term | As used here |
|---|---|
| Aib | 2-aminoisobutyric acid: alanine with a second methyl group on the α-carbon. Has no side-chain chirality and sharply restricts the backbone, favouring helical conformations. |
| Alanine scanning | Replacing each residue in turn with alanine and measuring what the substitution costs. Alanine is used because it deletes the side chain beyond the β-carbon without introducing anything new. |
| Amidation | Conversion of a peptide's C-terminal carboxyl to an amide. Removes the terminal negative charge, blocks carboxypeptidases, and is frequently required for receptor binding. |
| Amphipathic | Having separate polar and non-polar faces. In a helix this is a property of the order of residues rather than their composition, because the two faces are created by the 3.6-residue periodicity. |
| Amyloid | A fibrillar assembly in which the polypeptide backbone runs perpendicular to the fibre axis in extended β-strands. Defined by architecture, not by which protein formed it. |
| Conformational ensemble | The population of structures a chain actually occupies, with their relative weights. This document's controlling idea is that this, rather than any single structure, is what a sequence specifies. |
| Conformational selection | Binding in which the partner captures a conformation the free chain already visits. Distinguished from induced fit by its kinetics, not by its endpoint. |
| Cross-β | The amyloid architecture: β-sheets stacked face to face with strands roughly perpendicular to the fibre axis. |
| Cystine knot | A disulfide arrangement in which one bridge threads the macrocycle formed by two others. Among the most thermally and proteolytically stable arrangements available to a short peptide. |
| Developability | The collection of properties — solubility, chemical stability, aggregation propensity, manufacturability — that decide whether a good binder can become a product. |
| Hot spot | A residue or small cluster contributing a disproportionate share of an interface's binding energy. The set of hot-spot residues is generally smaller than the set of contacting residues. |
| Hydrophobic moment | A vector sum of side-chain hydrophobicities around a helical wheel. Large when the polar and non-polar residues segregate onto opposite faces. |
| Induced fit | Binding in which the partner elicits a conformation the free chain does not appreciably populate. |
| Intrinsically disordered | Lacking a single dominant folded structure under physiological conditions. A functional state with its own biophysics, not a failed experiment. |
| Macrocyclisation | Closing a peptide into a ring — head-to-tail, side-chain to side-chain, or side-chain to terminus. Removes free termini and restricts the conformational ensemble. |
| Message–address | The two-part architecture of many peptide hormones: one region binds and holds, another activates. Truncating from one end abolishes activity; from the other it can produce an antagonist. |
| Peptide | A chain of amino acids linked by peptide bonds. The upper boundary against "protein" is conventional, commonly placed near fifty residues, and has no natural basis. |
| Peptide bond | The amide linkage between one residue's carboxyl and the next one's amino group. Its partial double-bond character makes it planar, which is why a chain has two free rotations per residue rather than three. |
| pLDDT | AlphaFold's per-residue confidence score. Low values in a flexible region may indicate genuine disorder rather than a failed prediction, and the two readings are not distinguishable from the score alone. |
| Polyproline II | An extended left-handed backbone conformation with no internal hydrogen bonding. A major component of what used to be called "random coil", and the default state of many short peptides in water. |
| Ramachandran map | The plot of backbone φ against ψ showing which combinations are sterically allowed. Glycine and proline have their own maps. |
| Resolution | In crystallography and cryo-EM, the finest spacing at which features are separable. It decides what a structure can be used for, and a claim about a side-chain contact requires better resolution than a claim about a fold. |
| Retro-inverso | A peptide with the sequence reversed and every residue in the D configuration, intended to present the same side-chain topology on a protease-resistant backbone. Whether it does so is discussed in Section 13. |
| Scissile bond | The peptide bond a given protease cuts. Its position is determined by the surrounding sequence, which is why susceptibility is encoded exactly as activity is. |
| Secondary structure | Local backbone conformation defined by its hydrogen-bonding pattern — helix, sheet, turn. A property of several contiguous residues, not of one. |
| Stapling | Covalently bridging two side chains, classically with a hydrocarbon linker across one or two helical turns, to pre-organise a helix. |
| Statistical coil | The modern replacement for "random coil": an unfolded ensemble whose residues have distinct, sequence-dependent conformational preferences rather than sampling all allowed space equally. |
| Steric zipper | The dry, tightly interdigitated interface between two β-sheets in an amyloid spine. |
| TFE | 2,2,2-trifluoroethanol. A co-solvent that stabilises helical conformations. Helicity measured in TFE is a statement about the peptide in TFE. |
25Abbreviations
| Abbreviation | Expansion |
|---|---|
| Aib | 2-aminoisobutyric acid |
| Å | ångström, 10−10 m |
| ADA | anti-drug antibody |
| CD | circular dichroism |
| cryo-EM | cryogenic electron microscopy |
| DPP-4 | dipeptidyl peptidase 4 |
| GLP-1 | glucagon-like peptide 1 |
| GPCR | G-protein-coupled receptor |
| HDX | hydrogen–deuterium exchange |
| IDP · IDR | intrinsically disordered protein, intrinsically disordered region |
| ITC | isothermal titration calorimetry |
| JATS | Journal Article Tag Suite, the XML format the corpus is stored in |
| MD | molecular dynamics |
| MeSH | Medical Subject Headings, the NLM indexing vocabulary |
| MHC · HLA | major histocompatibility complex; human leukocyte antigen |
| MIC | minimum inhibitory concentration |
| NMR | nuclear magnetic resonance |
| NOE | nuclear Overhauser effect |
| PAM | peptidylglycine α-amidating monooxygenase |
| PDB | Protein Data Bank, and its structure identifier |
| PEG | polyethylene glycol |
| PK | pharmacokinetics |
| pLDDT | predicted local distance difference test — AlphaFold's per-residue confidence |
| PMC · PMCID | PubMed Central, and its article identifier |
| PMID | PubMed identifier |
| pPII | polyproline II conformation |
| PTM | post-translational modification |
| RMSD | root-mean-square deviation |
| RUO | research use only |
| SAR | structure–activity relationship |
| SAXS | small-angle X-ray scattering |
| SPPS | solid-phase peptide synthesis |
| SPR | surface plasmon resonance |
| TFE | 2,2,2-trifluoroethanol |
| ThT | thioflavin T |
26References
Generated from verified NCBI records rather than from recall. Author lists, journal names, volumes, pages and identifiers are taken from the PubMed record for each citation, and identifiers are read only from each record's own identifier list — a walk over every identifier node in a PubMed record also traverses its reference list and returns the identifiers of the last work that paper cited, which is how a monograph in this series once fetched the wrong articles under the right headers. Every entry carries an identifier. Two foundational works that carry none — Merrifield's 1963 solid-phase synthesis and du Vigneaud's 1953 total synthesis of oxytocin, both in a journal PubMed does not index that far back — are described in the text but not cited, because the claims beside them are carried by indexed papers. An uncited work does not belong in a reference list, and the fact that those two have no identifier is recorded in the unresolved-structure appendix where it belongs.
- A Thirumurthy M, Aguilar Díaz de León JS, Pan S, Ly N. In Vitro Characterization of Agonist and Antagonist Peptide Binding Interaction Kinetics to GLP-1R in HEK293T Cells Using Surface Plasmon Resonance Microscopy. ACS Med Chem Lett. 2026;17(5):963-972.
PMID 42157826 · doi:10.1021/acsmedchemlett.6c00091 · PMC13181479 - Abraham S, Zhu C, Le LHS, Alugubelli YR, Nonomura T, Huang Y, et al.. Phage Display Driven Identification and Computational Mapping of Macrocyclic Peptides Targeting RhoA G17V. Biochemistry. 2026;65(9):1465-1477.
PMID 41986244 · doi:10.1021/acs.biochem.6c00058 · PMC13151067 - Al Khzem AH, Gomaa MS. Peptide-Based Therapeutics for Alzheimer's Disease: Medicinal Chemistry, AI-Guided Computational Design, and Blood-Brain Barrier Delivery. Drug Des Devel Ther. 2026;20:597087.
PMID 42007396 · doi:10.2147/DDDT.S597087 · PMC13089476 - Al-Momani SS, Ramsbottom K, Collins A, Boswell E, Hendriks IA, Nielsen ML, et al.. A Landscape Analysis of Human SUMOylation. Mol Cell Proteomics. 2026;25(5):101571.
PMID 42019804 · doi:10.1016/j.mcpro.2026.101571 · PMC13213312 - Anfinsen CB. Principles that govern the folding of protein chains. Science. 1973;181(4096):223-30.
PMID 4124164 · doi:10.1126/science.181.4096.223 - Antfolk D, Ming Q, Manturova A, Goebel EJ, Thompson TB, Luca VC. Molecular mechanism of Activin receptor inhibition by DLK1. Nat Commun. 2025;16(1):5976.
PMID 40593645 · doi:10.1038/s41467-025-60634-3 · PMC12216052 - Anwar AF, Cano-Sampaio N, Del Valle JR. Backbone N-heteroatom substitution as a strategy to enhance peptide proteolytic stability. RSC Chem Biol. 2026;7(6):1042-1047.
PMID 42094782 · doi:10.1039/d6cb00061d · PMC13142275 - Argyle MJ, Chipman DM, Woolley AC, Bundy BC, Della Corte D. Albumin-Binding Domains in Therapeutic Protein Engineering: A Structural and Computational Perspective on Rational Design. SynBio. 2026;4(1).
PMID 42495264 · doi:10.3390/synbio4010005 · PMC13395267 - Arosa FA, Cardoso EM. The Cytoplasmic Domain of MHC Class I Molecules as a Molecular Switch: A Perspective from Short Linear Motifs and Intrinsically Disordered Regions. Biomolecules. 2026;16(7).
PMID 42509859 · doi:10.3390/biom16071067 · PMC13407152 - Arroyo-Urea S, Nazarova AL, Carrión-Antolí Á, Bonifazi A, Battiti FO, Lam JH, et al.. Structure of the dopamine D3 receptor bound to a bitopic agonist reveals a new specificity site in an expanded allosteric pocket. Res Sq. 2023.
PMID 38196573 · doi:10.21203/rs.3.rs-3433207/v1 · PMC10775388 - Babin KM, Kilinc C, Gostynska SE, Dickson A, Pioszak AA. Characterization of the Two-Domain Peptide Binding Mechanism of the Human CGRP Receptor for CGRP and the Ultrahigh Affinity ssCGRP Variant. Biochemistry. 2025;64(8):1770-1787.
PMID 40172014 · doi:10.1021/acs.biochem.4c00812 · PMC12004451 - Badia M, Batlle C, Bolognesi B. Massively parallel quantification of mutational impact on IAPP amyloid formation. Nat Commun. 2026;17(1).
PMID 41844597 · doi:10.1038/s41467-026-70611-z · PMC13144336 - Balamkundu S, Liu CF. Lysosomal-Cleavable Peptide Linkers in Antibody-Drug Conjugates. Biomedicines. 2023;11(11).
PMID 38002080 · doi:10.3390/biomedicines11113080 · PMC10669454 - Bauer V, Schmidtgall B, Gógl G, Dolenc J, Osz J, Nominé Y, et al.. Conformational editing of intrinsically disordered protein by α-methylation. Chem Sci. 2020;12(3):1080-1089.
PMID 34163874 · doi:10.1039/d0sc04482b · PMC8178997 - Bačić Toplek F, Scalone E, Stegani B, Paissoni C, Capelli R, Camilloni C. Multi-eGO: Model Improvements toward the Study of Complex Self-Assembly Processes. J Chem Theory Comput. 2024;20(1):459-468.
PMID 38153340 · doi:10.1021/acs.jctc.3c01182 · PMC10782439 - Bennett NR, Watson JL, Ragotte RJ, Borst AJ, See DL, Weidle C, et al.. Atomically accurate de novo design of antibodies with RFdiffusion. Nature. 2026;649(8095):183-193.
PMID 41193805 · doi:10.1038/s41586-025-09721-5 · PMC12727541 - Benson DR, Deng B, Kashipathy MM, Lovell S, Battaile KP, Cooper A, et al.. The N-terminal intrinsically disordered region of Ncb5or docks with the cytochrome b5 core to form a helical motif that is of ancient origin. Proteins. 2024;92(4):554-566.
PMID 38041394 · doi:10.1002/prot.26647 · PMC10932899 - Berggren A, Bakker M, Fisher H, Skepö M. Conformational flexibility and transient structure of the proline-rich domain in p53. Biophys J. 2026;125(8):1914-1925.
PMID 41832603 · doi:10.1016/j.bpj.2026.03.024 · PMC13351842 - Blackburn NJ. Metal-mediated peptide processing. How copper and iron catalyze diverse peptide modifications such as amidation and crosslinking. RSC Chem Biol. 2025;6(7):1048-1067.
PMID 40520143 · doi:10.1039/d5cb00085h · PMC12164859 - Blower RJ, Barksdale SM, van Hoek ML. Snake Cathelicidin NA-CATH and Smaller Helical Antimicrobial Peptides Are Effective against Burkholderia thailandensis. PLoS Negl Trop Dis. 2015;9(7):e0003862.
PMID 26196513 · doi:10.1371/journal.pntd.0003862 · PMC4510350 - Bousch C, Bérubé F, Babych M, Ongeri S, Bourgault S. Molecular Mechanisms of Islet Amyloid Polypeptide Aggregation: Towards Chemical Strategies to Prevent Amyloid Formation and to Design Non-Aggregating Peptide Therapeutics. Int J Mol Sci. 2026;27(6).
PMID 41898461 · doi:10.3390/ijms27062598 · PMC13026743 - Brahma B, Patra MC, Karri S, Chopra M, Mishra P, De BC, et al.. Diversity, Antimicrobial Action and Structure-Activity Relationship of Buffalo Cathelicidins. PLoS One. 2015;10(12):e0144741.
PMID 26675301 · doi:10.1371/journal.pone.0144741 · PMC4684500 - Briggs DC, Duffy RT, Ateaque S, Maslen S, Nagaraj H, Barde YA, et al.. Discovery of a sulfotyrosine-motif in the human TrkB extracellular domain required for agonist activation. bioRxiv. 2026.
PMID 42239242 · doi:10.64898/2026.05.19.725324 · PMC13228283 - Bugge K, Staby L, Salladini E, Falbe-Hansen RG, Kragelund BB, Skriver K. αα-Hub domains and intrinsically disordered proteins: A decisive combo. J Biol Chem. 2021;296:100226.
PMID 33361159 · doi:10.1074/jbc.REV120.012928 · PMC7948954 - Byerly-Duke J, Bernhard SM, Ibrahim R, Das S, Sharma KK, Panda C, et al.. Amidine isosteric modification tunes proteolytic stability and activity. RSC Chem Biol. 2026.
PMID 42454068 · doi:10.1039/d6cb00138f · PMC13367194 - Cao Z, Cao S, Wang L, Wang Z, Mao Q, Guo J, et al.. HighFold-MeD: a Rosetta distillation model to accelerate structure prediction of cyclic peptides with backbone N-methylation and D-amino acids. J Cheminform. 2025;17(1):167.
PMID 41214817 · doi:10.1186/s13321-025-01111-3 · PMC12604167 - Chang L, Mondal A, Singh B, Martínez-Noa Y, Perez A. Revolutionizing Peptide-Based Drug Discovery: Advances in the Post-AlphaFold Era. Wiley Interdiscip Rev Comput Mol Sci. 2024;14(1).
PMID 38680429 · doi:10.1002/wcms.1693 · PMC11052547 - Checco JW, Lee EF, Evangelista M, Sleebs NJ, Rogers K, Pettikiriarachchi A, et al.. α/β-Peptide Foldamers Targeting Intracellular Protein-Protein Interactions with Activity in Living Cells. J Am Chem Soc. 2015;137(35):11365-75.
PMID 26317395 · doi:10.1021/jacs.5b05896 · PMC4687753 - Chevrollier N, Dougha A, Ye C, Stratmann D, Moroy G, Rey J, et al.. PEP-EDIT: a web server for the 3D generation and interactive editing of complex peptides. Nucleic Acids Res. 2026;54(W1):W287-W294.
PMID 42129607 · doi:10.1093/nar/gkag455 · PMC13355081 - Cifuentes P, Adàlia R, Zamora I. Prediction of peptide cleavage sites using protein language models and graph neural networks. Sci Rep. 2025;15(1):38048.
PMID 41168342 · doi:10.1038/s41598-025-21801-0 · PMC12575701 - Cândido ES, Gasparetto LS, Maximiano MR, Rios TB, Franco OL. Cyclotides from Plants Driving the Next Generation of Antibacterial Agents. Antibiotics (Basel). 2026;15(6).
PMID 42353728 · doi:10.3390/antibiotics15060604 · PMC13295905 - D'Antona AM, Lee JM, Zhang M, Friedman C, He T, Mosyak L, et al.. Tyrosine Sulfation at Antibody Light Chain CDR-1 Increases Binding Affinity and Neutralization Potency to Interleukine-4. Int J Mol Sci. 2024;25(3).
PMID 38339208 · doi:10.3390/ijms25031931 · PMC10855961 - Dagbay KB, Treece E, Streich FC, Jackson JW, Faucette RR, Nikiforov A, et al.. Structural basis of specific inhibition of extracellular activation of pro- or latent myostatin by the monoclonal antibody SRK-015. J Biol Chem. 2020;295(16):5404-5418.
PMID 32075906 · doi:10.1074/jbc.RA119.012293 · PMC7170532 - Dahal A, Subramanian V, Shrestha P, Liu D, Gauthier T, Jois S. Conformationally constrained cyclic grafted peptidomimetics targeting protein-protein interactions. Pept Sci (Hoboken). 2023;115(5).
PMID 38188985 · doi:10.1002/pep2.24328 · PMC10769001 - De Zotti M, Sella L, Bolzonello A, Gabbatore L, Peggion C, Bortolotto A, et al.. Targeted Amino Acid Substitutions in a Trichoderma Peptaibol Confer Activity against Fungal Plant Pathogens and Protect Host Tissues from Botrytis cinerea Infection. Int J Mol Sci. 2020;21(20).
PMID 33053906 · doi:10.3390/ijms21207521 · PMC7589190 - Ding X, Li Y. Engineering Bispecific Peptides for Precision Immunotherapy and Beyond. Int J Mol Sci. 2025;26(20).
PMID 41155373 · doi:10.3390/ijms262010082 · PMC12563712 - Do TU, Kraft EJ, Chappell GF, Parnham S, Berlow RB. Sequence-encoded differences in the conformational ensembles of CITED transcriptional activation domains impact coactivator binding. Protein Sci. 2026;35(7):e70693.
PMID 42332380 · doi:10.1002/pro.70693 · PMC13286877 - DU VIGNEAUD V, RESSLER C, TRIPPETT S. The sequence of amino acids in oxytocin, with a proposal for the structure of oxytocin. J Biol Chem. 1953;205(2):949-57.
PMID 13129273 - Dudek D, Dzień E, Wątły J, Matera-Witkiewicz A, Mikołajczyk A, Hajda A, et al.. Zn(II) binding to pramlintide results in a structural kink, fibril formation and antifungal activity. Sci Rep. 2022;12(1):20543.
PMID 36446825 · doi:10.1038/s41598-022-24968-y · PMC9708664 - Fedorko A, Grynyuk I, Afonso Urich JA. Quality by Design to Mitigate Aggregation: Mechanistic Insights and Analytical Strategies for Biopharmaceutical Manufacturing. Ther Innov Regul Sci. 2026.
PMID 42467318 · doi:10.1007/s43441-026-01023-w - Ferková S, Froehlich U, Nepveu-Traversy MÉ, Murza A, Azad T, Grandbois M, et al.. Comparative Analysis of Cyclization Techniques in Stapled Peptides: Structural Insights into Protein-Protein Interactions in a SARS-CoV-2 Spike RBD/hACE2 Model System. Int J Mol Sci. 2023;25(1).
PMID 38203338 · doi:10.3390/ijms25010166 · PMC10778704 - Gallardo R, Iadanza MG, Xu Y, Heath GR, Foster R, Radford SE, et al.. Fibril structures of diabetes-related amylin variants reveal a basis for surface-templated assembly. Nat Struct Mol Biol. 2020;27(11):1048-1056.
PMID 32929282 · doi:10.1038/s41594-020-0496-3 · PMC7617688 - Ghosh C, Nagpal S, Muñoz V. Molecular simulations integrated with experiments for probing the interaction dynamics and binding mechanisms of intrinsically disordered proteins. Curr Opin Struct Biol. 2024;84:102756.
PMID 38118365 · doi:10.1016/j.sbi.2023.102756 · PMC11242915 - Gibadullin R, Suárez Ó, Lazaris FS, Gutiez N, Atondo E, Araujo-Aris S, et al.. Enhancing Cancer Vaccine Efficacy: Backbone Modification with β‑Amino Acids Alters the Stability and Immunogenicity of MUC1-Derived Glycopeptide Formulations. JACS Au. 2025;5(5):2270-2284.
PMID 40443897 · doi:10.1021/jacsau.5c00224 · PMC12117419 - Glossop HD, Gebretsadik G, Sultana S, Biswas D, Schacht NA, Yennawar NH, et al.. Retro-inversion imparts antimycobacterial specificity to host defense peptides. Nat Commun. 2025;17(1):469.
PMID 41354737 · doi:10.1038/s41467-025-67162-0 · PMC12800093 - Gnanaolivu RD, Hart SN. Enhancing missense variant classification in predicted intrinsically disordered regions. PLoS One. 2026;21(7):e0354365.
PMID 42507643 · doi:10.1371/journal.pone.0354365 · PMC13405113 - Gutte B, Merrifield RB. The total synthesis of an enzyme with ribonuclease A activity. J Am Chem Soc. 1969;91(2):501-2.
PMID 5782505 · doi:10.1021/ja01030a050 - Haile MT, Kaxiras DA, Zhen J, Lee CL, Jiang B, Small-Saunders JL, et al.. Structural basis for host membrane binding and remodeling by invading malaria parasites. Cell. 2026.
PMID 42379167 · doi:10.1016/j.cell.2026.06.012 · PMC13322224 - Han Y, Mei J, Li G, Dai E, Lu H, Zhang C, et al.. HighRes_Builder: improved access and modeling of noncanonical residues for protein structure prediction. Brief Bioinform. 2026;27(3).
PMID 42218720 · doi:10.1093/bib/bbag272 · PMC13222519 - Harake SNA, Btadini S, Qadri AH, Harb F, Faheem I. Beyond the structure-function paradigm: A comprehensive review of intrinsically disordered proteins. Biochem Biophys Rep. 2026;47:102706.
PMID 42437078 · doi:10.1016/j.bbrep.2026.102706 · PMC13355577 - Harikumar KG, Piper SJ, Christopoulos A, Wootten D, Sexton PM, Miller LJ. Impact of secretin receptor homo-dimerization on natural ligand binding. Nat Commun. 2024;15(1):4390.
PMID 38782989 · doi:10.1038/s41467-024-48853-6 · PMC11116414 - Hart P', Hommen P, Noisier A, Krzyzanowski A, Schüler D, Porfetye AT, et al.. Structure Based Design of Bicyclic Peptide Inhibitors of RbAp48. Angew Chem Int Ed Engl. 2021;60(4):1813-1820.
PMID 33022847 · doi:10.1002/anie.202009749 · PMC7894522 - Heaps WP, Packard AE, McCammon KM, Green TP, Talley JP, Bundy BC, et al.. Molecular Survival Strategies Against Kidney Filtration: Implications for Therapeutic Protein Engineering. Biophysica. 2026;6(1).
PMID 42524171 · doi:10.3390/biophysica6010004 · PMC13410728 - Heid LF, Agerschou ED, Orr AA, Kupreichyk T, Schneider W, Wördehoff MM, et al.. Sequence-based identification of amyloidogenic β-hairpins reveals a prostatic acid phosphatase fragment promoting semen amyloid formation. Comput Struct Biotechnol J. 2024;23:417-430.
PMID 38223341 · doi:10.1016/j.csbj.2023.12.023 · PMC10787225 - Hicks A, Escobar CA, Cross TA, Zhou HX. Fuzzy Association of an Intrinsically Disordered Protein with Acidic Membranes. JACS Au. 2021;1(1):66-78.
PMID 33554215 · doi:10.1021/jacsau.0c00039 · PMC7851954 - Ilina Y, Kaufmann P, Press M, Uba TI, Bergmann A. Enhancing Stability and Bioavailability of Peptidylglycine Alpha-Amidating Monooxygenase in Circulation for Clinical Use. Biomolecules. 2025;15(2).
PMID 40001527 · doi:10.3390/biom15020224 · PMC11853079 - Ishii M, Nakakido M, Caaveiro JMM, Kuroda D, Okumura CJ, Maruyama T, et al.. Structural basis for antigen recognition by methylated lysine-specific antibodies. J Biol Chem. 2021;296:100176.
PMID 33303630 · doi:10.1074/jbc.RA120.015996 · PMC7948472 - Jagrosse ML, Baliga UK, Jones CW, Russell JJ, García CI, Najar RA, et al.. Impact of Peptide Sequence on Functional siRNA Delivery and Gene Knockdown with Cyclic Amphipathic Peptide Delivery Agents. Mol Pharm. 2023;20(12):6090-6103.
PMID 37963105 · doi:10.1021/acs.molpharmaceut.3c00455 · PMC10698724 - Jahan I, Kumar SD, Ajish C, Lee CW, Shin SY, Yang S. Rational design of thrombin-derived VFR12 analogs with enhanced antimicrobial, antibiofilm, and anti-inflammatory properties. Sci Rep. 2025;15(1):40876.
PMID 41258221 · doi:10.1038/s41598-025-24789-9 · PMC12630807 - Jang YE, Kwon M, Kwon CW, Kim SG, Hwang JS, George NP, et al.. Integrative Peptide Drug Development: Chemical Engineering, AI-Driven Design, and Cell-Penetrating Peptides. Pharmaceutics. 2026;18(5).
PMID 42198231 · doi:10.3390/pharmaceutics18050537 · PMC13210002 - Jin T, Sass JI, Gao W, Hurst AA, Coley CW, Alexander-Katz A. Polymerized Short Sequences as a Template for Protein Folding and Evolution. Nano Lett. 2026;26(13):4471-4479.
PMID 41879619 · doi:10.1021/acs.nanolett.6c00527 - Juraszek J, Kadam RU, Branduardi D, van Ameijde J, Garg D, Dailly N, et al.. De novo design of D-peptide ligands: Application to influenza virus hemagglutinin. Proc Natl Acad Sci U S A. 2025;122(26):e2426554122.
PMID 40577121 · doi:10.1073/pnas.2426554122 · PMC12232713 - Kadam RU, Juraszek J, Brandenburg B, Garg D, Zhu X, Jongeneelen M, et al.. Small molecule-constrained paratope mimetic bicyclic peptides as potent inhibitors of group 1 and 2 influenza A virus hemagglutinins. Proc Natl Acad Sci U S A. 2026;123(13):e2537533123.
PMID 41875158 · doi:10.1073/pnas.2537533123 · PMC13037862 - Kannan S, Aronica PGA, Ng S, Gek Lian DT, Frosi Y, Chee S, et al.. Macrocyclization of an all-d linear α-helical peptide imparts cellular permeability. Chem Sci. 2020;11(21):5577-5591.
PMID 32874502 · doi:10.1039/c9sc06383h · PMC7441689 - Karlsson E, Paissoni C, Erkelens AM, Tehranizadeh ZA, Sorgenfrei FA, Andersson E, et al.. Mapping the transition state for a binding reaction between ancient intrinsically disordered proteins. J Biol Chem. 2020;295(51):17698-17712.
PMID 33454008 · doi:10.1074/jbc.RA120.015645 · PMC7762952 - Katoh T. EF-P Inhibits Ribosomal α-Hydroxy Acid Incorporation: Strategic tRNA Body Selection for Co-incorporating α-Hydroxy Acids and Nonproteinogenic Amino Acids into Depsipeptides. RNA. 2026.
PMID 42309672 · doi:10.1261/rna.081083.126 - Kawai T, Sun B, Yoshino H, Feng D, Suzuki Y, Fukazawa M, et al.. Structural basis for GLP-1 receptor activation by LY3502970, an orally active nonpeptide agonist. Proc Natl Acad Sci U S A. 2020;117(47):29959-29967.
PMID 33177239 · doi:10.1073/pnas.2014879117 · PMC7703558 - KENDREW JC, BODO G, DINTZIS HM, PARRISH RG, WYCKOFF H, PHILLIPS DC. A three-dimensional model of the myoglobin molecule obtained by x-ray analysis. Nature. 1958;181(4610):662-6.
PMID 13517261 · doi:10.1038/181662a0 - KENDREW JC, DICKERSON RE, STRANDBERG BE, HART RG, DAVIES DR, PHILLIPS DC, et al.. Structure of myoglobin: A three-dimensional Fourier synthesis at 2 A. resolution. Nature. 1960;185(4711):422-7.
PMID 18990802 · doi:10.1038/185422a0 - Kjaer LF, Ielasi FS, Winbolt T, Delaforge E, Tengo M, Nebl S, et al.. Alternative Splicing of a Structured Partner Alters the Folding-Upon-Binding Trajectory of an Intrinsically Disordered Protein. J Am Chem Soc. 2026;148(27):28401-28413.
PMID 42361232 · doi:10.1021/jacs.6c04072 · PMC13426257 - Koga R, Yamamoto M, Kosugi T, Kobayashi N, Sugiki T, Fujiwara T, et al.. Robust folding of a de novo designed ideal protein even with most of the core mutated to valine. Proc Natl Acad Sci U S A. 2020;117(49):31149-31156.
PMID 33229587 · doi:10.1073/pnas.2002120117 · PMC7739874 - Koren G, Meir S, Holschuh L, Mertens HDT, Ehm T, Yahalom N, et al.. Intramolecular structural heterogeneity altered by long-range contacts in an intrinsically disordered protein. Proc Natl Acad Sci U S A. 2023;120(30):e2220180120.
PMID 37459524 · doi:10.1073/pnas.2220180120 · PMC10372579 - Kuo ST, Xi Z, Cong X, Yan X, Russell DH. Dissecting Hidden Liraglutide Oligomerization Pathways via Direct Mass Technology, Electron-Capture Dissociation, and Molecular Dynamics. Anal Chem. 2025;97(25):13465-13473.
PMID 40521838 · doi:10.1021/acs.analchem.5c01851 · PMC12224166 - Kyte J, Doolittle RF. A simple method for displaying the hydropathic character of a protein. J Mol Biol. 1982;157(1):105-32.
PMID 7108955 · doi:10.1016/0022-2836(82)90515-0 - Lamouroux A, Tournier M, Iaculli D, Caufriez A, Rusiecka OM, Martin C, et al.. Structure-Based Design and Synthesis of Stapled 10Panx1 Analogues for Use in Cardiovascular Inflammatory Diseases. J Med Chem. 2023;66(18):13086-13102.
PMID 37703077 · doi:10.1021/acs.jmedchem.3c01116 · PMC10544015 - Lau J, Bloch P, Schäffer L, Pettersson I, Spetzler J, Kofoed J, et al.. Discovery of the Once-Weekly Glucagon-Like Peptide-1 (GLP-1) Analogue Semaglutide. J Med Chem. 2015;58(18):7370-80.
PMID 26308095 · doi:10.1021/acs.jmedchem.5b00726 - Li J, Chen G, Guo Y, Wang H, Li H. Single molecule force spectroscopy reveals the context dependent folding pathway of the C-terminal fragment of Top7. Chem Sci. 2020;12(8):2876-2884.
PMID 34164053 · doi:10.1039/d0sc06344d · PMC8179357 - Liu D, Jiang Y, Ma B, Li L. Structure-based artificial intelligence-aided design of MYC-targeting degradation drugs for cancer therapy. Biochem Biophys Res Commun. 2025;766:151870.
PMID 40288261 · doi:10.1016/j.bbrc.2025.151870 - Liu H, Xu H, Yang L, Zhang D, Wang Y, Shen J, et al.. Tetrazine-Enabled Peptide Macrocyclization and Bioorthogonal Cell Penetration Profiling. JACS Au. 2026;6(7):3903-3913.
PMID 42529336 · doi:10.1021/jacsau.6c00457 · PMC13417186 - Lobo S, Griem L, Shell MS, Shea JE. amyloid-predict and LLPS-predict: Predicting phase separation propensities in the intrinsically disordered proteome. Proc Natl Acad Sci U S A. 2026;123(22):e2531932123.
PMID 42190015 · doi:10.1073/pnas.2531932123 · PMC13229271 - Lohan S, Konshina AG, Mohammed EHM, Helmy NM, Jha SK, Tiwari RK, et al.. Impact of stereochemical replacement on activity and selectivity of membrane-active antibacterial and antifungal cyclic peptides. NPJ Antimicrob Resist. 2025;3(1):56.
PMID 40527928 · doi:10.1038/s44259-025-00121-3 · PMC12174330 - Lucana MC, Lucchi R, Gosselet F, Díaz-Perlas C, Oller-Salvia B. BrainBike peptidomimetic enables efficient transport of proteins across brain endothelium. RSC Chem Biol. 2024;5(1):7-11.
PMID 38179197 · doi:10.1039/d3cb00194f · PMC10763564 - López-Sánchez R, Pantoja-Uceda D, Mompeán M, Laurents DV. PolyProline Predictor: A web server for empirical sequence-based prediction of polyproline II helices. Protein Sci. 2026;35(7):e70675.
PMID 42252518 · doi:10.1002/pro.70675 · PMC13243192 - Madhu P, Benz C, Simonetti L, Winters MJ, Subbanna MS, Kliche J, et al.. An Atlas of Short Linear Motif-Mediated Human Protein-Protein Interactions. bioRxiv. 2026.
PMID 42465437 · doi:10.64898/2026.07.03.735260 · PMC13370473 - Maity B, Moorthy H, Govindaraju T. Intrinsically Disordered Ku Protein-Derived Cell-Penetrating Peptides. ACS Bio Med Chem Au. 2023;3(6):471-479.
PMID 38144254 · doi:10.1021/acsbiomedchemau.3c00032 · PMC10739243 - Mao KL, Fu JX, Chen J, Liao YL, Huang ML, Chen Z, et al.. From marine predator to pharmacology: Conotoxin diversity, discovery, and therapeutic potential. Zool Res. 2026;47(2):378-403.
PMID 41888060 · doi:10.24272/j.issn.2095-8137.2025.553 · PMC13425393 - Martin C, Gimenez LE, Williams SY, Jing Y, Wu Y, Hollanders C, et al.. Structure-Based Design of Melanocortin 4 Receptor Ligands Based on the SHU-9119-hMC4R Cocrystal Structure†. J Med Chem. 2021;64(1):357-369.
PMID 33190475 · doi:10.1021/acs.jmedchem.0c01620 - Martí-Aranda A, Lehner B. Seven complete comparative maps of allosteric mutations in a protein family. Nat Commun. 2026;17(1).
PMID 42014722 · doi:10.1038/s41467-026-71005-x · PMC13284222 - Matsumoto A, Hiroi M, Mori K, Yamamoto N, Ohmori Y. Differential Anti-Tumor Effects of IFN-Inducible Chemokines CXCL9, CXCL10, and CXCL11 on a Mouse Squamous Cell Carcinoma Cell Line. Med Sci (Basel). 2023;11(2).
PMID 37218983 · doi:10.3390/medsci11020031 · PMC10204432 - Meng Q, Han Y, Zhou H, He Z. Protocol for structural analysis of amyloid fibrils using hydrogen-deuterium exchange mass spectrometry. STAR Protoc. 2025;6(4):104247.
PMID 41353746 · doi:10.1016/j.xpro.2025.104247 · PMC12741275 - Merrifield RB. Solid-phase peptide synthesis. Adv Enzymol Relat Areas Mol Biol. 1969;32:221-96.
PMID 4307033 · doi:10.1002/9780470122778.ch6 - Merz ML, Habeshian S, Li B, David JGL, Nielsen AL, Ji X, et al.. De novo development of small cyclic peptides that are orally bioavailable. Nat Chem Biol. 2024;20(5):624-633.
PMID 38155304 · doi:10.1038/s41589-023-01496-y · PMC11062899 - Mitchell HM, Nocek B, Guinn EJ, Heng JYY. Crystallization and 1.6 Å resolution crystal structure of an acylated GLP-1/GIP analogue peptide. Acta Crystallogr F Struct Biol Commun. 2026;82(Pt 4):114-124.
PMID 41841205 · doi:10.1107/S2053230X26001937 · PMC13041627 - Mitchell W, Tamucci JD, Ng EL, Liu S, Birk AV, Szeto HH, et al.. Structure-activity relationships of mitochondria-targeted tetrapeptide pharmacological compounds. Elife. 2022;11.
PMID 35913044 · doi:10.7554/eLife.75531 · PMC9342957 - Mou X, Wang J, Shao Y, Li Y, Kwok CK. RNA G-quadruplex structure recognition by chimeric peptides. Nucleic Acids Res. 2026;54(14).
PMID 42531076 · doi:10.1093/nar/gkag745 · PMC13421873 - Mourenza Á, Llano-Verdeja J, Castañera P, Krishnan R, Vogelaar A, Lorente-Torres B, et al.. Design of a cyclic peptide targeting intracellular Staphylococcus aureus. Mol Biomed. 2026;7(1).
PMID 42525328 · doi:10.1186/s43556-026-00519-z · PMC13421713 - Müller E, Hackney CM, Ellgaard L, Morth JP. High-resolution crystal structure of the Mu8.1 conotoxin from Conus mucronatus. Acta Crystallogr F Struct Biol Commun. 2023;79(Pt 9):240-246.
PMID 37642664 · doi:10.1107/S2053230X23007070 · PMC10478764 - Nguyen HTN, Le BHN, Van NTH, Tran TTT, Nguyen MT. Targeting the Undruggable: Deep Learning-Driven Design of Peptide Therapeutics in Cancer. Pharmaceuticals (Basel). 2026;19(7).
PMID 42515682 · doi:10.3390/ph19070998 · PMC13416385 - Nguyen LT, Chau JK, Perry NA, de Boer L, Zaat SA, Vogel HJ. Serum stabilities of short tryptophan- and arginine-rich antimicrobial peptide analogs. PLoS One. 2010;5(9).
PMID 20844765 · doi:10.1371/journal.pone.0012684 · PMC2937036 - Nilsen J, Aaen KH, Benjakul S, Ruso-Julve F, Greiner TU, Bejan D, et al.. Enhanced plasma half-life and efficacy of engineered human albumin-fused GLP-1 despite enzymatic cleavage of its C-terminal end. Commun Biol. 2025;8(1):810.
PMID 40419755 · doi:10.1038/s42003-025-08249-8 · PMC12106674 - Nishida K, Yamashita H, Takahashi A, Miyako E. Intrinsically Disordered Polypeptides-Based Stealth Materials Enable Enhanced Photothermal Cancer Therapy Using an In Situ Fiber-Based Penetrating Laser System. Small Sci. 2026;6(7):e70342.
PMID 42483662 · doi:10.1002/smsc.70342 · PMC13387302 - Nussinov R, Liu Y, Zhang W, Jang H. Protein conformational ensembles in function: roles and mechanisms. RSC Chem Biol. 2023;4(11):850-864.
PMID 37920394 · doi:10.1039/d3cb00114h · PMC10619138 - Ogundele AV, Nongthombam GS, Nwagu AD, Silva HH, Fabiyi OA. Venom-Derived Enzyme Inhibitors as Anticancer Agents: Structure-Activity Relationships, Molecular Targets and Mechanistic Insights. Molecules. 2026;31(13).
PMID 42451765 · doi:10.3390/molecules31132398 · PMC13363293 - Pantelejevs T, Zuazua-Villar P, Koczy O, Counsell AJ, Walsh SJ, Robertson NS, et al.. A recombinant approach for stapled peptide discovery yields inhibitors of the RAD51 recombinase. Chem Sci. 2023;14(47):13915-13923.
PMID 38075664 · doi:10.1039/d3sc03331g · PMC10699557 - Parkkinen T, Heiniluoto H, Jänis J, Takkinen K, Rouvinen J. Structural Basis for Trivalent Cross-Linking of a Patient-Derived IgE Antibody by the Major Peanut Allergen Ara h 2.0201. Allergy. 2026;81(6):2036-2046.
PMID 41700105 · doi:10.1111/all.70258 · PMC13256284 - PAULING L, COREY RB, BRANSON HR. The structure of proteins; two hydrogen-bonded helical configurations of the polypeptide chain. Proc Natl Acad Sci U S A. 1951;37(4):205-11.
PMID 14816373 · doi:10.1073/pnas.37.4.205 · PMC1063337 - PAULING L, COREY RB. The pleated sheet, a new layer configuration of polypeptide chains. Proc Natl Acad Sci U S A. 1951;37(5):251-6.
PMID 14834147 · doi:10.1073/pnas.37.5.251 · PMC1063350 - PAULING L, COREY RB. The polypeptide-chain configuration in hemoglobin and other globular proteins. Proc Natl Acad Sci U S A. 1951;37(5):282-5.
PMID 14834151 · doi:10.1073/pnas.37.5.282 · PMC1063354 - Peterle D, Yan NL, Klimtchuk ES, Wales TE, Gursky O, Kelly JW, et al.. Small molecule stabilization of diverse amyloidogenic immunoglobulin light chains revealed by hydrogen-deuterium exchange mass spectrometry. bioRxiv. 2026.
PMID 41542593 · doi:10.64898/2026.01.07.698275 · PMC12803214 - Pinheiro I, Calo N, Paolini-Bertrand M, Hartley O. Arylsulfatases and neuraminidases modulate engagement of CCR5 by chemokines by removing key electrostatic interactions. Sci Rep. 2024;14(1):292.
PMID 38167636 · doi:10.1038/s41598-023-50944-1 · PMC10762049 - Puerta-González A, Soto-Ospina A, Montoya Osorio Y, Bustamante-Osorno J, Salazar-Peláez LM. Structural and immunogenic evaluation of silk proteins from Bombyx mori using advanced bioinformatics and deep learning for biomaterials applications. J Genet Eng Biotechnol. 2026;24(2):100692.
PMID 42309597 · doi:10.1016/j.jgeb.2026.100692 · PMC13137201 - Rahimi N, Molnár LA, Sur A, Dózsa Z, Molnár P, Kovács Z, et al.. Sorbitol modulates the structure and nanomechanics of κ-casein amyloid fibrils. Biophys J. 2026;125(9):2045-2058.
PMID 41865247 · doi:10.1016/j.bpj.2026.03.042 · PMC13351584 - Ramachandran GN. Need for nonplanar peptide units in polypeptide chains. Biopolymers. 1968;6(10):1494-6.
PMID 5685106 · doi:10.1002/bip.1968.360061013 - RAMACHANDRAN GN, RAMAKRISHNAN C, SASISEKHARAN V. Stereochemistry of polypeptide chain configurations. J Mol Biol. 1963;7:95-9.
PMID 13990617 · doi:10.1016/s0022-2836(63)80023-6 - Ramakrishnan C, Ramachandran GN. Stereochemical criteria for polypeptide and protein chain conformations. II. Allowed conformations for a pair of peptide units. Biophys J. 1965;5(6):909-33.
PMID 5884016 · doi:10.1016/S0006-3495(65)86759-5 · PMC1367910 - Ripka JF, Perez-Riba A, Chaturbedy PK, Itzhaki LS. Testing the length limit of loop grafting in a helical repeat protein. Curr Res Struct Biol. 2021;3:30-40.
PMID 34235484 · doi:10.1016/j.crstbi.2020.12.002 · PMC8244534 - Rroji M, Bob F, Lo Cicero L, Figurek A, Spasovski G. Biomarkers in Diabetic Kidney Disease: Early Detection, Prognostic Assessment, and Integration with Multi-Omics Signatures. Life (Basel). 2026;16(7).
PMID 42514233 · doi:10.3390/life16071164 · PMC13412889 - Ruan H, Yu C, Niu X, Zhang W, Liu H, Chen L, et al.. Computational strategy for intrinsically disordered protein ligand design leads to the discovery of p53 transactivation domain I binding compounds that activate the p53 pathway. Chem Sci. 2020;12(8):3004-3016.
PMID 34164069 · doi:10.1039/d0sc04670a · PMC8179352 - Rüdisser SH, Matabaro E, Sonderegger L, Güntert P, Künzler M, Gossert AD. Conformations of Macrocyclic Peptides Sampled by Nuclear Magnetic Resonance: Models for Cell-Permeability. J Am Chem Soc. 2023;145(50):27601-27615.
PMID 38062770 · doi:10.1021/jacs.3c09367 · PMC10739998 - Sachan S, Bhatia T, Baumann M, Weber M, Gordijenko I, Voigt B, et al.. The intrinsically disordered region of the human parathyroid hormone controls functional amyloid properties. J Biol Chem. 2026;302(4):111295.
PMID 41708001 · doi:10.1016/j.jbc.2026.111295 · PMC12993186 - Saino H, Sugiyabu T, Ueno G, Yamamoto M, Ishii Y, Miyano M. Crystal Structure of OXA-58 with the Substrate-Binding Cleft in a Closed State: Insights into the Mobility and Stability of the OXA-58 Structure. PLoS One. 2015;10(12):e0145869.
PMID 26701320 · doi:10.1371/journal.pone.0145869 · PMC4689445 - Salladini E, Jørgensen MLM, Theisen FF, Skriver K. Intrinsic Disorder in Plant Transcription Factor Systems: Functional Implications. Int J Mol Sci. 2020;21(24).
PMID 33371315 · doi:10.3390/ijms21249755 · PMC7767404 - Salvi S, Linciano P, Collina S, Rossino G. Cyclic Peptides as Modulators of Protein-Protein Interactions: A Survival Guide from Discovery Platforms to AI-Driven Design. Int J Mol Sci. 2026;27(13).
PMID 42450332 · doi:10.3390/ijms27136067 · PMC13361436 - Sancho-Vaello E, Kücükyildiz H, Gil-Carton D, Biarnés X, Zeth K. Structure of a barrel-stave pore formed by magainin-2 reveals anion selectivity and zipper-mediated assembly. Sci Rep. 2025;15(1):39830.
PMID 41233432 · doi:10.1038/s41598-025-23539-1 · PMC12615594 - Sanger F, Thompson EO. The amino-acid sequence in the glycyl chain of insulin. Biochem J. 1952;52(1):iii.
PMID 13018185 - SANGER F, SMITH LF, KITAI R. The disulphide bridges of insulin. Biochem J. 1954;58(330th Meeting):vi-vii.
PMID 13198880 - Sangwung P, Ho JD, Siddall T, Lin J, Tomas A, Jones B, et al.. Class B1 GPCRs: insights into multireceptor pharmacology for the treatment of metabolic disease. Am J Physiol Endocrinol Metab. 2024;327(5):E600-E615.
PMID 38984948 · doi:10.1152/ajpendo.00371.2023 · PMC11559640 - Schweitzer-Stenner R. The multifaceted character of water as solvent for proteins: From poor for folded proteins to good for (some) intrinsically disordered proteins and protein segments. Q Rev Biophys. 2026;59:e5.
PMID 41725507 · doi:10.1017/S0033583526100080 - Sgourakis NG, Merced-Serrano M, Boutsidis C, Drineas P, Du Z, Wang C, et al.. Atomic-level characterization of the ensemble of the Aβ(1-42) monomer in water using unbiased molecular dynamics simulations and spectral algorithms. J Mol Biol. 2011;405(2):570-83.
PMID 21056574 · doi:10.1016/j.jmb.2010.10.015 · PMC3060569 - Smith AB, Charnley AK, Hirschmann R. Pyrrolinone-based peptidomimetics. "Let the enzyme or receptor be the judge". Acc Chem Res. 2011;44(3):180-93.
PMID 21175156 · doi:10.1021/ar1001186 · PMC3078624 - Streit JO, Invernizzi M, Bottaro S, Tamiola K, Lindorff-Larsen K. Transient tertiary structure in intrinsically disordered proteins revealed by multithermal enhanced sampling. Nat Commun. 2026;17(1).
PMID 42129207 · doi:10.1038/s41467-026-73067-3 · PMC13291234 - Suleman M, Said A, Khan H, Rehman SU, Alshammari A, Crovella S, et al.. Mutational analysis of SARS-CoV-2 ORF6-KPNA2 binding interface and identification of potent small molecule inhibitors to recuse the host immune system. Front Immunol. 2023;14:1266776.
PMID 38283360 · doi:10.3389/fimmu.2023.1266776 · PMC10811244 - Sun H, Tan H, Chu Y, Li J, Wang R, Wei DQ. Design of permeability-optimized target-binding macrocycles via direct preference optimization. Chem Sci. 2026;17(20):10223-10236.
PMID 41982931 · doi:10.1039/d6sc01722c · PMC13074632 - Suresh A, Schweitzer-Stenner R, Urbanc B. Nearest-Neighbor Effects in Short Unfolded Peptides: An Assessment of Molecular Dynamics Force Fields. J Chem Inf Model. 2026;66(12):7190-7206.
PMID 42240685 · doi:10.1021/acs.jcim.6c00438 · PMC13292223 - Svensson O, Gerelli Y, Skepö M. Multidimensional Decomposition and Ensemble Modeling of Histatin 1 and Its Siblings: Detailing Structure and Biological Function Using an Integrative Approach. J Chem Inf Model. 2025;65(13):7089-7101.
PMID 40600658 · doi:10.1021/acs.jcim.5c00854 · PMC12264947 - Szarszoń K, Kachnowicz J, Janek T, Domínguez-Martin A, Jezierska A, Wątły J. Dual Amino Acid Swap in MUC7-Derived Peptide Enhances Resistance and Modulates Zn(II) and Cu(II) Complex Stability, Secondary Structure and Antimicrobial Activity. Inorg Chem. 2026;65(10):5611-5626.
PMID 41781350 · doi:10.1021/acs.inorgchem.5c05849 · PMC13298914 - Taherali F, Chouhan N, Wang F, Lavielle S, Baran M, McCoubrey LE, et al.. Impact of Peptide Structure on Colonic Stability and Tissue Permeability. Pharmaceutics. 2023;15(7).
PMID 37514143 · doi:10.3390/pharmaceutics15071956 · PMC10384666 - Takashima H, Du Vigneaud V, Merrifield RB. The synthesis of deamino-oxytocin by the solid phase method. J Am Chem Soc. 1968;90(5):1323-5.
PMID 5636536 · doi:10.1021/ja01007a038 - Tang Z, Jiang W, Li S, Huang X, Yang Y, Chen X, et al.. Design and evaluation of tadpole-like conformational antimicrobial peptides. Commun Biol. 2023;6(1):1177.
PMID 37980400 · doi:10.1038/s42003-023-05560-0 · PMC10657444 - Taylor AIP, Radford SE. Amyloid fibril polymorphism: Structural mechanisms of assembly and the links to disease. Curr Opin Struct Biol. 2026;98:103245.
PMID 41830661 · doi:10.1016/j.sbi.2026.103245 · PMC7618959 - Thacker D, Sanagavarapu K, Frohm B, Meisl G, Knowles TPJ, Linse S. The role of fibril structure and surface hydrophobicity in secondary nucleation of amyloid fibrils. Proc Natl Acad Sci U S A. 2020;117(41):25272-25283.
PMID 33004626 · doi:10.1073/pnas.2002956117 · PMC7568274 - Thakkar H, Eerla R, Sharma L, Shah RP. A rapid discriminative hydrogen-deuterium exchange and LC-HRMS/MS strategy for primary and higher order structural mapping of therapeutic proteins: a case study using filgrastim. Anal Methods. 2023;15(12):1527-1535.
PMID 36880166 · doi:10.1039/d2ay01788a - Thorpe MP, Hopkins CR, Johnston JN. End-to-End Backbone Cyclization Enhances Passive Permeability of bRo5 Oligomeric Depsipeptides with Nonlinear Size Dependence. ACS Med Chem Lett. 2025;16(4):638-645.
PMID 40236530 · doi:10.1021/acsmedchemlett.5c00037 · PMC11995216 - Timchenko MA, Galzitskaya OV, Chulkov AV, Likhachev IV, Glyakina AV, Molchanov MV, et al.. YB-1 AP-CSD Forms Cross-β Amyloid Fibrils Without Secondary-Structure Conversion In Vitro. Int J Mol Sci. 2026;27(8).
PMID 42074193 · doi:10.3390/ijms27083553 · PMC13116819 - Tino AS, Quagliata M, Schiavina M, Attanasio L, Santos BPO, Pacini L, et al.. Peptide ligands to explore interactions with intrinsically disordered multidomain proteins: the case of SARS-CoV-2 nucleocapsid protein. Sci Rep. 2026;16(1).
PMID 42000821 · doi:10.1038/s41598-026-46442-9 · PMC13254127 - Trinh TB, Upadhyaya P, Qian Z, Pei D. Discovery of a Direct Ras Inhibitor by Screening a Combinatorial Library of Cell-Permeable Bicyclic Peptides. ACS Comb Sci. 2016;18(1):75-85.
PMID 26645887 · doi:10.1021/acscombsci.5b00164 · PMC4710893 - Tripathi T, Uversky VN, Giuliani A. Extending the classical sequence-structure-function paradigm through protein dynamics and context-dependent behavior. FEBS Lett. 2026.
PMID 42387899 · doi:10.1002/1873-3468.70404 - Tu T, Tseng CY, Islam MD, Hsu WL, Ou SC, Yamazaki T, et al.. Adjuvant-free pH-controlled aggregates of E. coli-expressed H1N1-RBD enhance neutralizing antibody responses and confer protection against influenza virus. Vaccine. 2026;88:128898.
PMID 42435621 · doi:10.1016/j.vaccine.2026.128898 - Uttarkar A, Niranjan V, Saxena A, Kumar V. QuPepFold: A python package for hybrid quantum-classical protein folding simulations with CVaR-optimized VQE. PLoS One. 2026;21(2):e0342012.
PMID 41671215 · doi:10.1371/journal.pone.0342012 · PMC12893577 - Vogel A, Blakely A, Dao Y, Lin NP, Chou D, Hill CP. Structural basis of insulin receptor antagonism by bivalent site 1-site 2 ligands S961 and Ins-AC-S2. Nat Commun. 2026;17(1).
PMID 42265100 · doi:10.1038/s41467-026-73851-1 · PMC13249828 - Voronko OE, Khotina VA, Kashirskikh DA, Lee AA, Gasanov VAO. Antimicrobial Peptides of the Cathelicidin Family: Focus on LL-37 and Its Modifications. Int J Mol Sci. 2025;26(16).
PMID 40869425 · doi:10.3390/ijms26168103 · PMC12386566 - Wekalao J, Topisia T. Aquatic-derived antimicrobial peptides and their strategically modified analogues as prospective anticancer therapeutics: a comprehensive systematic review of enhancement methodologies and mechanistic insights. Front Chem. 2026;14:1876776.
PMID 42482696 · doi:10.3389/fchem.2026.1876776 · PMC13385044 - Welch EF, Rush KW, Eastman KAS, Bandarian V, Blackburn NJ. The binuclear copper state of peptidylglycine monooxygenase visualized through a selenium-substituted peptidyl-homocysteine complex. Dalton Trans. 2025;54(12):4941-4955.
PMID 39981625 · doi:10.1039/d5dt00082c · PMC12088448 - Werner HM, Cabalteja CC, Horne WS. Peptide Backbone Composition and Protease Susceptibility: Impact of Modification Type, Position, and Tandem Substitution. Chembiochem. 2016;17(8):712-8.
PMID 26205791 · doi:10.1002/cbic.201500312 · PMC4721950 - Wiggenhorn AL, Abuzaid HZ, Coassolo L, Li VL, Tanzo JT, Wei W, et al.. A class of secreted mammalian peptides with potential to expand cell-cell communication. Nat Commun. 2023;14(1):8125.
PMID 38065934 · doi:10.1038/s41467-023-43857-0 · PMC10709327 - Wu H, Wang W, Hu Z, Yin Y, Peng D, Zhou Q, et al.. Structural basis for inhibition of the voltage-gated sodium channel NaV1.7 by the tarantula toxin HWTX-I. J Biol Chem. 2026;302(6):113130.
PMID 42107643 · doi:10.1016/j.jbc.2026.113130 · PMC13266018 - Yang F, Chen W, Dabbour M, Kumah Mintah B, Xu H, Pan J, et al.. Preparation of housefly (Musca domestica) larvae protein hydrolysates: Influence of dual-sweeping-frequency ultrasound-assisted enzymatic hydrolysis on yield, antioxidative activity, functional and structural attributes. Food Chem. 2024;440:138253.
PMID 38150897 · doi:10.1016/j.foodchem.2023.138253 - Yoshimura M, Teramoto T, Asano H, Iwamoto Y, Kondo M, Nishimoto E, et al.. Crystal structure of tick tyrosylprotein sulfotransferase reveals the activation mechanism of the tick anticoagulant protein madanin. J Biol Chem. 2024;300(3):105748.
PMID 38354785 · doi:10.1016/j.jbc.2024.105748 · PMC10951654 - Zhang J, Yin Z, Li Y, Ge C, Zhang Z, Yuan P, et al.. Deep learning-driven discovery and mechanism of action study of a minimalist conopeptide targeting α7 nicotinic acetylcholine receptor. Acta Pharm Sin B. 2026;16(7):4147-4165.
PMID 42453416 · doi:10.1016/j.apsb.2025.12.035 · PMC13366377 - Zhang X, Li C, Deng Z, Liang C, Li J. Development of a Mechanism of Action-Reflective Cell-Based Reporter Gene Assay for Measuring Bioactivities of Therapeutic Glucagon-like Peptide-2 Analogues. Molecules. 2025;30(9).
PMID 40363720 · doi:10.3390/molecules30091915 · PMC12073449 - Zhivnov A, Ottochian C, El-Bouri K, Francis R, Campbell R, Macpherson A, et al.. Modulator-induced conformational changes in complement C5, implications for function and drug design. Front Immunol. 2026;17:1834455.
PMID 42220504 · doi:10.3389/fimmu.2026.1834455 · PMC13219244
27How this document was assembled
The evidence base has three layers, and they are reported as three different numbers because a single figure would imply coverage this document does not have.
Layer A, the local library. Every JATS asset in project 05, the Therapeutic Peptide Research Library, was parsed and scored: 10,204 documents, about 110,631 printed-page equivalents. Scoring was not by keyword frequency but by how many of eighteen concept families — drawn from this monograph's coverage requirements — each document engages. A monograph about how levels of description connect is served by documents that connect them, and a paper naming one residue four hundred times engages one family.
Three of those eighteen families are ubiquitous across biochemistry: structural methods, non-covalent forces, and computational structure. A paper that ran an NMR assay and named a hydrogen bond is not a paper about how sequence produces conformation. Those three may raise a document's score but may not carry it into the corpus alone, and 552 local documents were rejected on exactly that basis — counted separately from the 3,789 that engaged nothing at all and the 312 with no readable body text. Reporting one rejection figure would have credited the screen with work it did not do.
Layer B, a targeted external harvest. Scoring Layer A showed the library is dense where peptide action is and sparse where the physics underneath it is: of the local reading corpus, 711 documents engage receptor recognition and 567 aggregation, but only 78 touch intrinsic disorder, 168 secondary structure and 187 conformational constraint. Those are the foundations this document is built on. Sixteen PubMed axes built on MeSH descriptors produced an indexed surface of 187,274 records, partitioned into six publication windows, of which 7,146 were retrieved as verified records and 2,707 as open-access full text. The concept must appear in the indexed record — title, abstract or MeSH — which is the indexer's judgement that a paper is about the concept rather than merely mentioning it.
Layer C, the verified historical record. Neither literature layer can supply the discovery history: the local library holds almost nothing before 2005, and a recency-weighted query returns commentary on a classic paper rather than the paper. Forty-three foundational works were resolved individually against NCBI and each author–journal–volume line was read before acceptance. Two of them are in a journal PubMed does not index that far back, and are described in the text rather than cited for that reason.
Merging Layers A and B keyed by identifier rather than summed — twelve documents appear in both and are counted once — gives the corpus this monograph was drawn from: 2,183 documents, 16,357,598 words, approximately 32,715 printed-page equivalents, together with 161 verified references. Of those, 227 documents were placed on the reading lists and read; the other 1,956 were screened and not read. Both figures belong in the same sentence, because "drawn from" and "read" are different claims and the larger number is the less informative one.
A structural limitation of the instrument, rather than of any one search. Both literature layers are built from open-access full text, and an open-access corpus is a poor place to look for foundational constants — those live in textbooks and in closed classics. This document therefore declines to print the pitch of an α-helix, any per-residue helical propensity scale, hydrogen-bond and salt-bridge energies in kilocalories per mole, and the amide rotational barrier, because the evidence base read for it does not carry them. That is a consequence of how the corpus was built, and naming it once is more honest than presenting each absence as a separate discovery.
| Stage | What it does | Result |
|---|---|---|
| 01 | Parse and score every project-05 JATS asset on eighteen concept families | 10,204 scanned |
| 02 | Sixteen MeSH-built axis queries against PubMed | 187,274 surface |
| 02b | Individual verification of the foundational papers | 43 resolved |
| 02d | Re-harvest partitioned by publication window | 7,146 records |
| 03 | Open-access full-text retrieval and substantive-use screen | 2,707 full texts |
| 04 | Bounded reading list per narrative theme, so breadth is forced rather than hoped for | 227 read |
| 05 | Keyed merge and inventory | 2,183 unique |
| 06 | Assembly, with the build gates | 1 deliverable |
What was counted and not read. Of the local surface, 4,503 documents engaged one to four concept families and were classed peripheral, 552 engaged only ubiquitous families, 3,789 engaged none, and 312 carry no readable body text. Of the external layer, 1,455 documents fell below the reading threshold or carried no retrievable body. Of the 2,183 documents in the read tiers, 227 were placed on the per-theme reading lists and read; the remaining 1,956 were counted and not read. Naming these is what makes the smaller number trustworthy. A corpus figure that silently absorbed them would be a claim about coverage this document has not earned.
Two defects in the build machinery are worth recording, because both produced output that looked correct. The first external harvest sorted its results by date with a two-hundred-record limit against surfaces of up to forty-seven thousand records; of the 884 documents it retained, 763 were published in 2026 and the remaining 121 in 2025. Sixteen years of the literature the queries were written to cover had been truncated away, silently, while every count in the log looked healthy. Re-running each axis partitioned into six publication windows, with relevance ordering inside each window, recovered coverage back to 1997. And the assembly gate that refuses an unrenumbered figure marker looked only for the exact form the renumbering stage always rewrites; an inline cross-reference written with the same placeholder survives it, and a previous monograph in this series shipped a caption promising a figure that did not exist. The gate now refuses on any surviving placeholder anywhere in the document and prints its surrounding text. It proved itself immediately: the first build of this monograph was refused because this paragraph quoted the placeholder while describing it. A stage that runs is not a stage that is right.
28Evidence handling
Findings are labelled by the kind of study that produced them, in the sentence that reports them. For a structural monograph that taxonomy is not animal, human and in-vitro alone — it is also how the structure was obtained, and every structural statement carries both. A crystal structure is called a crystal structure and its resolution is given; a cryo-EM reconstruction is called a reconstruction; an NMR result is reported with its solvent; a predicted structure is reported with its predictor and its confidence; a docking pose is called a pose. These are not interchangeable, and the difference between them is the subject of Section 17 rather than a footnote to it.
The solvent is part of a conformation claim and is always named. A peptide that is helical in trifluoroethanol or in a detergent micelle may be a random coil in water, and reporting the helix without the solvent is the most common overstatement in this literature. Where a source gives a helicity without its conditions, this document says the conditions were not given.
Recency is weighted but not blindly, and the rule has three parts. A newer measurement of the same quantity supersedes an older one where the method is at least as good and the population comparable, and both numbers are printed. A newer null does not automatically supersede an older positive; it is reported as a conflict with the reason it does or does not overturn the earlier finding. And recency does not apply to foundations — Pauling and Corey's helix geometry has been refined, not superseded, and a 2026 paper does not outrank it on what an α-helix is.
One axis of this document has an unusual freshness problem and is treated separately. Computational structure prediction has changed more between 2021 and 2026 than in the thirty years before it, so every claim in Section 17 about what prediction can and cannot do is dated explicitly, and its limits are stated as of this corpus rather than as permanent.
Three kinds of absence are reported rather than written around. Where a number does not exist in this evidence base, no number is printed. Several figures planned for this document were not drawn because the matched dataset they would have needed — one peptide measured in three solvents, a conservation profile and an energetic-contribution profile for the same molecule — is not in the corpus; those omissions are listed in the unresolved-structure appendix rather than filled with plausible values. Where two sources disagree, both are given with their designs. And where a widely repeated claim is not supported, it is named as unsupported rather than quietly omitted.
The evidence base is open-access biased. Both literature layers are built from PubMed Central, so work in journals that deposit no open-access full text is reachable only through its abstract. No claim in this document should be read as implying that the closed literature was searched.
Every number in this document was read twice
The reading that produced this monograph was done once, by one reader per theme. That is not enough, and saying so is not a formality: a reader who misreads a value puts it into the prose with a correct identifier beside it, and no count, gate or audit anywhere in the build can tell. So after the first complete draft, 182 passages — every paragraph in the running prose carrying both a number and a citation — were re-read against the source full text by readers instructed to try to refute them.
About one passage in five needed correction, and the dominant fault was not arithmetic but attribution. Nine citations pointed at entirely unrelated papers, every one by the same mechanism: a common first-author surname and a year matching whichever record happened to be indexed first. Three passages in Section 07, including this document's central argument about solvent dependence, cited a membrane-transporter structure containing no peptide, no inhibitory concentration and no circular-dichroism spectrum. The numbers were right; the source could not have produced them. A citation that resolves is not a citation that is correct, and the audit that checks resolution cannot see the difference.
Nine numbers were wrong outright. Two are worth naming here because they bear on how this document should be read. In Section 04 — the section warning against confusing modelled values with measured ones — a pair of distances from computational models was reported as though measured in a crystal. And in Section 14 an argument about non-additivity rested on a figure attributed to authors who never stated it; the correct figure, computed here from their own table and labelled as this document's arithmetic, makes the argument stronger rather than weaker. Two further passages counted their independent support as three sources and two reviews where there was one of each, in places where the number of independent sources was the argument.
Nine passages remain unverifiable, because the source full text is not held locally or the claim is qualitative; they are marked as such and were not corrected, since correcting an unread source from inference is the failure this pipeline exists to prevent. Three claims were searched for properly in their cited source and were not there; each has been removed or restated as unsupported rather than quietly re-sourced. The full record, with the exact clause, what the source says, the correct value and a severity for each, is filed with the project.
The forty-four figures are of two kinds running in one numbered series: twenty-nine authored for this document as inline vector graphics, and fifteen commissioned plates. No third-party figure has been reproduced, adapted or redrawn. Each figure declares its epistemic class in its caption from a vocabulary fixed before drafting. Because no molecular-graphics or molecular-dynamics software was available to this session, no figure here is a rendering of coordinates, and none claims to be; the commissioned programme reached the same position independently, classifying all fifteen of its plates as hypothetical illustrations drawn for explanation.
The plates were checked against this evidence base rather than accepted as delivered, and that check changed four printed values and three captions. Two plates print numbers that are simply wrong — a mass increment and a charge sign on one, two entries of a hydropathy scale on the other — and both are corrected in place. Both were caught by arithmetic internal to the plate, before any source was opened: one plate's own two panels contradict each other. Two further plates assert that aggregates are markedly more immunogenic than monomer, which this corpus asserts three times and measures never, and their captions say so.
Two figures the commission specified were never supplied, and their absence is the largest unmet requirement in this document. One was a worked structural case study — a single therapeutic peptide shown in four states, free-state NMR ensemble, receptor-bound experimental structure, predicted model with confidence colouring and a simulation snapshot, rendered from deposited coordinates with full metadata. The other was a historical plate of early physical molecular models and first-generation sequencing apparatus. The first would have been the only figure in the document derived from real coordinates, and the argument of Section 17 would have been considerably stronger for it. Both require resources this session did not have: deposited-structure accessions with a rendering pipeline, and licensed or verified public-domain photographs.
An honest estimate of what remains wrong on the plates. Four errors were found among roughly thirty printed values that this corpus can check. Three plates print values it cannot check — six interaction energies, a helical propensity ranking and an unlabelled stability chart — and those are captioned as shown-as-supplied. If the error rate on the uncheckable values resembles the rate on the checkable ones, one or two of them are also wrong. "Not independently verified" and "probably contains an error" are different statements, and the second is the one the arithmetic supports.
Continue exploring