Medical Biology · Year 1 · Medical University of Sofia
02
Proteins. Protein domains. Protein families
Free notes for topic 02 of the Medical Biology syllabus, open without an account. Written by a senior student against the syllabus question and checked line by line by a second student before publishing. How content is made
Updated
What this topic covers
note
The structure of proteins; how they are classified according to their shape; protein motifs and repetitive and non-repetitive secondary structures; protein domains; how protein domains are duplicated; how domains from different proteins are combined into one molecule; and what protein families are and how they arise.
The first part is revision from high school. The new material, and the part the university examines, begins at protein families.
Diagram of the structural hierarchy: primary sequence, secondary structure, protein motif, protein domain, tertiary structure and quaternary structure, each shown as a ribbon model
1. Protein composition
note
Proteins are polymers of 20 types of α-amino acid, which differ by their side chains:
H₂N - CH(R) - COOH, where R is the side chain.
Amino acids bind together through peptide bonds to form polypeptide chains, running from the N terminus to the C terminus:
H₂N - CH(R₁) - CO - NH - CH(R₂) - CO - ... - NH - CH(Rₙ) - COOH
The important observation is that the polypeptide chain has a monotonous backbone, (-NH-CH-CO-)ₙ, from which the side chains protrude. Everything that distinguishes one protein from another is in the side chains and their order; the backbone is the same in all of them.
Gallery of dozens of differently shaped and coloured protein and macromolecular structures, illustrating the great diversity built from the same amino acid building blocks, unlabelled
2. Primary structure
note
There are four levels of organisation of protein molecules: primary, secondary, tertiary and quaternary.
Primary structure is the amino acid sequence of the polypeptide chain. It includes the number, type and order of the amino acids.
It is based on peptide bonds.
It is determined by the respective gene.
It determines all the other structural levels of the protein, and therefore its function.
Example: protamine, a DNA packaging protein of the sperm nucleus. Its primary structure is fitted to its job in two ways at once. It is small, 58 amino acids, so that it fits in the groove of the double helix, and it is rich in basic amino acids so that it binds acidic DNA: 38 basic amino acids against only 1 acidic one.
Space-filling model of a small basic protein (blue) wound around a DNA double helix (white), illustrating protamine's primary structure fitted to packaging DNA
3. Spatial structure
note
The polypeptide chain is a three-dimensional structure. It can fold in space in different ways, called conformations.
Every protein has one conformation, called the native state, which is thermodynamically stable and suited to its biological functions.
Spatial structure is determined by primary structure.
The general rules of folding in aqueous solution:
polar amino acids end up at the molecule's surface;
hydrophobic ones are buried deep inside;
side chains carrying opposite charges interact, positive with negative.
Spatial structure is subdivided into secondary, tertiary and quaternary.
Schematic of an unfolded polypeptide chain with mixed surface side chains folding into a compact globule, with one type of side chain clustering in the buried core and the other left on the surface
Secondary structure
note
Secondary structure is the regular folding of parts of the polypeptide chain, based on numerous hydrogen bonds between non-adjacent >NH and >CO groups of the backbone:
> N - H ..... O = C <
Note that these bonds are between backbone groups, not side chains. That is what makes the structure regular and repeatable regardless of which protein it occurs in.
There are two basic types of secondary structure:
the α-helix;
the β-sheet.
Atomic diagrams of an alpha helix and a beta sheet showing backbone N, C and O atoms, R side chains and the hydrogen bonds (dashed) between non-adjacent backbone NH and CO groups
Tertiary structure
note
Tertiary structure is the final three-dimensional folding of the polypeptide chain, resulting from the irregular but non-random bending of the parts not included in secondary structure.
It is based on various non-covalent bonds - ionic, hydrogen, Van der Waals, hydrophobic - and in some cases on covalent disulfide bonds between the side chains of distant amino acids.
Disulfide bonds (bridges) are formed between two residues of the amino acid cysteine brought close together by the three-dimensional folding. Two -SH groups are joined by oxidation into an -S-S- bridge.
Diagram of a folded polypeptide backbone labelled with a hydrogen bond, hydrophobic and van der Waals interactions, a disulfide bridge (S-S) and an ionic bond between side chains
Protein classes based on shape
note
Based on their tertiary structure and solubility, proteins are subdivided into three classes:
globular;
fibrous;
membrane.
Membrane proteins must have hydrophobic parts, because part of the molecule sits inside the lipid bilayer.
Cartoon contrasting fibrous proteins (several coiled strands bundled in parallel) with globular proteins (a single chain folded into a compact tangled shape)Diagram of membrane protein types in a lipid bilayer: single and multiple transmembrane helices, a beta-barrel porin, a monotopic amphipathic helix, GPI-anchored and lipid-anchored proteins, and a peripheral extrinsic protein
Quaternary structure
note
Many proteins function in tertiary structure. In others, two or more polypeptide chains, called subunits, must bind together. This assembly of polypeptide chains into a functional complex is called quaternary structure.
It is based on the same types of bond between amino acid side chains that are found in tertiary structure.
The classic example.Myoglobin, which binds O₂ in muscles, has only tertiary structure. The related haemoglobin has 4 subunits, 2 α and 2 β, bound non-covalently. That quaternary structure is not decoration: it facilitates O₂ binding in the lungs and its release in the tissues, which a single chain cannot do.
Ribbon model of myoglobin, a single alpha-helical chain with one bound haem group (red), unlabelledRibbon model of haemoglobin showing its four subunits, two coloured red and two blue, each holding a haem group (green), illustrating its quaternary structure
To sum up: which bonds hold what together
note
Many weak, non-covalent bonds act in parallel to hold polypeptide regions tightly together.
The stability of the protein is determined by the combined strength of large numbers of such non-covalent bonds, not by any single strong one.
Where those bonds are found differs by level:
in secondary structure, between groups of the polypeptide backbone;
in tertiary structure, between side chains of the same polypeptide;
in quaternary structure, between side chains of different polypeptides.
Diagram of a folded polypeptide backbone labelled with a hydrogen bond, hydrophobic and van der Waals interactions, a disulfide bridge (S-S) and an ionic bond between side chains
4. Protein families
note
When a gene undergoes duplication, sometimes the two copies adapt to slightly different functions by accumulating point mutations.
By this process of duplication and divergence, gene and protein families and superfamilies have evolved. The standard example is the globins: five proteins make up the β-globin family, and there are further proteins in the wider globin superfamily. Comparing them, the haem group is conserved while the amino acids differ.
The beta-globin gene cluster on chromosome 11 (HBB, HBD, HBBP1, HBG1, HBG2, HBE1) above surface models of beta, delta, gamma-1, gamma-2 and epsilon globins and, separately, alpha globin, myoglobin, cytoglobin and neuroglobin, with the conserved haem pocket in red
5. Protein motifs
note
Within families and superfamilies, different proteins display the same folding peculiarities that adapt them to their function and underlie their success.
Such structural patterns, suited to particular functions and found across a number of proteins, are called protein structural motifs.
Example: the coiled coil. A keratin dimer consists of two α-helices coiled together. Keratins are fibrous cytoskeletal proteins responsible for the mechanical strength of epithelia. The coiled coil motif is found in keratins and other intermediate filament components, and also in myosin, tropomyosin and many other proteins.
Two long alpha helices (labelled N and C termini) wound around each other along their whole length, showing the coiled-coil motif of a keratin dimer
Non-repetitive secondary structure
note
Approximately one half of an average globular protein is organised into repetitive structures, the α-helix and the β-sheet.
The other half consists of non-repetitive secondary structures - random coils, loops and turns - which are irregular but still structured segments where the folding is not periodic.
These are not filler. They permit flexibility and enable the protein to fold into its unique three-dimensional tertiary structure.
The spatial arrangement matters: the repetitive structures form primarily the interior of the molecule, and they are connected by loop regions at the surface. The loops and turns, together with the α-helices and β-sheets, form the motifs that can be seen in many different proteins.
Small recurring motifs built from beta strands and an alpha helix, including a beta-alpha-beta unit, a beta hairpin and a beta-X crossover, plus a four-stranded Greek key topology diagram with strands numbered 1-4 and N and C termini markedRibbon structure of a small domain with a curved beta-sheet barrel (green) wrapped by alpha helices (blue) and connecting surface loops, N terminus marked
DNA-binding motifs and zinc fingers
note
Proteins that bind to DNA contain a limited number of motifs. They recognise and contact specific DNA sequences, and in this way they act as regulators of gene activity.
A DNA-binding domain (DBD) contains at least one structural motif that recognises double- or single-stranded DNA. The helix-loop-helix and helix-loop-sheet motifs are examples, found in a number of proteins that function as transcription factors and other regulatory proteins.
The zinc fingers are helix-loop-helix motifs stabilised by zinc atoms, and they contact the major groove of the DNA. Zinc fingers are very common among transcription regulators.
Structure of a single zinc finger labelled N terminus, beta-pleated sheet, alpha helix and C terminus, with a zinc ion coordinated by a Cys side chain and a His side chain
They are also becoming a tool: proteins with zinc fingers designed to bind almost any DNA sequence will soon be available to any laboratory that wants them.
Crystal structure of a five-finger GLI zinc-finger protein wrapped around DNA, with Finger1 to Finger5 and DNA base pairs bp2, bp10 and bp20 labelled
6. Protein domains
note
At the level of tertiary structure, most globular proteins consist of two or more distinct domains - structural units that fold more or less independently. Such proteins are called multi-domain.
The domains of a protein can have similar or different structure and functions.
Each domain has its own hydrophobic core, which is what makes it able to fold on its own.
Some small proteins fold as a whole and are called single-domain.
Examples.Ubiquitin, a small protein that labels other proteins for lysis, is single-domain; it has both types of secondary structure, whereas protein domains often have just one. DnaG, the primase of Escherichia coli and an enzyme participating in DNA replication, has three domains.
Ribbon model of ubiquitin, a small single-domain protein with both a beta sheet and an alpha helix, N and C termini labelledRibbon model of the DNA primase DnaG shown as three separately coloured domains (red, blue and green) connected in one chain
Example: pore-forming membrane proteins
note
The pore-forming membrane proteins are abundant in nature and show clearly how motifs are put to work.
Haemolysins are secreted by bacteria as soluble molecules. Afterwards they can be integrated into animal cell membranes, forming tunnels, and that causes lysis of erythrocytes, leukocytes and platelets.
The example is α-haemolysin, the main toxin produced by Staphylococcus aureus.
Ribbon model of the heptameric alpha-haemolysin pore, shown from above as a ring of seven subunits around a central channel and from the side as the membrane-spanning beta-barrel stemModel of the alpha-haemolysin beta-barrel pore embedded in a lipid bilayer, 1.5 nanometres wide, letting K+, Na+ and Ca2+ ions pass through
Example: the domain structure and function of steroid receptors
note
Receptors for steroid hormones are intracellular globular proteins with two named domains:
a ligand-binding domain (LBD);
a DNA-binding domain (DBD).
Their ligand, the respective steroid hormone, enters the cell. When the ligand-binding domain binds it, the receptor moves to the nucleus, its DNA-binding domain binds to target genes and activates them.
This is the cleanest illustration of what domains are for: one domain does the sensing and a physically separate one does the acting, and the connection between them is the whole mechanism of the hormone.
Diagram of steroid hormone action: hormone binds the nuclear receptor's ligand-binding domain (LBD), the receptor dimer's DNA-binding domain (DBD) binds nuclear DNA at the target gene, and the resulting mRNA is translated into protein that changes cell functionStructure of a nuclear receptor bound to DNA, with its N-terminal domain (NTD), DNA-binding domain (DBD), hinge and ligand-binding domain (LBD) labelled
Membrane proteins have domains too
note
Membrane proteins have three parts, each composed of one or more domains:
an extracellular part;
a transmembrane part;
an intracellular (cytoplasmic) part.
The transmembrane part is often a single α-helical domain of about 20 hydrophobic amino acids spanning the lipid bilayer.
The example is the receptor for epidermal growth factor (EGF). Its ligand is a secreted protein serving as a signal for epithelial cells to divide. As in other surface receptors, the extracellular part is ligand-binding and the intracellular part is signal-transducing.
Model of the EGF receptor spanning a lipid bilayer, with the ligand-binding extracellular domains above the membrane, the transmembrane helix crossing it, and the signal-transducing intracellular domain below
7. Domains can be duplicated within one protein
note
Sometimes the part of a gene encoding a protein domain is duplicated, leading to an elongated polypeptide chain with two similar adjacent domains.
This happens in all organisms but is easier in eukaryotes, where there is a tendency for protein domains to be encoded by an exon or a group of exons.
Because of such rearrangements, members of the same protein family differ in their length, that is, in their number of domains.
The example is the antibodies.IgG and IgE are similar, but IgE has one additional domain in the lower part of the molecule. They are the result of exon duplication.
Because the copies stay close together along the length of the molecule, the duplication is called tandem.
Domain diagrams of IgG (VH, VL, CL and Cgamma1-3 domains) and IgE (VH, VL, CL and Cepsilon1-4 domains), showing IgE with one extra constant domain per heavy chain
The structure of immunoglobulins
note
Immunoglobulins have quaternary structure. Their molecule contains four polypeptide chains: 2 heavy and 2 light, connected by disulfide bonds.
The light chain has 2 domains.
The heavy chain has 4 or 5 domains.
All the domains are similar, having evolved by duplication of an ancestral domain. Each of them is supported by an intrachain disulfide bond and has a characteristic β-sheet structural motif called the immunoglobulin fold.
Schematic Y-shaped immunoglobulin with two heavy chains and two light chains as loops linked by S-S disulfide bonds between domainsMolecular ribbon model of an IgG antibody, its two heavy chains (blue) and two light chains (pink/green) folded into separate immunoglobulin domains
The immunoglobulin superfamily
note
The immunoglobulin fold motif is found across many domains of many proteins, which together form the large immunoglobulin superfamily.
All of its members take part in the recognition and binding of molecules in the immune system, and most of them are met again in the immunology course.
Diagram of immunoglobulin superfamily receptors at the cell surface: Thy-1, CD4, CD8, CD28, class I MHC, class II MHC, the T-cell receptor, the Fc receptor and IgM, each built from immunoglobulin-fold domains
8. Exon shuffling: bringing domains of different origin together
note
In some proteins we can see domains which also exist in other proteins. That means that the genes of these proteins contain the same exons.
The mechanism that mixes exons taken from different genes is called exon shuffling.
Exon shuffling is a result of errors in meiotic recombination or of the action of transposons.
Exon shuffling diagram: a double crossover moves exon 2 and flanking intron sequence from gene 1 into gene 2, so gene 1 loses domain 2 of its protein while gene 2's protein gains it
Example: tissue plasminogen activator (TPA)
note
TPA is an extracellular protein that helps control blood clotting.
It has four domains of three types, each encoded by an exon, and one of those exons is present in two copies. Because each type of exon is also found in other proteins, it is supposed that the TPA gene arose by exon shuffling.
Diagram of the TPA gene assembled by exon shuffling and duplication from an EGF exon, a fibronectin finger (F) exon and two copies of a plasminogen kringle (K) exonRibbon model of tissue plasminogen activator (TPA) showing its several separately coloured domains linked in one chain, with attached sugar chains in white
9. What is the difference between a motif and a domain?
note
These two are constantly confused, and the distinction is examinable.
A protein MOTIF is a supersecondary structure containing several secondary structures in a stable arrangement. It is commonly found in other proteins, not necessarily having any similarity in function.
A protein DOMAIN contains a conserved polypeptide sequence which can be seen in many proteins with similar function. It can fold independently and can function relatively independently from the rest of the molecule. A domain has its own tertiary structure.
Put shortly: a motif is a shape that recurs; a domain is a working part that recurs.
Diagram of the structural hierarchy: primary sequence, secondary structure, protein motif, protein domain, tertiary structure and quaternary structure, each shown as a ribbon model
10. Tendencies in protein evolution
note
Eukaryotic proteins are generally larger than prokaryotic ones: about 360 amino acids average length in eukaryotes against about 260 in prokaryotes.
This difference is partly due to eukaryotes having more multi-domain proteins: 80% of eukaryotic proteins have two or more domains, against 65% of prokaryotic ones.
In addition, multicellular eukaryotes, unlike unicellular organisms, have a tendency towards tandem duplication of protein domains.
Table of median protein length in amino acids: about 375 in H. sapiens, 373 in D. melanogaster, 344 in C. elegans, 379 in S. cerevisiae, 356 in A. thaliana, versus 267 in 67 bacteria and 247 in 15 archaea3D bar chart of protein families by number of consecutive tandem domains (2, 3-5, 6-10, 11-30, 31-50) in multicellular eukaryotes, unicellular eukaryotes, bacteria and archaea, with only multicellular eukaryotes reaching the higher tandem-repeat counts
11. A modern approach to studying protein structure
note
A protein's three-dimensional structure can now be studied from its DNA sequence, using computational methods that translate the DNA into an amino acid sequence which is then used for modelling.
Diagram of a DNA sequence transcribed and translated into a protein amino acid sequence, which then folds into a 3D protein structure shown as a rainbow-coloured ribbon model
Tools can predict the three-dimensional structure from that sequence using techniques such as:
homology modelling, based on similar known structures;
de novo modelling, from scratch.
The predicted structures can then be visualised and analysed with molecular viewers and specialised servers that compare sequences and structures.
AlphaFold is an AI system developed by Google DeepMind that predicts a protein's three-dimensional structure from its amino acid sequence, regularly achieving accuracy competitive with experiment. DeepMind and the EMBL European Bioinformatics Institute have partnered to make hundreds of thousands, and eventually many millions, of AlphaFold structure predictions freely available through AlphaFold DB, which provides open access to over 200 million protein structure predictions.
AlphaFold2 model of a protein with Domain 1 (Macro), Domain 2 (Macro) and Domain 3 (ART) plus a disordered region, coloured by model confidence from very high to very low
The scale of what this changed is worth stating plainly: the initial release of the database included structure predictions for 98.5% of the proteins in the human proteome, whereas only 11% of human proteins have had their structure determined experimentally.
Newer versions of AlphaFold can also predict changes in structure under the influence of surrounding molecules and ions, as well as changes in ligand binding.
The most important things to know
note
Four levels: primary (sequence, peptide bonds, set by the gene), secondary (regular, hydrogen bonds between backbone groups), tertiary (irregular, bonds between side chains of one chain), quaternary (between side chains of different chains).
Folding rule: polar out, hydrophobic in, opposite charges attract. The native state is the one thermodynamically stable conformation.
Three classes by shape: globular, fibrous, membrane - and membrane proteins must have hydrophobic parts.
Myoglobin has no quaternary structure; haemoglobin has four subunits, and the quaternary structure is what lets it load in the lung and unload in the tissue.
Families arise by gene duplication and divergence; the globins are the standard example.
About half a globular protein is α-helix and β-sheet, forming the interior; the loops and turns at the surface give flexibility and complete the motifs.
Motif = recurring shape, no shared function required. Domain = independently folding working unit with its own hydrophobic core and its own tertiary structure.
Domains grow by tandem duplication (IgE has one domain more than IgG) and are combined across genes by exon shuffling (TPA), caused by meiotic recombination errors or transposons.
Eukaryotic proteins are longer (360 versus 260 amino acids) because more of them are multi-domain (80% versus 65%).
The four levels of protein structure: primary structure as a beaded amino acid sequence, secondary structure as an alpha helix and beta sheet, tertiary structure as one folded chain, and quaternary structure as haemoglobin's four chains
All 105 Medical Biology topics · exam study support, not clinical guidance. SuperMed is not affiliated with the Medical University of Sofia.
Start free. No card needed.
Every account starts free, with free topics in Cytology and Medical Biology. Super opens the rest. Super is €15 a month.