...
Skip to content
Pure Lab Peptides
Analytical Methods

Peptide Sequence: Notation, Methods and Identity

TL;DR · The short version

A peptide sequence records residue order, usually from the N-terminus to the C-terminus. Keep stereochemistry, terminal groups and other modifications with it. A listed sequence is a specification; evidence from a tested sample depends on the method and its coverage.

GHK and KHG contain the same three amino acids. They are not the same sequence, and a matching molecular weight alone cannot tell them apart.

A peptide sequence records the order of residues in a chain. Read it alongside the end groups and modifications, then distinguish the intended structure from what an analytical test actually supports. Those checks prevent a short string of letters from losing important chemistry.

Start with the two ends

TL;DR: Read the chain from N to C unless the source explicitly defines another notation.

The N-terminus is the amino end and the C-terminus is the carboxyl end. Sequences are conventionally written N to C. IUPAC peptide nomenclature distinguishes the residue chain and its terminal groups. For a modified or cyclic peptide, the source should define how those features are represented.

Consider Gly-His-Lys and Lys-His-Gly. They contain the same three residues but in different positions. A composition-based molecular formula can be the same even though the sequences differ. That is one reason a molecular weight alone cannot resolve every identity question.

Read a short peptide sequence step by step

TL;DR: GHK means Gly-His-Lys: three residues in that order, with two peptide bonds in the ordinary linear chain.

Using GHK as an example, the first residue is glycine (G), the second is histidine (H), and the third is lysine (K). It is a tripeptide because it has three residues. For the ordinary unmodified linear chain, there are two peptide bonds connecting those residues.

This ordered list is part of the molecule’s primary structure. It is different from a three-dimensional conformation: the letters do not show every way the chain can bend in solution. If a record supplies only a fragment of a longer protein, retain the residue numbering and the identity of the parent sequence.

Amino-acid notation shows GHK as three residues and distinguishes the sequence DF from D-phenylalanine.
Read the notation before interpreting the sequence. A one-letter D denotes aspartate; the D- prefix in D-Phe specifies stereochemistry. Enlarge illustration

Keep modifications attached to the sequence

TL;DR: Retain stereochemistry, terminal chemistry, ring connections and attached groups.

Feature Why it matters when comparing records
D/L configuration The stereochemical form is part of residue identity; do not remove a D- prefix.
Terminal modification Acetylated or amidated ends are different from the corresponding free termini.
Cyclization or disulfide bond Connectivity may not be fully represented by the linear string.
Attached group A linker, lipid or other conjugate belongs in the complete identity.
Counterion or associated water These affect the material description and mass basis without becoming chain residues.

How is a peptide sequence determined?

TL;DR: Edman sequencing and MS/MS obtain order-related evidence. An intact-mass measurement asks a different question.

Sequencing methods obtain evidence about residue order. They differ from simply calculating a mass from a proposed sequence or measuring which amino acids are present.

Edman degradation reads from the N-terminus

Edman chemistry identifies successive residues by removing one at a time from the N-terminal end. It requires an accessible, unmodified terminal amino group; a blocked end can prevent the ordinary reaction. The Wistar Institute authors’ N-terminal sequencing protocol explains this cycle and its practical limits. A partial read establishes only the segment successfully examined, not automatically the entire chain.

Tandem mass spectrometry examines fragments

In MS/MS, a selected peptide ion is fragmented and the resulting ions are measured. Those fragments can provide evidence about the sequence. The University of Washington’s targeted proteomics guide describes the distinction between a peptide precursor ion and its sequence-related fragment ions. Measuring the intact precursor alone is a different experiment.

Evidence in a record What to check
Intact molecular mass Whether it agrees with the proposed complete chemical form; agreement alone does not establish residue order.
N-terminal sequence read Which positions were read and whether a blocked terminus or incomplete read limits the result.
MS/MS fragment assignments Which parts of the proposed sequence the assigned fragments support, and which distinctions remain unresolved.

Separate a specified sequence from an experimentally supported sequence

TL;DR: A specified structure and an experimentally supported assignment are different entries in the record.

A catalog sequence is a specification: it states what the material is intended to be. An analytical result provides evidence about the submitted sample. A report may contain an intact-mass result, a chromatographic purity estimate or fragment-ion information, but those measurements answer different questions.

A fragment-based assignment does not automatically resolve every residue isomer, stereochemical distinction or structural ambiguity. The peptide reference-standard characterization study describes the value and limits of complementary methods. Interpret the stated method rather than upgrading a simple mass match to “complete sequence verified.”

Four laboratory questions separated by method: identity, purity, amount and structure.
Choose methods for the question being asked. No single result establishes every property of a peptide sample. Enlarge illustration

Compare like with like

TL;DR: Check fragments, termini, form and transcription before declaring a mismatch.

When two records appear inconsistent, first check whether one shows only an active fragment, one includes a terminal modification, or one reports a salt form while the other reports the parent molecule. Then check sequence direction and transcription. Different names can sometimes refer to the same specified structure, but similar names can also refer to different structures.

If a supplier name is used inconsistently in the literature, retain that uncertainty until the actual supplied identity is documented. Do not select whichever published sequence gives the closest mass and treat the issue as resolved.

What to record in a research material entry

TL;DR: Keep the notation, source and lot together, with unresolved details clearly marked.

Keep the exact sequence notation, modifications, material form, source document and lot reference together. Where a detail is not provided, leave it unresolved instead of filling it from a related product. This creates a record that another researcher can check without reconstructing your assumptions.

For a code reference beside the sequence, use our amino acid abbreviations guide. To understand why a matching mass is not a complete sequence assignment, see peptide molecular weight.