Kozak sequence libraries for systematically characterizing transgenes across expression levels.

Publication date: Jul 17, 2026

Typical mammalian overexpression systems test protein sequence variants with little control over expression levels and steady-state protein abundances, hindering interpretations of how protein sequence and expression converge to yield phenotypic outcomes. We explored the translation initiation sequence, commonly referred to as the Kozak sequence, as a means to systematically modulate protein steady-state abundance and cellular function. We performed sort-seq on a randomized library of the 6 nt preceding the start codon, amounting to 4042 sequences, with a single sequence integrated at a shared genomic integration site in each HEK 293T cell. Calibrating the scores revealed a ∼100-fold range of protein steady-state abundances possible through manipulation of the Kozak sequence. We identified human germline variants with predicted expression-reducing Kozak substitutions in disease-associated genes. Modulating the cell surface abundance of the host cell receptor ACE2 controlled the rate at which those cells became infected by SARS-like coronavirus spike pseudotyped particles. We demonstrated the potential of the approach by simultaneously testing Kozak libraries with a small panel of coding variants for ACE2 and STIM1. This approach lays the methodological groundwork for linking the causal relationships between protein sequence, abundance, and functional outcome.

Open Access PDF

Concepts Keywords
ACE2 protein, human
Angiotensin-Converting Enzyme 2
Angiotensin-Converting Enzyme 2
Codon, Initiator
Codon, Initiator
Gene Library
HEK293 Cells
Humans
SARS-CoV-2
Spike Glycoprotein, Coronavirus
Spike Glycoprotein, Coronavirus
spike protein, SARS-CoV-2
Transgenes

Original Article

(Visited 4 times, 1 visits today)

Leave a Comment

Your email address will not be published. Required fields are marked *