Deciphering Regulatory Syntax: The Interplay of Core Promoter and Transcription Factor Motifs in Gene Transcription
Principal Investigator: Charles Danko
DESCRIPTION (provided by applicant):
The activity of cis-regulatory elements is driven by transcriptional activators, repressors, and general transcription factors (GTFs) that interact with DNA sequence motifs. By analogy, these motifs function as words in the regulatory language of the genome. The organizational syntax of these motifs—akin to sentences—is a critical determinant of regulatory function at classical cis-regulatory elements. Most research on regulatory syntax has focused on interactions between different transcriptional activators, yet few generalizable rules have emerged. It remains debated whether syntax is a fundamental feature of most cis-regulatory elements. Our central hypothesis is that regulatory syntax arises from steric constraints, reflecting the need for transcription factors to bind in the correct position and orientation relative to Pol II or nucleosome complexes, which they direct enzymatic activity toward. To investigate this, we developed a machine learning model, CLIPNET, to predict transcription start sites from DNA sequence. CLIPNET identified DNA sequence motifs encoding binding sites for GTFs, transcriptional activators, and repressors, as well as intricate motif syntax connecting them. Collectively, these motifs and syntax relationships accurately explain how DNA sequence governs transcription initiation at virtually all human cis-regulatory elements. This proposal will analyze CLIPNET and related computational models using advanced interpretation tools. We will uncover syntax rules governing interactions between transcriptional activators, repressors, and DNA sequence motifs that direct PIC assembly and Pol II pausing, as well as the positioning of the +1 and -1 nucleosomes (Aim 1). We will then test whether syntax constraints are conserved among transcription factors with the same activation domains or those targeting the same step in the transcription cycle (Aim 2). Finally, we will validate syntax rules learned from CLIPNET-like models using a novel genome-integrated massively parallel reporter assay system. Our work aims to develop a "Rosetta Stone" that decodes the regulatory language hidden in our genome.
