Research

Physics-based mechanistic models of how one-dimensional epigenetic information becomes three-dimensional structure, and how that structure regulates genes.

The genome encodes a substantial amount of regulatory information in one dimension: nucleosome positions, histone modifications, and the binding sites of proteins such as CTCF and cohesin. How that information shapes three-dimensional chromatin structure, and how that structure in turn regulates gene expression, is the question my research addresses. I build physics-based mechanistic models in which structure is generated by physical processes rather than fitted to data.

The themes below run from my PhD to current work. Two of them have live versions in the Learning Lab, if you would rather watch a model run than read about it.

Current work, New York University

How epigenetic marks reshape the X-inactivation center

Two chromosomes, one sequence, opposite regulatory outcomes. The difference is epigenetic, and it acts through structure.

In female mammals one X chromosome is transcriptionally silenced and stays silenced through the divisions that follow. The decision is made in a region of roughly one megabase called the X-inactivation center, where the gene Xist sits among its regulators: Tsix, which represses it, the enhancer Xite, which supports Tsix, and Jpx, Ftx and Rnf12, which promote it. Both X chromosomes carry identical copies of all of this. Whatever separates them is epigenetic, so it has to act by changing structure.

I model the region at two resolutions at once. A nucleosome-resolution model covers a 90 kb window around Xist with explicit linker DNA, histone tails, linker histones, and the effects of H3K27 acetylation and methylation. A coarse-grained polymer model covers the surrounding megabase with loop extrusion running explicitly. The contacts driven by H3K27me3 are assigned by a machine-learning model trained to predict contact enrichment from one-dimensional epigenomic tracks, which is the point where the 1D data enters the physics.

Running both lets us follow a change from a single histone modification up to the fold of the whole locus, and to compare three states: undifferentiated stem cells, the active X, and the inactive X.

What emerges is a rewiring rather than a uniform compaction. On the inactive X, Xist groups with its activators Jpx and Ftx into one accessible domain, while the long-range contact between Linx and the Xite/Tsix region is lost and Xite itself ends up buried in a compact, methylated environment. A single architecture therefore favours Xist expression and suppresses Tsix at the same time. At nucleosome resolution the changes are locus-specific: Xist adopts a fragmented, open structure near its transcription start site, while Xite forms large clutches of methylated nucleosomes.

This is the argument in miniature. The input is one-dimensional, the mechanism is physical, and the regulatory outcome follows from the structure those mechanisms produce.

The X-inactivation center: a genetic map of the 1 Mb locus, and the two simulation models, a coarse-grained polymer of the full region and a nucleosome-resolution model of the 90 kb core
The X-inactivation center, and the two resolutions we model it at.
Summary of chromatin reorganisation during X inactivation, comparing undifferentiated stem cells with the active and inactive X at both domain level and nucleosome level
What changes, and where: undifferentiated stem cells against the active and inactive X, at the scale of domains and at the scale of individual nucleosomes.

PhD work, IIT Bombay

Which mechanism folds a domain: structure against dynamics

Loop extrusion and epigenetic self-attraction can produce the same contact map. Motion tells them apart.

Most measurements of genome folding are contact maps, which record how often pairs of loci are found close together, averaged over millions of cells. Two quite different mechanisms generate similar maps. Motor proteins such as cohesin translocate along chromatin and extrude loops that grow, stall at CTCF sites and eventually release. Independently of that, segments in similar epigenetic states attract one another and separate into domains. Both produce domain structure, so a contact map on its own underdetermines the mechanism that built it.

We simulated a chromatin polymer with both processes active and measured structure and dynamics on the same system. They can disagree. Over a substantial range of parameters the contact map is set almost entirely by self-attraction, while the dynamics, meaning how quickly two loci come together and how far a locus diffuses in a given time, is set almost entirely by extrusion. Reading the map alone would identify the wrong mechanism.

Dynamics also discriminates between models that structure cannot separate. Fractal globule descriptions of chromatin predict particular scaling exponents for the relative motion of two loci. With active extrusion those exponents change, and they move towards the values reported experimentally.

The conclusion is methodological. Structure and dynamics have to be measured on the same system before a mechanism can be assigned. Where CTCF sites sit along the sequence is one-dimensional input; whether they actually shape the fold is a question only the dynamics answers.

Rendering of a coarse-grained chromatin polymer, coloured by domain

PhD work, IIT Bombay

From nucleosome-resolution data to a calibrated polymer model

A mechanistic model is only as good as its parameters. We measured them rather than assuming them.

Any model that turns one-dimensional information into three-dimensional structure has to commit to a representation. Chromatin is almost always drawn as beads joined by springs, but the parameters underneath that picture, meaning the size of a bead, the stiffness of the springs, the energy cost of bending, and whether two beads may overlap, were largely chosen rather than measured. Different groups used different values, and the choice propagates into every prediction the model makes.

We derived them from data instead. Micro-C resolves chromatin contacts at the level of individual nucleosomes. Starting there, we coarse-grained systematically, grouping 5, 10 and 25 nucleosomes into single beads, and at each level measured what the resulting beads actually look like: their size distributions, the fluctuations in bond length between neighbours, the effective spring constants, and the distribution of angles between consecutive bonds.

The central result is that coarse-grained chromatin beads are soft. They interpenetrate rather than excluding one another the way the hard spheres in most models do. We derived an effective soft potential and quantified how much overlap to expect at each scale.

Two further results bear on the 1D-to-3D question directly. Bead sizes, bond lengths and bond angles all differ systematically between TAD boundaries and TAD interiors, so local mechanical properties carry a signature of the large-scale organisation above them. And the angle distribution is bimodal, which indicates more than one preferred local conformation rather than a single average shape.

The output is a coarse-grained model with a measured value attached to every parameter, meant as a starting point for simulations that would otherwise have to guess.

Schematic of systematic coarse-graining from nucleosome resolution to a bead-spring chromatin polymer

Ongoing

Learning contacts from one-dimensional epigenomic signal

Statistical models supply the interactions. The mechanistic simulation then has to make predictions with them.

A mechanistic model has to be told what interacts with what. In the X-inactivation work, the contacts mediated by H3K27me3 are not assigned by hand. They come from a random-forest regression trained on 200 bp Micro-C maps together with genome-wide epigenomic tracks: H3K27ac, H3K4me1, H3K4me3, H3K36me3, H3K9me3, H3K27me3, ATAC-seq and CTCF, plus genomic distance.

One detail is what makes it useful. The model is trained on the contact enrichment that remains once the average distance dependence has been subtracted, so it learns the part of the signal associated with epigenetic state rather than the trivial fact that nearby loci contact each other often.

My wider interest is in keeping models of this kind interpretable. A network that reproduces a contact map but yields no mechanism is of limited use here, because the purpose is to hand a simulation an interaction it can act on, and whose consequences an experiment can then contradict. So the work favours representations that can be read back as physical quantities: interaction strengths, boundary elements, compartment identity.