Clustered CTCF binding is an evolutionary mechanism to maintain topologically associating domains, bioRxiv, 2019-06-13

ABSTRACTCTCF binding contributes to the establishment of higher order genome structure by demarcating the boundaries of large-scale topologically associating domains (TADs). We have carried out an experimental and computational study that exploits the natural genetic variation across five closely related species to assess how CTCF binding patterns stably fixed by evolution in each species contribute to the establishment and evolutionary dynamics of TAD boundaries. We performed CTCF ChIP-seq in multiple mouse species to create genome-wide binding profiles and associated them with TAD boundaries. Our analyses reveal that CTCF binding is maintained at TAD boundaries by an equilibrium of selective constraints and dynamic evolutionary processes. Regardless of their conservation across species, CTCF binding sites at TAD boundaries are subject to stronger sequence and functional constraints compared to other CTCF sites. TAD boundaries frequently harbor rapidly evolving clusters containing both evolutionary old and young CTCF sites as a result of repeated acquisition of new species-specific sites close to conserved ones. The overwhelming majority of clustered CTCF sites colocalize with cohesin and are significantly closer to gene transcription start sites than nonclustered CTCF sites, suggesting that CTCF clusters particularly contribute to cohesin stabilization and transcriptional regulation. Overall, CTCF site clusters are an apparently important feature of CTCF binding evolution that are critical the functional stability of higher order chromatin structure.

biorxiv genomics 100-200-users 2019

A robust benchmark for germline structural variant detection, bioRxiv, 2019-06-10

AbstractNew technologies and analysis methods are enabling genomic structural variants (SVs) to be detected with ever-increasing accuracy, resolution, and comprehensiveness. Translating these methods to routine research and clinical practice requires robust benchmark sets. We developed the first benchmark set for identification of both false negative and false positive germline SVs, which complements recent efforts emphasizing increasingly comprehensive characterization of SVs. To create this benchmark for a broadly consented son in a Personal Genome Project trio with broadly available cells and DNA, the Genome in a Bottle (GIAB) Consortium integrated 19 sequence-resolved variant calling methods, both alignment- and de novo assembly-based, from short-, linked-, and long-read sequencing, as well as optical and electronic mapping. The final benchmark set contains 12745 isolated, sequence-resolved insertion and deletion calls ≥50 base pairs (bp) discovered by at least 2 technologies or 5 callsets, genotyped as heterozygous or homozygous variants by long reads. The Tier 1 benchmark regions, for which any extra calls are putative false positives, cover 2.66 Gbp and 9641 SVs supported by at least one diploid assembly. Support for SVs was assessed using svviz with short-, linked-, and long-read sequence data. In general, there was strong support from multiple technologies for the benchmark SVs, with 90 % of the Tier 1 SVs having support in reads from more than one technology. The Mendelian genotype error rate was 0.3 %, and genotype concordance with manual curation was >98.7 %. We demonstrate the utility of the benchmark set by showing it reliably identifies both false negatives and false positives in high-quality SV callsets from short-, linked-, and long-read sequencing and optical mapping.

biorxiv genomics 100-200-users 2019

The murine transcriptome reveals global aging nodes with organ-specific phase and amplitude, bioRxiv, 2019-06-07

Aging is the single greatest cause of disease and death worldwide, and so understanding the associated processes could vastly improve quality of life. While the field has identified major categories of aging damage such as altered intercellular communication, loss of proteostasis, and eroded mitochondrial function1, these deleterious processes interact with extraordinary complexity within and between organs. Yet, a comprehensive analysis of aging dynamics organism-wide is lacking. Here we performed RNA-sequencing of 17 organs and plasma proteomics at 10 ages across the mouse lifespan. We uncover previously unknown linear and non-linear expression shifts during aging, which cluster in strikingly consistent trajectory groups with coherent biological functions, including extracellular matrix regulation, unfolded protein binding, mitochondrial function, and inflammatory and immune response. Remarkably, these gene sets are expressed similarly across tissues, differing merely in age of onset and amplitude. Especially pronounced is widespread immune cell activation, detectable first in white adipose depots in middle age. Single-cell RNA-sequencing confirms the accumulation of adipose T and B cells, including immunoglobulin J-expressing plasma cells, which also accrue concurrently across diverse organs. Finally, we show how expression shifts in distinct tissues are highly correlated with corresponding protein levels in plasma, thus potentially contributing to aging of the systemic circulation. Together, these data demonstrate a similar yet asynchronous inter- and intra-organ progression of aging, thereby providing a foundation to track systemic sources of declining health at old age.

biorxiv genomics 100-200-users 2019

 

Created with the audiences framework by Jedidiah Carlson

Powered by Hugo