Prof. Dr. Anna Poetsch
Research Area: Computational Biology in Ageing Research
Branches: Computational BiologyGeneticsMolecular Biology
Website: CECAD Profile Page
1. Research Background:
The Poetsch group is interested in how the DNA sequence contributes to ageing in the soma, to the development of ageing associated diseases and their treatment. The sequence contributes to gene regulation and protein coding, and also to processes like DNA replication, 3D genome organization, genome maintenance. It therefore impacts on its own stability, where it breaks, gets base adducts, how it is repaired, and where it mutates. However, despite having almost the same DNA sequence, these processes differ between tissues and circumstances, because the DNA interacts with the epigenome and the environment, which leads to tissue-specific genome biology, somatic genome evolution and routes towards ageing and disease development. Using a mixture of computational biology, machine learning and experimental data, the group dissects these processes.
To understand the DNA sequence, the Poetsch lab has built the DNA language model GROVER (Genome Rules Obtained Via Extracted Representations)1, a foundation model that learns language rules in the human genome. GROVER can be fine-tuned to investigate how different genome biology is dependent on the DNA sequence and what other processes contribute, for example in the context of DNA double strand break distribution in the genome2 or DNA replication timing in the cell cycle3. A current focus of the group lies with causes and consequences of small adduct DNA damage distribution in the genome4–6 and on dissecting cellular complexities with clonal reconstructions, multi-omics and single cell data integration strategies7,8.
2. Research questions addressed by the group:
Combining the above-mentioned methodology, the group currently focusses on three main research questions:
What determines how DNA damage distributes over the genome in different tissues and conditions and with which consequences?
How do somatic mutations arise in their genomics context and how do they change genome function?
How does the genome encode its different functions beyond protein coding and gene regulation and how does it create balance or loose it during ageing?
3. Possible project(s):
Somatic mutations distribute heterogeneously over genomes, dependent on immediate sequence context, the epigenome, and non-genetic factors. Mutations build the foundation for somatic genome evolution towards age-related diseases. We have built EAGLE-MUT, a DNA language model trained on sparse single base substitution data from individual somatic cancer genomes. Its predicted mutation probabilities show variation across several orders of magnitude, depending on their sequence context. We have defined mutational cold- and hotspots and find distinct up to 200 base pair DNA motifs, associated for example with nucleosome positioning. EAGLE-MUT provides new links between DNA sequence, genome function, and stability in individual samples, a foundation for understanding somatic genome evolution, and more personalized strategies for risk-assessment, prevention, and treatment.
This project aims to investigate personalized genetic predisposition to somatic mutagenesis by diversifying the genomes used for output generation. Carrying different germline single nucleotide polymorphisms (SNP) can have an impact on somatic mutagenesis. This may in different tissues lead to different frequencies of driver mutations for somatic mosaicism, like clonal hematopoiesis, or cancer development. This potentially leads to individually diverse timing for somatic cell evolution and represents a so far unappreciated route of individual differences for ageing and ageing-related disease development. We have so far found SNPs that have the potential to raise the probability of specific driver mutations in the TP53 gene by > 50%, despite the SNP and the driver mutation being > 40 base pairs apart.
The project aims to investigate the sequence context dependent genetic predisposition to somatic mutagenesis systematically and genome-wide with four main aims:
- Aim1: Identify germline genetic variation that impacts somatic mutagenesis in multiple tissues
- Aim2: Model the impact of diverse frequencies of driver mutations on somatic mosaicism and cancer development
- Aim3: Investigate the mechanistic link between the SNP and the driver mutation
- Aim4: Validate the impact of the SNP on driver mutations through orthogonal evidence
Project 2: How local mutational burden relates to species life span
Analogously to diverse humans, the genomes of different species lead to differences in the accumulation of somatic mutations9. Different environments led to the adaptation of DNA repair pathways and genome maintenance mechanisms, which contribute to a species’ life span. It is however unexplored, whether also the susceptibility of the genome to damage and mutation has also evolved, via adaptation of base composition towards different impact of genome damage.
Using data on somatic base substitution in healthy tissues from different mammals, for example from the colon9, the project aims to investigate, whether drivers of somatic genome evolution differ between species, based on their genomes’ base composition with three main aims:
- Aim1: Building EAGLE-MUT models for somatic base substitutions in different species.
- Aim2: Investigating the impact of species-dependent differences on drivers of somatic genome evolution.
- Aim3: Model the impact of diverse frequencies of driver mutations on somatic mosaicism and investigate relationships with species life span.
4. Applied Methods and model organisms:
- Cancer Genomics
- Somatic tissue sequencing data
- Machine Learning
- Deep Learning
- Human and possibly multi-species data
5. Desirable skills and qualifications:
- Computational Biology
- Good understanding of Genome Biology
- Experience with Machine Learning
- Good programming skills in R and/or Python
6. References:
Sanabria, M., Hirsch, J., Joubert, P. M. & Poetsch, A. R. DNA language model GROVER learns sequence context in the human genome. Nature Machine Intelligence 1-13 (2024). www.nature.com/articles/s42256-024-00872-0
Joubert, P. M. & Poetsch, A. R. Using the DNA language model, GROVER, to parse sequence and epigenetic effects on genome stability. bioRxiv 2025.07. 23.666402 (2025). www.biorxiv.org/content/10.1101/2025.07.23.666402.abstract
Janakievski, N., Joubert, P. M. & Poetsch, A. R. Using Deep Learning to predict replication timing reveals baseline control of genomic DNA sequence. bioRxiv (2026). www.biorxiv.org/content/10.64898/2026.07.18.739304.abstract
Poetsch, A. R., Boulton, S. J. & Luscombe, N. M. Genomic landscape of oxidative DNA damage and repair reveals regioselective protection from mutagenesis. Genome Biol19, 215 (2018). pubmed.ncbi.nlm.nih.gov/30526646
Poetsch, A. R. The genomics of oxidative DNA damage, repair, and resulting mutagenesis. Comput Struct Biotechnol J18, 207-219 (2020). pubmed.ncbi.nlm.nih.gov/31993111
Takhaveev, V. et al. Click-code-seq reveals strand biases of DNA oxidation and depurination in human genome. Nat Chem Biol22, 716-727 (2026). pubmed.ncbi.nlm.nih.gov/41174235
Dang, T. T. V. et al. Tumor clone dynamics in gastro-esophageal cancer organoids reveal a non-genetic memory of neoadjuvant chemotherapy via downregulation of NFκB signaling. bioRxiv (2025). www.biorxiv.org/content/10.1101/2025.07.29.667467
Cosacak, M. I. et al. Baseline regulatory programs in larval and adult neural progenitors converge towards an injury-induced state after spinal cord injury (Cold Spring Harbor Laboratory, 2026). www.biorxiv.org/content/10.64898/2026.07.24.740552v1
Cagan, A. et al. Somatic mutation rates scale with lifespan across mammals. Nature604, 517-524 (2022). pubmed.ncbi.nlm.nih.gov/35418684
