Skip to content
Sufyan Suleman

Academic journey

My own account, in order, of moving from pharmaceutical quality control in Lahore through two labs in Oulu and a PhD in Copenhagen to statistical genetics of ADHD and obesity at Aarhus University and the Broad Institute.

Updated September 2026

I am a computational biologist and geneticist who started at the bench. This is how I got from a quality-control laboratory in Lahore, through protein science and two labs in Oulu, to statistical genetics in Copenhagen and Aarhus.

Lahore, biotechnology and pharmaceutical quality control (2006-2013)

I began with a BSc (Hons) in Biotechnology at Government College University Lahore, from 2006 to 2010. From 2010 to 2013 I worked as a quality control inspector at a pharmaceutical company in Lahore: microbiological and chemical testing of raw materials and finished drugs, plus the GMP documentation that accompanies every batch. It was my first real exposure to the discipline that regulated data demands, checking and documenting a result before trusting it, and that habit followed me into every lab I worked in afterwards.

Oulu, protein science and two labs at once (2014-2019)

Studying protein science

In September 2014 I moved to Finland on a fully funded scholarship for the International Master’s Degree Programme in Protein Science and Biotechnology at the University of Oulu, with Biochemistry as my major subject, completing 123 ECTS credits and graduating on 18 May 2017. I flew north from a city where summers reach plus 45 degrees Celsius into a Finnish winter where the temperature drops to minus 35, and the first year carried real culture shock alongside that physical adjustment. Every year in Finland also came with a residence permit renewal, a yearly reminder that my right to stay and study was never entirely settled.

The coursework was a broad grounding in protein science: protein production and analysis, biochemical methodologies, protein crystallographic methods, structural enzymology, the biochemistry of protein folding, and an introduction to structure based drug discovery, alongside a systems biology course taught by Gong-Hong Wei, tumour cell biology, virology, yeast genetics and molecular biology. On the practical side it meant hands-on biomolecular work: recombinant expression of proteins in bacteria, protein production and purification, and downstream processing.

Four labs

Biocenter Oulu ran its MSc research training through short rotations, and I went through four labs. In Peppi Karppinen’s lab I worked on HIF-1α expression and contributed to mouse experiments with Anna Laitakari, my first real contact with in vivo work. In Johanna Myllyharju’s lab, whose group works on extracellular matrix and collagen modifying enzymes, I worked on extracellular matrix proteins. In Seppo Vainio’s lab I optimised protocols for extracellular vesicle isolation, my first contact with the lab I would return to a year later. In Gong-Hong Wei’s lab the rotation became the start of my move into cancer genomics and turned into my MSc thesis, “Functional annotation of prostate cancer risk variants identifies RGS17 as a disease susceptibility gene,” submitted as my pro gradu in 2017.

Cancer genomics with Gong-Hong Wei

After the MSc I stayed in Oulu as a research assistant in the Wei lab, working on the functional genomics of prostate and breast cancer susceptibility loci. My role changed paper by paper. On the 19q13 aggressive prostate cancer susceptibility locus paper, published in Cell in 2018, I am credited among the authors who assisted with and performed experiments, broad support rather than one named assay. On the SNPs-seq prostate cancer screening paper, published in Nature Communications in 2018, I am credited as having performed the experiments, running SNPs-seq, a high-throughput allele-specific protein-binding screen, alongside EMSA and CRISPR/Cas9-based gene editing. On the breast cancer exome-wide association study, published in Cancer Research in 2018, my contribution was development of methodology, statistical and biostatistical analysis, and manuscript writing and revision.

The lab as a whole used ChIP-qPCR and ChIP-seq reanalysis, luciferase enhancer reporter assays, siRNA and shRNA knockdown via lentiviral constructs, 3D Matrigel culture, 3C-qPCR, and Sanger and allele-specific qPCR genotyping. During this period I also wrote a first-author commentary on combined immunotherapy for advanced prostate cancer, published in the Asian Journal of Urology in 2017.

Back to the Vainio lab: Trisk 95

In the summer of 2018 I went back to Seppo Vainio’s lab, this time with Nsrein Ali, on a project asking whether a skin protein could mirror systemic blood glucose. This became the most hands-on wet-lab work of my career. My specific, credited contribution was the Triadin, or Trisk 95, promoter luciferase reporter study behind the Scientific Reports paper published in 2020: restriction-enzyme molecular cloning of the promoter fragment into a pGL3 vector, transient transfection using Turbofect, and the Dual-Luciferase reporter assay. The wider project involved primary mouse keratinocyte isolation and culture, qRT-PCR, Western blotting, immunofluorescence and confocal microscopy, live-cell calcium imaging, flow cytometry, transmission electron microscopy, an in vivo STZ-induced type 1 diabetic mouse model with glucose tolerance testing, and ELISA. That mouse model work meant formal training in animal procedures under the European directive for animal handling, training I still hold, though I no longer want to be the person at the bench and am comfortable overseeing others instead. The lab’s broader programme extended into stem-cell and genome-editing biosensor technology, connected to a related patent, without me as a named inventor.

Copenhagen, research assistant to PhD at CBMR (2019-2024)

This is the period I look back on as the most exciting stretch of my career so far.

Arriving with no code

In September 2019 I moved to Denmark to join the Novo Nordisk Foundation Center for Basic Metabolic Research, CBMR, at the University of Copenhagen, as a research assistant working with Niels Grarup and Torben Hansen. I arrived with essentially no coding skills; everything I know now in R, in GWAS pipelines, in handling large genetic datasets, was learned on the job in that group. It felt like a completely new world with no coding skills, compared to the wet-lab biochemistry and cancer genomics I had come from in Oulu.

Our computing ran on the Centre’s own cluster, first called Porus and later renamed Esrum, together with the national resource Computerome. COVID hit partway through this research assistant period, changing how the group worked day to day, though the research on insulin-resistance phenotypes continued.

Three applications

Getting into a PhD took three tries. I applied for a fellowship from the Novo Nordisk Foundation and was rejected, applied again to the Danish Diabetes Academy and was rejected again, then took what I had learned from both applications into a third, to the Danish Independent Research Fund, DFF. People said DFF was the hardest of the three. It was the one that was funded. My PhD started in June 2021, co-funded by the Novo Nordisk Foundation and the DFF, with Niels Grarup and Torben Hansen as my supervisors.

The PhD

My thesis was titled “Genetics of insulin sensitivity and implications for T2D, obesity and cardiovascular disease.” At its core was a genome-wide association study of 21 distinct insulin sensitivity indices across roughly 5,000 Danish individuals, alongside heritability and genetic and phenotypic correlations, presented at Metabolism Day 2023 at the University of Copenhagen.

Three first-author papers followed, in order. The first, in the Journal of Clinical Endocrinology and Metabolism in 2024, “Genetic underpinnings of fasting and oral glucose-stimulated based insulin sensitivity indices” (doi 10.1210/clinem/dgae275), showed that fasting and post-glucose insulin sensitivity indices have distinct genetic architectures. That finding later fed into an interactive browser I built, searching effect sizes of T2D risk variants, from Suzuki et al., Nature, 2024, across 23 insulin sensitivity indices. The second, also JCEM 2024, “Adult-based genetic risk scores for insulin resistance associate with cardiometabolic traits in children and adolescents” (doi 10.1210/clinem/dgae882), built three genetic risk scores for insulin-resistance subtypes and validated them in the Holbæk paediatric cohort, with the supplementary data released openly. The third, in Scientific Reports in 2025, “Exploring the genetic intersection between obesity-associated genetic variants and insulin sensitivity indices” (doi 10.1038/s41598-025-98507-w), examined where the genetics of obesity and insulin sensitivity overlap and diverge.

Consortia, courses and Steno

Alongside my thesis, I contributed to the MAGIC consortium’s GWAS of glycaemic and insulin-sensitivity traits, drawing on roughly 20,000 to 25,000 Danish participants and published in Nature Genetics in 2023, and to the SpiroMeta consortium’s age-stratified lung function GWAS: 120 analyses across five age bands, eight cohorts and three phenotypes, still in preparation. I was a co-applicant on successful grants to the Novo Nordisk Foundation, the Danish Diabetes and Endocrine Academy, and the DFF, and coordinated GDPR data-access applications for UK Biobank and Generation Scotland. Formal training through this period included Basic Statistics for Health Researchers, an introduction to statistical analysis of GWAS, the Wellcome Connecting Science course on computational systems biology, the Danish Diabetes Academy’s Basal Cardio-Metabolic course and Summer School, intermediate and advanced Reproducible Research in R, and responsible conduct of research. In April 2024 I spent two months as a visiting researcher at Steno Diabetes Center Copenhagen, on proteomics and genomics of diabetic complications.

The last months

In March 2024, Niels Grarup left CBMR for a position elsewhere. The group was winding down, with two other PhD students finishing in parallel, so the final months were thin on day-to-day supervision. I asked Torben Hansen for two extra months to finish the writing and start looking for what came next, and he gave me that time.

At that point I had two realistic postdoc options: one in Sweden, which would have meant living in Copenhagen and commuting across the Øresund, an arrangement my residence situation did not allow; the other in Aarhus, with Ditte Demontis and Anders Børglum, on the shared genetics of ADHD and obesity. I chose Aarhus.

Choosing Aarhus

I moved to Aarhus in September 2024 to start that postdoc, and returned to Copenhagen in November 2024 to defend my thesis, with the degree conferred in December 2024. The obesity and insulin sensitivity overlap I had spent three years mapping had already raised a related question: obesity’s genetic and phenotypic links reach well beyond cardiometabolic disease, and ADHD is strongly comorbid with it. The genetic tools I had built during the PhD, polygenic scoring, cross-cohort validation, consortium-scale GWAS, applied directly to that new question.

Aarhus and the Broad, shared genetics of ADHD and obesity (2024-present)

Since September 2024 I have held a joint position: postdoctoral researcher in the Department of Biomedicine at Aarhus University, working with Ditte Demontis and Anders Børglum, and affiliate scientist at the NNF Center for Genomic Mechanisms of Disease at the Broad Institute of MIT and Harvard. Both are ongoing. The question at the centre is the shared genetic architecture and biology between ADHD and obesity, conditions that are epidemiologically linked but not well understood mechanistically at the genetic level.

I waited around a year for data access agreements before I could start the substantive work, then delivered the analysis within about a year of access being granted.

This project is where I learned most of the statistical genetics toolkit I now use, largely from a standing start. For polygenic architecture and genetic correlation: MiXeR, LAVA and LDSC. For cross-trait meta-analysis: ASSET and METASOFT, alongside MungeSumstats, a Bioconductor package I contribute to, for harmonisation and genome-build alignment. For locus annotation and gene prioritisation: FUMA, MAGMA, FINEMAP, PoPS and FLAMES. For single-cell work: scDRS, building gene sets from FLAMES-prioritised genes weighted by MAGMA scores. For polygenic scoring: SBayesRC, feeding PheWAS analyses across large phecode sets. I also work with LipocyteProfiler data from the Broad Institute, a high-content imaging assay measuring cellular morphology, to ask whether shared genetic risk shows up in cellular biology as well as statistical association. All of this runs on GenomeDK, the Aarhus University HPC cluster, distinct from the Centre’s infrastructure I had used in Copenhagen.

I built and lead a course called Integrative Post-GWAS Methods, registered in the Aarhus University PhD catalogue, coordinating a team of contributing teachers and bioinformaticians. The ADHD and obesity work was presented as an oral presentation at the ESHG Annual Meeting in 2025, and I will present it again at the World Congress of Psychiatric Genetics in Glasgow in September 2026.

Building tools and teaching

Throughout these years, alongside the research, I built and maintained open-source tools and teaching materials.

Computing 21 or more insulin sensitivity indices by hand, consistently, across cohorts, is slow and error-prone. Re-deriving the same formulas during the PhD pushed me to write InsuSensCalc, which started as a lab tool and grew into a package published on CRAN. HealthMarkers, a broader clinical and metabolic biomarker toolkit, and later IsletCalc, calculating insulin release, glucagon release and glucagon resistance indices, followed the same logic. I also contribute to MungeSumstats, a widely used package for GWAS summary statistics standardisation.

Alongside Integrative Post-GWAS Methods, I authored two open, Zenodo-archived courses from scratch: Getting Started with R and Basic Statistics for Researchers. I am building CDISC with R, teaching CDISC clinical data standards, SDTM, ADaM, Define-XML, Dataset-JSON, and tables, listings and figures, using a simulated Phase III trial; it is still in development and I begin teaching it in late September 2026. Separately, I built a demonstration data-management and central-monitoring pipeline modelled on an adaptive platform trial at Rigshospitalet, an independent educational project using entirely synthetic data, not commissioned work. The pipeline conforms heterogeneous exports from 25 simulated sites, holds its validation rules as configuration rather than code, and scores its own detection performance against injected defects with known ground truth. Earlier, during the PhD, I was a teaching assistant for two Novo Nordisk Foundation funded Danish Diabetes Academy R courses in 2022, and I run recurring Study Cafe genetics sessions for 50 to 60 bachelor medical students at a time. I have supervised four bachelor medical students through full research-project cycles.

Laboratory technique inventory

The table below covers only the molecular biology and wet-lab work, listed per paper with the role credited in each author-contribution statement. The computational and statistical-genetics methods are described in the Copenhagen and Aarhus sections above.

Paper Lab Credited role Techniques
Ali N et al., Trisk 95 as a novel skin mirror for normal and diabetic systemic glucose level, Scientific Reports 2020 Seppo J. Vainio, Laboratory of Developmental Biology, Biocenter Oulu Performed the Triadin promoter luciferase reporter study Restriction-enzyme molecular cloning (promoter fragment into pGL3 vector); transient transfection (Turbofect); Dual-Luciferase reporter assay; primary mouse keratinocyte isolation and culture; qRT-PCR; Western blotting; immunofluorescence staining and confocal microscopy; live-cell calcium imaging (Fluo-4 AM); flow cytometry; transmission electron microscopy; in vivo STZ-induced type 1 diabetic mouse model, glucose tolerance testing, IP injections; ELISA (rodent insulin); metabolite quantification (YSI 2950); genetic knockout mouse model use (Trisk 95 KO); GraphPad Prism statistical analysis
Gao P et al., Biology and clinical implications of the 19q13 aggressive prostate cancer susceptibility locus, Cell 2018 Gong-Hong Wei, Biocenter Oulu Assisted / performed experiments (broad experimental support) CRISPR/Cas9 genome editing (incl. single-nucleotide editing, CRISPR interference/activation); ChIP-qPCR and ChIP-seq data reanalysis; luciferase enhancer reporter assays; siRNA transfection; shRNA knockdown with lentiviral constructs; 3D Matrigel cell culture and phenotyping; chromosome conformation capture (3C-qPCR); Sanger sequencing and allele-specific qPCR genotyping; cancer cell line culture with mycoplasma testing; proliferation/migration/invasion assays
Zhang P et al., High-throughput screening of prostate cancer risk loci by single nucleotide polymorphisms sequencing, Nature Communications 2018 Gong-Hong Wei, Biocenter Oulu Performed the experiments SNPs-seq (high-throughput allele-specific protein-binding screen); EMSA (electrophoretic mobility shift assay); gene reporter assays; CRISPR/Cas9-based gene editing
Zhang B et al., A large-scale, exome-wide association study of Han Chinese women identifies three novel loci predisposing to breast cancer, Cancer Research 2018 Gong-Hong Wei, Biocenter Oulu Methodology development; statistical/biostatistical/computational analysis; manuscript writing and revision Exome-array coding-variant association analysis; eQTL analysis; protein-protein interaction and pathway enrichment analysis (STRING)

The through-line

Looking back across Lahore, Oulu, Copenhagen and Aarhus, the pattern is less about any single disease area and more about repeatedly starting again: pharmaceutical QC to protein biochemistry, wet-lab cancer genomics to statistical genetics, insulin sensitivity to ADHD, each move meaning a new country, lab or field, and its toolkit learned from zero.

What I am working towards

My research direction is to understand shared genetic and biological mechanisms across conditions usually studied separately, most immediately the metabolic and psychiatric comorbidity between obesity and ADHD, but the same integrative approach applies to other cross-disease overlaps such as depression and BMI. I want to keep combining large-scale statistical genetics, polygenic architecture, causal inference, fine-mapping, gene prioritisation, with functional and cellular readouts, so that genetic findings connect to a plausible biological mechanism rather than stopping at a locus. Alongside that, I intend to keep building open, reproducible software and teaching materials, because methods that are not usable or auditable by others do not move a field forward on their own.

If any of this overlaps with your work, I would like to hear from you. Email me