Capabilities
What I can do with human genetic and clinical data, from statistical analysis in large cohorts to the functional experiments that test what an association means, for research groups, biotech and pharma teams, and companies building platforms on this kind of evidence.
01Human genetics at scale
Human genetics is my primary domain. I work with family studies, population cohorts and consortium-scale datasets, and with the methods that turn them into defensible estimates, heritability, genetic and phenotypic correlation, common and rare variant association testing, pleiotropy across traits, and linkage disequilibrium modelling and fine-mapping.
I have run analyses in cohorts ranging from a few thousand participants to half a million, and contributed age-stratified and trait-level meta-analyses to international consortia. Study design, phenotype definition and quality control are part of that work, not a preliminary to it.
02From association to target
An association is a starting point, not an answer. I take signals through fine-mapping, gene prioritisation and functional annotation to a ranked shortlist, with the reasoning for each rank stated, and I distinguish candidates the evidence supports from candidates it merely fails to exclude.
The same stack supports polygenic scores, phenome-wide scans across large phecode sets and single-cell disease relevance, so a target can be examined from several directions rather than one.
03Biomarkers from routine clinical data
I write and maintain the R packages that turn a routine clinical table into several hundred documented markers and indices per participant, over 290 validated biomarkers in HealthMarkers, 38 insulin sensitivity indices across fasting, OGTT, adipose and tracer families in InsuSensCalc, and surrogate islet function indices in IsletCalc.
Each formula carries its published source and its unit conventions, which is what makes the output auditable rather than merely computed. The packages are open source and can be embedded as a library inside another platform or pipeline.
04Disease areas I know as data
Metabolic disease is the core of my own research, insulin sensitivity and its subtypes, obesity, type 2 diabetes and cardiometabolic traits. Psychiatric genetics is the second, ADHD and depression, and the overlap between them and body weight.
I know these areas as datasets and phenotype definitions, not only as literature, which is usually what decides whether an analysis is possible at all.
05Populations beyond the discovery sample
My work so far is in European ancestry cohorts, and I treat that as a limitation rather than a default. Prediction accuracy falls with genetic distance from the sample a tool was built in, so a score or a threshold carried into a new population needs to be re-estimated, not reused.
The same applies to clinical cut-offs. Population-specific thresholds can be implemented in the biomarker packages where the reference data exist, and where they do not, that is the first thing worth saying.
06Data, pipelines and reproducibility
I build analyses that another person can run again and get the same answer, with dependencies pinned, steps declared, rules held as configuration rather than buried in code, and tests that fail when the data change shape.
This extends to clinical data standards and to the paperwork that precedes the analysis, data access applications, GDPR documentation, and the quality control that decides whether a dataset is usable.
07Molecular biology and functional genomics
I began at the bench, and that is where my understanding of what a genetic association can mean was formed. In the Wei laboratory in Oulu I worked on the functional characterisation of prostate and breast cancer susceptibility loci, including allele-specific protein binding screens, reporter assays, CRISPR and Cas9 editing, chromatin immunoprecipitation and expression analysis, contributing to papers in Cell, Nature Communications and Cancer Research.
In the Vainio laboratory I ran a promoter reporter study on a candidate skin biomarker of systemic glucose, from molecular cloning to the dual luciferase assay, within a project that extended to primary cell culture, imaging and an in vivo diabetic mouse model.
I no longer work at the bench, and I state that plainly. What remains is the ability to read experimental evidence critically, to judge whether a proposed validation experiment will answer the question, and to design a computational analysis that an experimentalist can act on.
How I work
- Say first what the data can and cannot support, before proposing what to do with it.
- Separate missing evidence from negative evidence. They are not the same claim.
- Report the estimate, its uncertainty and the sample size. A p-value on its own is not a result.
- Leave the analysis in a state where someone else can run it, read it and disagree with it.
Where to see the work
Working on something that needs this?
I am open to collaboration, consulting and applied projects with research groups and companies working on human genetic and clinical data.