Computational Biology & AI

From raw data to a validated biological signal

Our computational biology team handles the dry-lab side of every project, from bioinformatics pipelines and statistical analysis to AI-driven models designed to uncover the signals that matter in your data. Every result is interpreted by our scientists and connected back to the biology.

What the computational team covers

Bioinformatics pipelines

Alignment, quantification and QC, version-controlled end to end.

  • Expression: bulk RNA-seq, 3’ mRNA-seq, small RNA and miRNA, single-cell & spatial transcriptomics.
  • Genome and variants: WGS, WES, targeted panels, copy number and structural variants, long-read sequencing.
  • Epigenome: DNA methylation (arrays, RRBS and WGBS), ChIP-seq, ATAC-seq.
  • Microbiome: 16S and ITS amplicon, shotgun metagenomics, metatranscriptomics.
  • Targeted and non-sequencing: qPCR and dPCR, LC-MS/MS proteomics.
  • Reproducible, containerised runs from raw reads to results.

Techniques not listed above are scoped on request. The question we answer first is whether the data and the design support the analysis, not whether we run the instrument.

Multi-omics integration

Two or more layers analysed together on the same samples.

  • Joint analysis across genomics, transcriptomics, proteomics and metagenomics.
  • Batch correction and cross-platform harmonisation, including layers generated at different sites or at different times.
  • Multi-layer correlation and network analysis.

Biomarker discovery

Identifying the markers worth taking to the next stage.

  • Feature selection to reduce thousands of candidates to a testable set.
  • Classification, regression and representation learning, adapted to the constraints of omics data.
  • Robustness testing on unseen data.

Analysis by research area

How this computational work applies in each sector.

Have a dataset that needs interpreting?

Let's talk about your data and your biological question.

Write to us

Questions we get asked

Working from your data, what you receive, and how a signal is held to be real.

Can you work on data we generated elsewhere?

Yes, and it is a large part of what the computational team does. We start from your raw or processed files, re-run the quality control so that the starting point is documented, and tell you what the data can and cannot support before proposing an analysis.

What do we receive at the end?

A report that answers the biological question, the processed data and result tables, and the figures. The pipeline versions and reference genome are stated so the analysis can be repeated. Anything beyond that (code, notebooks, an interactive dashboard) is agreed in the study documentation rather than assumed.

How do you avoid reporting a signal that does not hold up?

By deciding the analysis plan before looking at the outcome, and by separating discovery from validation. A signal found on one set is tested on data that did not contribute to finding it. Where the cohort is too small for that, we say so and report the finding as exploratory, which is a result, not a failure.

Which reference genome and pipeline versions do you use?

They are fixed per study and written down, because two analyses run on different annotation versions are not comparable and the difference does not show in the report. We state the reference build, the annotation release and the pipeline version with the results.