Models that read the genome.

XS1 Biosciences researches genomic language models — models trained on DNA and RNA sequence — and uses them to study how a change in sequence changes what a gene does.

Illustrative

In-silico mutagenesis: every single-base change, scored

Illustrative in-silico mutagenesis map over 48 positions. Larger predicted effects appear in the regulatory motif and at the splice sites around the exon; third codon positions inside the exon score lower. Values are generated for illustration, not model output.

Hover a cell to read it. Each column is one position; each row, the base it is changed to.

Illustrative values, not model output.

Variant-effect research asks which changes matter. A model can score every possible single-base change in a region at once; the pattern points researchers to the positions most worth testing.

Designing a change at one of those positions? See Genome editing

01What the work covers

From reading DNA to predicting function.

Four strands of research, from models that learn the language of the genome to models of the cells it runs in.

Genomic language models

Models pretrained on DNA and RNA from many genomes, reading long stretches of sequence to learn the patterns that control genes.

Variant effect prediction

Estimating whether a variant is likely to change function, including zero-shot scores from a pretrained model, as a signal for which variants to study next.

Regulatory DNA

Predicting how sequence in promoters and enhancers shapes when and how much a gene is expressed, and how splice sites shape which transcripts are made.

Single-cell and multi-omics

Models over gene expression and other measurements taken cell by cell, for grouping cells, integrating datasets and predicting responses to perturbation.

02Single-cell models

From genome to cell.

Single-cell data measures gene expression one cell at a time. Models trained on it can tell cell states apart, bring experiments together, and estimate how cells might respond before an experiment is run.

  • 01

    Group cells by state

    Cells with similar expression profiles cluster together, whatever experiment they came from.

  • 02

    Integrate datasets

    Align measurements from different labs, platforms and modalities into one shared space.

  • 03

    Estimate a perturbation's effect

    Predict how expression might change after a gene is knocked down or a compound is added, to choose which experiments to run.

Illustrative

Cells grouped by state

Each dot is one cell, placed by its gene expression. Illustrative positions, not a dataset.

03How outputs are used

Predictions guide experiments. They do not replace them.

Signals for prioritization

Predictions help decide which variants, regions or perturbations are worth testing first.

Checked against measurement

Where experimental data exists, such as deep mutational scans, predictions are compared against it before anyone relies on them.

Research use only

Outputs support research. They are not intended for decisions about individuals.

Work with XS1 Biosciences

Bring XS1 Biosciences a question about the genome.

A variant set to prioritize, a regulatory question, a single-cell dataset to model: describe it in general terms and we will follow up.

Please do not send unpublished data, sequences or other confidential material through this website.