From sequence to structure, and back.

XS1 Biosciences researches protein language models that learn structure and function from sequence, and generative methods that propose new protein sequences in silico for partner labs to test.

Illustrative

Sequence → structure

Illustrative: a 42-residue chain folds into three antiparallel strands and a helix. Its contact map shows the strand pairings as anti-diagonal stripes and the helix as a band along the diagonal.

Contact map · residue × residue

An unfolded chain: just a sequence.

01What the work covers

Reading proteins, and proposing new ones.

Prediction runs from sequence to structure and function. Design runs the other way: from a structure or function a researcher wants, back to a sequence that might have it.

01

Protein language models

Models trained on large collections of protein sequences. They learn which positions matter for structure and function, and can score mutations without task-specific training.

  • Zero-shot mutation scoring
  • Embeddings for downstream models
02

Structure prediction

Estimating a protein's three-dimensional structure from its sequence, with a confidence for each region, so researchers know which parts of a prediction to rely on.

  • Per-region confidence
  • Complexes and interfaces
03

Generative design

Proposing new protein sequences in silico — binders, enzymes and scaffolds — conditioned on a target structure or function, for researchers to evaluate.

  • Binders and scaffolds
  • Enzyme design
04

Fitness and variant ranking

Ranking variants of a protein for a property of interest, checked against deep mutational scanning and other experimental data where it exists.

  • Deep mutational scan benchmarks
  • Ranking with stated uncertainty

02How a protein model learns

Twenty letters. Nearly every protein is written in them.

Many protein language models learn by filling in hidden residues. A model that has learned which residues fit where can also say how surprising a mutation is.

Illustrative

Predicting a masked residue

Position 4 hidden · original residue V (Valine)

Illustrative: one position of a protein sequence is hidden and a protein language model ranks the twenty amino acids for it. The original residue ranks first and residues with similar side chains rank next.

Bars show relative scores for the hidden position. Illustrative values, not model output.

03Designed, then tested

Designed in silico. Tested in the lab.

Models propose; experiments decide. XS1 Biosciences works on the computational side and hands candidates to partner labs with the reasoning and confidence behind each one.

In silico

XS1 Biosciences

  1. 1Generate candidate sequences
  2. 2Predict structure and score fitness
  3. 3Filter and rank, with confidence

In the lab

Partner lab

  1. 1Express and purify
  2. 2Measure the property of interest
  3. 3Share results back
Measurements come back and refine the next round of models and candidates.

Work with XS1 Biosciences

Bring XS1 Biosciences a protein problem.

A family to model, a property to rank variants for, a design target: describe it in general terms and we will follow up.

Please do not send unpublished data, sequences or other confidential material through this website.