TL;DR: Key Takeaways
- Neoantigens are mutation-derived protein fragments unique to a tumour
- Algorithms rank candidates on expression, HLA presentation and predicted immunogenicity
- Only a few dozen targets survive into a manufactured panel
- Predicting presentation is far more reliable than predicting a T-cell response
- A positive trial validates the whole system, not the algorithm in isolation
Why selection is the hard part
Sequencing a melanoma typically surfaces far more candidate mutations than any vaccine can carry. The constraint is not finding mutations — it is choosing the handful most likely to be visible to that individual's immune system. That ranking problem is where machine learning is applied.
The pipeline, stage by stage
| Stage | Input | Question being answered | Main limitation |
|---|---|---|---|
| 1. Variant calling | Tumour and matched normal sequencing | Which mutations belong to the tumour rather than the person? | Sample quality, tumour purity and coverage all change what is detectable. |
| 2. Expression filtering | Tumour RNA sequencing | Is the mutated gene actually being expressed? | A mutation present in DNA may never be translated into protein. |
| 3. HLA typing | Patient sequencing | Which presentation molecules does this individual carry? | Prediction quality is uneven across less-studied HLA alleles. |
| 4. Binding and presentation prediction | Candidate peptides plus HLA type | Which peptides are likely to be processed and displayed on the cell surface? | Models are trained on available datasets and do not generalise equally to all alleles. |
| 5. Immunogenicity ranking | Predicted presented peptides | Which displayed peptides are likely to be recognised by T cells? | Presentation does not guarantee a T-cell response; this remains the weakest prediction step. |
| 6. Practical selection | Ranked shortlist | Which of these can be clonal, safe and manufacturable within the panel size? | Panel size is capped, so most candidates are discarded regardless of score. |
1. Variant calling
- Input
- Tumour and matched normal sequencing
- Question being answered
- Which mutations belong to the tumour rather than the person?
- Main limitation
- Sample quality, tumour purity and coverage all change what is detectable.
2. Expression filtering
- Input
- Tumour RNA sequencing
- Question being answered
- Is the mutated gene actually being expressed?
- Main limitation
- A mutation present in DNA may never be translated into protein.
3. HLA typing
- Input
- Patient sequencing
- Question being answered
- Which presentation molecules does this individual carry?
- Main limitation
- Prediction quality is uneven across less-studied HLA alleles.
4. Binding and presentation prediction
- Input
- Candidate peptides plus HLA type
- Question being answered
- Which peptides are likely to be processed and displayed on the cell surface?
- Main limitation
- Models are trained on available datasets and do not generalise equally to all alleles.
5. Immunogenicity ranking
- Input
- Predicted presented peptides
- Question being answered
- Which displayed peptides are likely to be recognised by T cells?
- Main limitation
- Presentation does not guarantee a T-cell response; this remains the weakest prediction step.
6. Practical selection
- Input
- Ranked shortlist
- Question being answered
- Which of these can be clonal, safe and manufacturable within the panel size?
- Main limitation
- Panel size is capped, so most candidates are discarded regardless of score.
What the models are actually predicting
- Binding: will this peptide bind the patient's HLA molecule?
- Processing and presentation: will it survive antigen processing and reach the cell surface?
- Clonality: is the mutation present in most tumour cells rather than one subclone?
- Expression: is the gene transcribed at a meaningful level?
- Immunogenicity: is a T-cell receptor likely to recognise it?
- Manufacturability: can it be produced reliably within the construct?
Where the prediction is weakest
Binding and presentation prediction is comparatively mature. Predicting whether a presented peptide will actually provoke a useful T-cell response is not. This is why programmes measure immune responses in patients rather than treating the algorithm's output as proof, and why published immune-monitoring data matters more than a described platform.
How this is validated in trials
A trial cannot isolate the algorithm. Any result reflects sequencing quality, the selection model, the construct format, the formulation, the manufacturing process and the companion therapy together. The clearest evidence that computational selection works comes from studies that measure T-cell responses against the specific predicted targets — the comparison in V940 vs EVX-01 sets out which programmes have published that link.
See the whole melanoma AI picture
Screening, monitoring, genomic analysis and personalised treatment research in one referenced hub.
A note on what this is not
This pipeline never sees a photograph. The AI used in photo-based mole checks analyses images for risk awareness and is described, with its measured limitations, in our accuracy report. The two use the word "AI" and share almost nothing else.
Frequently Asked Questions
AI is used to rank candidates. Algorithms predict which mutation-derived peptides are likely to be processed and presented by that patient's HLA molecules, scoring them on predicted binding, expression, clonality, immunogenicity and manufacturability. The AI does not create the immune response or treat the cancer; it helps researchers choose which targets to include.
A neoantigen is a protein fragment produced by a mutation that exists in tumour cells but not in the person's healthy cells. Because it is not part of normal self, the immune system can in principle recognise it as foreign.
Panels are deliberately small relative to the number of candidates found. Published programmes have described panels in the range of tens of targets — for example up to 34 in one personalised mRNA programme and up to 40 in a personalised DNA programme — selected from a much larger candidate pool.
Prediction of peptide-HLA binding and presentation is far stronger than prediction of whether a T cell will actually respond. That final step remains the hardest part of the pipeline, which is why programmes measure immune responses in trials rather than assuming them.
No. Neoantigen selection works on genomic sequence data from a tumour sample. Image-based AI, such as a photo mole check, is an entirely separate technology used for risk awareness and monitoring, not diagnosis or treatment design.