Contents Menu Expand Light mode Dark mode Auto light/dark, in light mode Auto light/dark, in dark mode Skip to content
ALF documentation
ALF documentation

Documentation

  • Explanation
    • Why ALF?
    • Intro to Active Learning
    • Core Concepts
  • Tutorials
  • How-to / Recipes
    • Add your own model
    • Add a dataset
    • Add an acquisition function
    • Add a search function
    • Switch offline to online
    • Run with Docker
  • Reference
    • Glossary
  • ALF Installation Guide

API Reference

  • Core
    • Dataclasses
      • Candidate
      • Labelled Candidates
      • Predictions
      • Results
      • Round Metrics
      • State
      • Surrogate Epoch Metrics
    • Dataset
      • Base Dataset
      • Splitting Utils
    • Model
      • Base Model
      • Normalisers
    • Optimizer
      • Optimizer
      • Acquisition Function
      • Search
    • Oracle
    • Surrogate
    • Tasks
      • Base Task
      • Design Task
      • Supervised Task
      • Zero-Shot Task
    • Utils
      • Enums
      • Metrics
        • Acquisition Batch
        • Base
        • Calibration
        • Classification
        • Design Task Metrics
        • Regression
      • State Logger
  • Tools
    • Datasets
      • GFP Dataset
      • ProteinGym Dataset
      • FLIP Dataset
      • GuacaMol Dataset
      • Matbench Dataset
    • Models
      • CNN Model
      • MLP Model
      • ESMFold Model
      • GP Model
      • ESM2 Model
      • Chemprop Model
      • MLIP Model
      • PyRosetta Model
      • Ensemble Wrapper
      • GuacaMol Oracle
      • Utilities
    • Optimizer
      • Acquisition Functions
        • BoTorch Acquisition Wrapper
        • BoTorch MC Samplers
        • CoreSet
        • Expected Improvement
        • Greedy
        • Random
        • Thompson Sampling
        • Uncertainty Sampling
        • Upper Confidence Bound (UCB)
      • Search Strategies
        • Botorch Continuous Search
        • Single Mutant Search
        • SMILES Mutation Search
        • Element Substitution Search
    • Utils
      • Constants
Back to top
View this page

Single Mutant Search¶

Single Mutant Search is a mutation-based search strategy that generates candidate pools by applying single-point mutations to existing sequences. This is particularly useful for local exploration around known high-performing sequences in protein engineering.

The top_k parameter controls how many training sequences are mutated. With the default of 1, only the best-labelled sequence is mutated, so the search explores a single neighbourhood at a time. With a larger value, the top_k best-labelled sequences are each mutated and the results pooled together, so the search covers several local optima at once rather than stalling when no neighbour of the current best improves. A mutant reachable from more than one sequence appears only once in the pool. Sequences with equal labels are ranked by their position in the training set (earliest first), so top_k=1 always selects the same sequence as single-best selection.

class alf_tools.optimizer.search.single_mutant_search.SingleMutantSearch(alphabet='ARNDCQEGHILKMFPSTWYV', top_k=1)[source]¶

Bases: SearchProtocol

Search protocol that enumerates single-point mutants of the top-K training sequences.

For each of the top_k highest-labelled training sequences, every single-position substitution over alphabet is enumerated. Different seeds can produce the same mutant, so duplicates are dropped while keeping the first one generated, which makes the pool order deterministic. top_k=1, the default, is a pure hill-climb on a single neighbourhood; higher values keep several local optima under exploration at once.

Next
SMILES Mutation Search
Previous
Botorch Continuous Search
Copyright © 2026, InstaDeep
Made with Sphinx and @pradyunsg's Furo
On this page
  • Single Mutant Search
    • SingleMutantSearch