Contents Menu Expand Light mode Dark mode Auto light/dark, in light mode Auto light/dark, in dark mode Skip to content
ALF documentation
ALF documentation

Documentation

  • Explanation
    • Why ALF?
    • Intro to Active Learning
    • Core Concepts
  • Tutorials
  • How-to / Recipes
    • Add your own model
    • Add a dataset
    • Add an acquisition function
    • Add a search function
    • Switch offline to online
    • Run with Docker
  • Reference
    • Glossary
  • ALF Installation Guide

API Reference

  • Core
    • Dataclasses
      • Candidate
      • Labelled Candidates
      • Predictions
      • Results
      • Round Metrics
      • State
      • Surrogate Epoch Metrics
    • Dataset
      • Base Dataset
      • Splitting Utils
    • Model
      • Base Model
      • Normalisers
    • Optimizer
      • Optimizer
      • Acquisition Function
      • Search
    • Oracle
    • Surrogate
    • Tasks
      • Base Task
      • Design Task
      • Supervised Task
      • Zero-Shot Task
    • Utils
      • Enums
      • Metrics
        • Acquisition Batch
        • Base
        • Calibration
        • Classification
        • Design Task Metrics
        • Regression
      • State Logger
  • Tools
    • Datasets
      • GFP Dataset
      • ProteinGym Dataset
      • FLIP Dataset
      • GuacaMol Dataset
      • Matbench Dataset
    • Models
      • CNN Model
      • MLP Model
      • ESMFold Model
      • GP Model
      • ESM2 Model
      • Chemprop Model
      • MLIP Model
      • PyRosetta Model
      • Ensemble Wrapper
      • GuacaMol Oracle
      • Utilities
    • Optimizer
      • Acquisition Functions
        • BoTorch Acquisition Wrapper
        • BoTorch MC Samplers
        • CoreSet
        • Expected Improvement
        • Greedy
        • Random
        • Thompson Sampling
        • Uncertainty Sampling
        • Upper Confidence Bound (UCB)
      • Search Strategies
        • Botorch Continuous Search
        • Single Mutant Search
        • SMILES Mutation Search
        • Element Substitution Search
    • Utils
      • Constants
Back to top
View this page

SMILES Mutation Search¶

The molecule-domain counterpart to Single Mutant Search: a search protocol that proposes novel SMILES each round by mutating the top-K best-labelled training molecules, filtered for RDKit validity. Unlike protein sequences, most single-character SMILES edits are structurally invalid, so this filtering step is required. RDKit’s error logger is silenced, since invalid mutations are expected and handled internally.

Mutating only the single current best (top_k=1) makes the search a pure hill-climb: once no neighbour of the incumbent beats it, the same neighbourhood is regenerated every round and the loop stalls in that local optimum. Setting top_k higher keeps several regions of the search space open at once, so a stall in one neighbourhood doesn’t stall the whole search.

Next
Element Substitution Search
Previous
Single Mutant Search
Copyright © 2026, InstaDeep
Made with Sphinx and @pradyunsg's Furo