Matbench Dataset

The Matbench benchmark dataset implementation. Matbench provides 13 materials-property prediction tasks with predefined 5-fold cross-validation splits, covering both composition-based (chemical formula) and structure-based (crystal structure) inputs.

Both input kinds are stored under the MATERIALS modality. Composition values are plain chemical-formula strings (e.g. "Fe0.62C0.01Mn0.37"), used as-is. Structure values are pymatgen Structure objects, serialised to an equivalent JSON string via their MSONable .to_json() interface before being stored in Candidate.data.

Supported tasks:

Task

Input type

Problem type

Samples

Target

matbench_steels

composition

regression

312

yield strength (MPa)

matbench_expt_gap

composition

regression

4,604

experimental gap (eV)

matbench_expt_is_metal

composition

classification

4,921

is_metal

matbench_glass

composition

classification

5,680

glass-forming ability

matbench_dielectric

structure

regression

4,764

refractive index

matbench_jdft2d

structure

regression

636

exfoliation energy

matbench_log_gvrh

structure

regression

10,987

log10(shear modulus)

matbench_log_kvrh

structure

regression

10,987

log10(bulk modulus)

matbench_mp_e_form

structure

regression

132,752

formation energy

matbench_mp_gap

structure

regression

106,113

band gap (eV)

matbench_mp_is_metal

structure

classification

106,113

is_metal

matbench_perovskites

structure

regression

18,928

formation energy

matbench_phonons

structure

regression

1,265

last phonon DOS peak

problem_type is derived automatically from the task — ProblemType.REGRESSION for regression tasks, ProblemType.BINARY for the three classification tasks (all are two-class) — and does not need to be set in MatbenchConfig.

Fold mode vs. merged mode:

MatbenchConfig.fold_number selects between two ways of using Matbench’s predefined 5-fold cross-validation:

  • Fold mode (fold_number set to 0-4): uses Matbench’s predefined train/test split for that fold directly. train_ratio controls what fraction of the Matbench train rows form the initial labelled training set (the remainder becomes candidate_pool, capped at max_candidate_pool if set); validation_frac carves a validation set out of that. test_ratio and split_type are ignored — the full Matbench test set for that fold is used as test, since benchmark- comparable results require Matbench’s exact predefined test rows.

  • Merged mode (fold_number=None): all 5 folds are combined into one dataset and split using the standard ratio-based train_ratio/validation_frac/ test_ratio/split_type. This loses Matbench’s benchmark integrity guarantees — results are no longer directly comparable to published Matbench leaderboard scores.

In both modes, every candidate’s features["fold_id"] records which of Matbench’s 5 folds (0-4) it originally belonged to, for traceability.

Dependencies:

Requires the optional matbench extra (pip install "alf-tools[matbench]" or alf_tools[materials]), which installs matbench and pymatgen. Data is downloaded and cached automatically by the matbench package on first use.

class alf_tools.datasets.matbench.Matbench(config)[source]

Bases: BaseDataset

Matbench benchmark dataset class.

Matbench provides 13 materials-property prediction tasks, each with predefined 5-fold cross-validation splits, covering both composition-based and structure-based inputs. Composition and structure inputs (pymatgen Composition and Structure objects respectively) are both MSONable, so both are serialised identically via .to_json() into a JSON string stored in Candidate.data; no alf_core changes are needed. pymatgen/matbench are optional: importing this module never requires them, and constructing a MatbenchConfig raises a clear ImportError if they’re missing (see MatbenchConfig.validate_config).

config: MatbenchConfig
load_dataset()[source]

Load Matbench data via the Matbench API.

In fold mode, only the configured fold’s predefined train/test rows are loaded (each candidate tagged with a matbench_split feature of “train” or “test” for use by _split_dataset). In merged mode, all rows across all 5 folds are loaded. In both modes, every candidate’s features["fold_id"] records which Matbench fold it belongs to, and features["input_type"] records whether data is a composition formula string or a serialised structure (constant across a given task’s candidates).

Return type:

LabelledCandidates

Returns:

LabelledCandidates with JSON-string composition/structure data and float labels.

class alf_tools.datasets.matbench.MatbenchConfig(**data)[source]

Bases: BaseDatasetConfig

Configuration for Matbench benchmark datasets.

Both composition and structure task inputs are stored under Modality.MATERIALS — Matbench has no data that needs any other modality, so modality is fixed and should not be overridden. Since MATERIALS covers both a composition formula string and a JSON-serialised crystal structure, every candidate’s features["input_type"] records which one ("composition" or "structure") its data string actually holds. problem_type is likewise auto-set from task_name (REGRESSION for Matbench regression tasks, BINARY for Matbench classification tasks) and should not be set explicitly.

When fold_number is set (0-4), the predefined Matbench train/test split for that fold is used. Of the Matbench train pool, train_ratio sets aside a slice for train + validation combined; validation_frac then carves validation out of that slice (not out of the whole Matbench train pool, and not out of the total dataset) — the rest of the slice becomes train. Whatever remains of the Matbench train pool beyond that slice becomes candidate_pool (same denominator FLIP uses). test_ratio and split_type are ignored; the full Matbench test set for that fold is used directly as test. Matbench’s predefined folds use a fixed internal seed (18012019) that cannot be overridden.

When fold_number is None, all 5 folds are merged into a single dataset and split using the standard ratio-based train_ratio/validation_frac/test_ratio/ split_type — this loses Matbench’s benchmark integrity guarantees (results are no longer directly comparable to published Matbench leaderboard scores). The fold each candidate originally belonged to (0-4) is recorded in features["fold_id"] regardless of mode, for traceability.

task_name

Name of the Matbench task, e.g. "matbench_steels".

fold_number

0-4 to use a single predefined Matbench fold; None to merge all 5 folds and split by ratio instead.

Example:

# Fold mode
config = MatbenchConfig(
    name="matbench_mp_e_form_fold0",
    task_name="matbench_mp_e_form",
    fold_number=0,
    seed=42,
    train_ratio=0.1,
    validation_frac=0.1,
    test_ratio=0.2,  # ignored in fold mode
)

# Merged mode
config = MatbenchConfig(
    name="matbench_mp_e_form_merged",
    task_name="matbench_mp_e_form",
    fold_number=None,
    seed=42,
    train_ratio=0.1,
    validation_frac=0.1,
    test_ratio=0.2,
)
fold_number: int | None
modality: Modality
model_config: ClassVar[ConfigDict] = {}

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

problem_type: ProblemType
split_type: Literal['random', 'low_vs_high']
task_name: str
validate_config()[source]

Override base class validator.

Looks up task_name in Matbench’s own task metadata (raising a clear error for unknown tasks), validates fold_number, and auto-sets problem_type from the task’s Matbench problem type. The base class’s train_ratio + test_ratio <= 1 check is intentionally not carried over: in fold mode the two ratios apply to separate pools (Matbench train vs. Matbench test), so their sum is allowed to exceed 1 — the same reasoning FLIPConfig uses.

Return type:

Self

Returns:

The validated configuration instance.

Raises:
  • ImportError – If the matbench package is not installed.

  • ValueError – If task_name is not a recognised Matbench task, or if fold_number is not None and not in 0-4.