Matbench Dataset¶
The Matbench benchmark dataset implementation. Matbench provides 13 materials-property prediction tasks with predefined 5-fold cross-validation splits, covering both composition-based (chemical formula) and structure-based (crystal structure) inputs.
Both input kinds are stored under the MATERIALS modality.
Composition values are plain chemical-formula strings (e.g. "Fe0.62C0.01Mn0.37"),
used as-is. Structure values are pymatgen Structure
objects, serialised to an equivalent JSON string via their MSONable
.to_json() interface before being stored in Candidate.data.
Supported tasks:
Task |
Input type |
Problem type |
Samples |
Target |
|---|---|---|---|---|
|
composition |
regression |
312 |
yield strength (MPa) |
|
composition |
regression |
4,604 |
experimental gap (eV) |
|
composition |
classification |
4,921 |
is_metal |
|
composition |
classification |
5,680 |
glass-forming ability |
|
structure |
regression |
4,764 |
refractive index |
|
structure |
regression |
636 |
exfoliation energy |
|
structure |
regression |
10,987 |
log10(shear modulus) |
|
structure |
regression |
10,987 |
log10(bulk modulus) |
|
structure |
regression |
132,752 |
formation energy |
|
structure |
regression |
106,113 |
band gap (eV) |
|
structure |
classification |
106,113 |
is_metal |
|
structure |
regression |
18,928 |
formation energy |
|
structure |
regression |
1,265 |
last phonon DOS peak |
problem_type is derived automatically from the task — ProblemType.REGRESSION for
regression tasks, ProblemType.BINARY for the three classification tasks (all are
two-class) — and does not need to be set in MatbenchConfig.
Fold mode vs. merged mode:
MatbenchConfig.fold_number selects between two ways of using Matbench’s predefined
5-fold cross-validation:
Fold mode (
fold_numberset to0-4): uses Matbench’s predefined train/test split for that fold directly.train_ratiocontrols what fraction of the Matbench train rows form the initial labelled training set (the remainder becomescandidate_pool, capped atmax_candidate_poolif set);validation_fraccarves a validation set out of that.test_ratioandsplit_typeare ignored — the full Matbench test set for that fold is used astest, since benchmark- comparable results require Matbench’s exact predefined test rows.Merged mode (
fold_number=None): all 5 folds are combined into one dataset and split using the standard ratio-basedtrain_ratio/validation_frac/test_ratio/split_type. This loses Matbench’s benchmark integrity guarantees — results are no longer directly comparable to published Matbench leaderboard scores.
In both modes, every candidate’s features["fold_id"] records which of Matbench’s 5
folds (0-4) it originally belonged to, for traceability.
Dependencies:
Requires the optional matbench extra (pip install "alf-tools[matbench]" or
alf_tools[materials]), which installs matbench and pymatgen. Data is
downloaded and cached automatically by the matbench
package on first use.
- class alf_tools.datasets.matbench.Matbench(config)[source]¶
Bases:
BaseDatasetMatbench benchmark dataset class.
Matbench provides 13 materials-property prediction tasks, each with predefined 5-fold cross-validation splits, covering both composition-based and structure-based inputs. Composition and structure inputs (pymatgen
CompositionandStructureobjects respectively) are both MSONable, so both are serialised identically via.to_json()into a JSON string stored inCandidate.data; noalf_corechanges are needed.pymatgen/matbenchare optional: importing this module never requires them, and constructing aMatbenchConfigraises a clearImportErrorif they’re missing (seeMatbenchConfig.validate_config).- config: MatbenchConfig¶
- load_dataset()[source]¶
Load Matbench data via the Matbench API.
In fold mode, only the configured fold’s predefined train/test rows are loaded (each candidate tagged with a
matbench_splitfeature of “train” or “test” for use by_split_dataset). In merged mode, all rows across all 5 folds are loaded. In both modes, every candidate’sfeatures["fold_id"]records which Matbench fold it belongs to, andfeatures["input_type"]records whetherdatais a composition formula string or a serialised structure (constant across a given task’s candidates).- Return type:
- Returns:
LabelledCandidates with JSON-string composition/structure data and float labels.
- class alf_tools.datasets.matbench.MatbenchConfig(**data)[source]¶
Bases:
BaseDatasetConfigConfiguration for Matbench benchmark datasets.
Both composition and structure task inputs are stored under
Modality.MATERIALS— Matbench has no data that needs any other modality, somodalityis fixed and should not be overridden. SinceMATERIALScovers both a composition formula string and a JSON-serialised crystal structure, every candidate’sfeatures["input_type"]records which one ("composition"or"structure") itsdatastring actually holds.problem_typeis likewise auto-set fromtask_name(REGRESSIONfor Matbench regression tasks,BINARYfor Matbench classification tasks) and should not be set explicitly.When
fold_numberis set (0-4), the predefined Matbench train/test split for that fold is used. Of the Matbench train pool,train_ratiosets aside a slice fortrain+validationcombined;validation_fracthen carvesvalidationout of that slice (not out of the whole Matbench train pool, and not out of the total dataset) — the rest of the slice becomestrain. Whatever remains of the Matbench train pool beyond that slice becomescandidate_pool(same denominator FLIP uses).test_ratioandsplit_typeare ignored; the full Matbench test set for that fold is used directly astest. Matbench’s predefined folds use a fixed internal seed (18012019) that cannot be overridden.When
fold_numberisNone, all 5 folds are merged into a single dataset and split using the standard ratio-basedtrain_ratio/validation_frac/test_ratio/split_type— this loses Matbench’s benchmark integrity guarantees (results are no longer directly comparable to published Matbench leaderboard scores). The fold each candidate originally belonged to (0-4) is recorded infeatures["fold_id"]regardless of mode, for traceability.- task_name¶
Name of the Matbench task, e.g.
"matbench_steels".
- fold_number¶
0-4to use a single predefined Matbench fold;Noneto merge all 5 folds and split by ratio instead.
Example:
# Fold mode config = MatbenchConfig( name="matbench_mp_e_form_fold0", task_name="matbench_mp_e_form", fold_number=0, seed=42, train_ratio=0.1, validation_frac=0.1, test_ratio=0.2, # ignored in fold mode ) # Merged mode config = MatbenchConfig( name="matbench_mp_e_form_merged", task_name="matbench_mp_e_form", fold_number=None, seed=42, train_ratio=0.1, validation_frac=0.1, test_ratio=0.2, )
- fold_number: int | None¶
- model_config: ClassVar[ConfigDict] = {}¶
Configuration for the model, should be a dictionary conforming to [
ConfigDict][pydantic.config.ConfigDict].
- problem_type: ProblemType¶
- split_type: Literal['random', 'low_vs_high']¶
- task_name: str¶
- validate_config()[source]¶
Override base class validator.
Looks up
task_namein Matbench’s own task metadata (raising a clear error for unknown tasks), validatesfold_number, and auto-setsproblem_typefrom the task’s Matbench problem type. The base class’strain_ratio + test_ratio <= 1check is intentionally not carried over: in fold mode the two ratios apply to separate pools (Matbench train vs. Matbench test), so their sum is allowed to exceed 1 — the same reasoning FLIPConfig uses.- Return type:
Self- Returns:
The validated configuration instance.
- Raises:
ImportError – If the
matbenchpackage is not installed.ValueError – If
task_nameis not a recognised Matbench task, or iffold_numberis notNoneand not in0-4.