Сегодня MIT + Merck выпустили на ChemRxiv CheMeleon для реакций, предобученный на 2 млн реакциях.
https://chemrxiv.org/doi/full/10.26434/chemrxiv.15008692/v1
🔥Код полностью открытый: https://github.com/MSDLLCpapers/chemeleon-rxn
А вот веса будут опубликованы позже:
https://chemrxiv.org/doi/full/10.26434/chemrxiv.15008692/v1
We introduce CheMeleon-Rxn, a pre-trained reaction graph neural network that adapts the CheMeleon framework to condensed graphs of reaction (CGR) and demonstrates that descriptor-regression is an effective pre-training strategy for low-data reaction property prediction.
An O(10M)-parameter D-MPNN encoder was pre-trained on #2 million reactions by regressing onto dense reaction descriptors pooled from classical molecular descriptors, then fine-tuned on small reaction property prediction datasets. Across regression tasks spanning gasphase activation energies, reaction enthalpies, experimental yields, and rate coefficients, CheMeleon-Rxn is best or statistically tied-for-best on seven of eight tasks. It outperforms the same encoder trained from scratch, as well as classical-descriptor, fingerprint, and pre-trained-fingerprint baselines.
We further tested the pre-training strategy across various graph neural network architectures and found that its benefit holds for other edge-aware backbones such as GINE, and that it is robust to the choice of descriptor set, pooling operation, and other pre-training hyperparameters. We visualized and analyzed the CheMeleon-Rxn embeddings, which reveal chemically coherent structure in the learned reaction representation space
🔥Код полностью открытый: https://github.com/MSDLLCpapers/chemeleon-rxn
А вот веса будут опубликованы позже:
The model weights will be released upon publication.