Skip to content
English
FikirPilot content

Pharmaceutical companies’ confidential data is powering AI protein models

Updated: 15 Eyl 2026 · 1 min read · 199 words

Published: · Story reached us: · Processing time: 10 h 46 min

Pharmaceutical companies’ confidential data is powering AI protein models
A 3D protein model in a laboratory

A consortium formed by pharmaceutical companies trained OpenFold3 using more than 20,000 protein structures from company archives. The new model performed better than models trained solely on publicly available data or on data from a single company. The study, which had not undergone peer review, was announced in a blog post; the model was not made publicly available.

Why it matters

The results suggest that in protein structure modeling, private data that can be shared among companies may affect performance comparisons just as much as the breadth of the data pool. This raises the question of how inaccessible corporate archives alter the competitive conditions for researchers working with publicly available data. For pharmaceutical companies, it is becoming important to determine whether jointly used data produce different results from studies based on the resources of a single organization. However, because the study has not yet been reviewed by peers and the model has not been made available for use, it is currently not possible to independently verify the findings, evaluate the method in detail, or test it with other teams. Both circumstances leave open whether the performance gap stems from combining the data or from other characteristics of the model.

Source: Nature Machine Learning