AlphaFold, the revolutionary AI model for predicting protein structures, is facing a significant challenge due to a shortage of data. While AlphaFold 3 has proven to be a game-changer in drug discovery by modeling protein interactions with molecules, it struggles to predict how proteins interact with drugs. This is because the publicly available protein databases, such as the Protein Data Bank (PDB), predominantly feature proteins interacting with biological molecules like ATP, rather than drug compounds. The lack of drug-related data in these databases limits AlphaFold’s ability to model drug-protein interactions effectively.
In response to this limitation, a group of leading pharmaceutical companies has announced plans to create their own version of an AI model inspired by AlphaFold 3. The companies, which include AbbVie, Johnson & Johnson, and Sanofi, intend to use proprietary data that has been locked away in their internal vaults. These data sets, which contain protein structures bound to various drug candidates, are not typically shared with the public or other companies. The consortium aims to leverage these secret data to enhance the model’s ability to predict drug-protein interactions, though access to the new AI model will be restricted to consortium members only.
The initiative will build upon OpenFold 3, an open-source version of AlphaFold 3 developed by academic researchers. However, the pharmaceutical companies will not share their data with one another or with external researchers. Instead, they will use a secure platform developed by the Berlin-based start-up Apheris, which allows companies to train the model without exposing their proprietary data. This platform ensures that the secret structures are kept secure and cannot be reverse-engineered by others.
Despite the potential improvements this private data could offer, there is uncertainty about how much it will actually boost AlphaFold’s performance. While some experts believe the inclusion of more drug-related data could improve predictions of drug interactions, others are more skeptical. Even modest improvements in predicting how drugs bind to proteins could be a valuable advancement for drug discovery, particularly in reducing the trial and error involved in experimental processes.
One of the ongoing discussions in the field is the question of whether pharmaceutical companies should share more of their structural data with the broader scientific community. Currently, just 6% of the structures in the PDB come from drug companies, and there is significant reluctance within the industry to make this data publicly available. Some experts remain hopeful that the success of AlphaFold and similar models might encourage greater openness in the future, which could benefit drug discovery as a whole. However, others are more pessimistic, given the long history of secrecy in the pharmaceutical sector.

