DeepMind's AlphaFold Database Expands to 200 Million Protein Structures, Covering Nearly All Known Life

In July 2022, Google DeepMind and EMBL-EBI released predicted 3D structures for over 200 million proteins — a 200-fold expansion that covers nearly every organism with a sequenced genome. The freely available database has become one of the most widely used scientific resources in history, transforming drug discovery, crop science, and basic biology.

FE
FIRAT Editorial BoardInstitutional Research Desk
Jul 28, 2022
7 min read
Share:
STEM
FIRAT
FIRATSTEM
Research & Translation

DeepMind's AlphaFold Database Expands to 200 Million Protein Structures, Covering Nearly All Known Life

Hinxton, United Kingdom · 28 July 2022 — Google DeepMind, in partnership with the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), announced a massive expansion of the AlphaFold Protein Structure Database, releasing predicted three-dimensional structures for over 200 million proteins. The expansion made the database the most complete and accessible protein structure resource in the world, covering nearly every organism whose genome has been sequenced.

The release represented a 200-fold increase from the database's previous size of approximately one million structures, which had been available since its launch in July 2021. The expanded database now spans the entire UniProt reference proteome — encompassing proteins from plants, bacteria, animals, archaea, fungi, and virtually every other branch of the tree of life.

The Protein Folding Problem

Proteins are the molecular machines of life. Each protein is assembled from a chain of amino acids — its sequence — which folds into a specific three-dimensional shape that determines its function. For over 50 years, predicting a protein's 3D structure from its amino acid sequence was one of biology's grand challenges, known as the "protein folding problem."

Experimental determination of protein structures — primarily through X-ray crystallography, nuclear magnetic resonance (NMR) spectroscopy, and, more recently, cryo-electron microscopy — is labour-intensive, expensive, and time-consuming. A single structure can take months or years to resolve, and many proteins resist experimental determination entirely. This left a vast gap: while genome sequencing projects had catalogued hundreds of millions of protein sequences, only a tiny fraction had corresponding structural data.

AlphaFold, the AI system developed by DeepMind, addressed this gap using deep learning. Trained on known protein structures and sequences, AlphaFold predicts a protein's 3D structure from its 1D amino acid sequence with accuracy that frequently approaches experimental methods. The system's performance at the Critical Assessment of Protein Structure Prediction (CASP14) competition in 2020 — where it achieved a median Global Distance Test Total Score (GDT_TS) of 92.4 across all targets — demonstrated that the protein folding problem had, for practical purposes, been solved for a large class of proteins.

How the Database Works

The AlphaFold Protein Structure Database is hosted by EMBL-EBI at the European Bioinformatics Institute in Hinxton, near Cambridge, UK. It is freely accessible to researchers worldwide, with no registration or login required. Users can search for proteins by UniProt accession number, gene name, organism, or sequence, and download predicted structures in standard formats (PDB and mmCIF) along with confidence metrics.

Each prediction includes a per-residue confidence score called pLDDT (predicted Local Distance Difference Test), which ranges from 0 to 100. Regions with pLDDT scores above 90 are considered high confidence, while regions below 50 are treated as low confidence and may correspond to disordered or flexible segments of the protein. Additionally, the database provides Predicted Aligned Error (PAE) scores, which estimate the expected position error at every residue pair, helping researchers assess the reliability of domain-level arrangements.

Impact Across Science

The release of 200 million predicted structures has had immediate and far-reaching effects across multiple scientific disciplines:

Drug Discovery and Neglected Diseases

Pharmaceutical research depends on understanding protein structures to design drugs that bind to specific molecular targets. Before AlphaFold, structural data was available for only a fraction of potential drug targets. The expanded database has been particularly valuable for neglected tropical diseases, where commercial incentives had historically limited structural biology research. Researchers studying malaria, leishmaniasis, and Chagas disease have used AlphaFold predictions to identify new drug targets and understand parasite biology.

Agriculture and Food Security

The database includes predicted structures for proteins from major crop species — rice, wheat, maize, soybean, and cassava, among others. Plant scientists have used these predictions to study enzymes involved in photosynthesis, disease resistance, and stress responses, with implications for crop improvement and food security.

Evolutionary Biology

With structures now available for proteins across the tree of life, evolutionary biologists can trace structural changes over billions of years of evolution, identifying conserved folds and domains that reveal deep evolutionary relationships invisible at the sequence level alone.

Enzyme Engineering

Industrial biotechnology relies on engineering enzymes for applications ranging from biofuel production to plastic degradation. AlphaFold predictions help researchers understand enzyme active sites and design modifications to improve catalytic efficiency or alter substrate specificity.

Limitations and Caveats

Despite its power, AlphaFold has important limitations that researchers must consider:

  • Static structures: AlphaFold predicts a single, static conformation. Many proteins are dynamic, adopting multiple conformations that are essential to their function. The predictions do not capture this conformational flexibility.
  • Protein complexes: The original AlphaFold (versions 1 and 2) predicts individual protein chains, not multi-protein complexes. While some complexes can be approximated by combining individual predictions, the interactions between subunits are not modelled. (This limitation was partially addressed by AlphaFold-Multimer and later by AlphaFold 3, released in 2024.)
  • Ligands and cofactors: AlphaFold predictions do not include bound small molecules, ions, or cofactors that are often essential to protein function.
  • Confidence varies: While many predictions are highly accurate, some — particularly for disordered regions or proteins with unusual compositions — have low confidence scores. Researchers are advised to treat low-confidence regions with appropriate caution.

Open Science and Accessibility

DeepMind made the decision to release the AlphaFold code and database freely, without commercial restrictions. The database is accessible through a web interface and an API, and bulk downloads are available for researchers who need the entire dataset. DeepMind also published the AlphaFold source code and methodology in the journal Nature, enabling other groups to build upon and extend the work.

This commitment to open access has been widely praised. The resource has levelled the playing field for researchers in low- and middle-income countries, who now have access to the same structural data as their counterparts in well-funded institutions — requiring only an internet connection.

From AlphaFold to AlphaFold 3

The July 2022 release represented the full expansion of AlphaFold 2 predictions across known protein space. In May 2024, DeepMind and Isomorphic Labs introduced AlphaFold 3, which expanded the system's capabilities beyond individual proteins to predict the structures and interactions of all life's molecules — including DNA, RNA, ligands, and ions — within a single model. This further broadened the tool's applicability to drug design and molecular biology.

Sources

  • Google DeepMind, "AlphaFold reveals the structure of the protein universe," deepmind.google, 28 July 2022
  • EMBL-EBI, "AlphaFold Protein Structure Database," ebi.ac.uk, accessed 2022
  • Jumper, J. et al., "Highly accurate protein structure prediction with AlphaFold," Nature, 596, 583–589 (2021)
  • Varadi, M. et al., "AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models," Nucleic Acids Research, 50, D439–D444 (2022)
  • NIH National Library of Medicine, AlphaFold documentation and resources, nih.gov
Filed Under:#AI#Protein Structures#Bioinformatics#DeepMind#Structural Biology

Share this research insight

Help circulate peer-reviewed evidence and institutional briefings.

Share:
FE
Author SpotlightDivision: ReMIT

FIRAT Editorial Board

Institutional Research Desk · Foresight Institute of Research and Translation

The collective editorial and research translation board of FIRAT, synthesising peer-reviewed evidence, policy briefs, and division milestones across our seven foundational research pillars.

Focus:Institutional PolicyResearch StrategyAfrican DevelopmentInnovation
More Research

Related Articles in Science Technology Engineering and Mathematics (STEM)

View all in STEM