Google DeepMind has found a way to stamp a secret signature into AI-designed proteins, and the proteins still do their jobs.
The company introduced SynthID Bio on Wednesday as a proof of concept for tracing biological designs created with AI. The technology embeds detectable patterns into protein sequences and predicted 3D structures. Laboratory tests found that watermarked protein binders retained their function, while the structural approach largely preserved prediction accuracy.
The research, published in Nature, tackles a growing biosecurity problem: AI can generate biological designs that may differ substantially from known threats, making their origins harder to determine using conventional screening methods.
How SynthID Bio works
SynthID Bio builds on Google’s existing SynthID watermarking technology, which identifies AI-generated digital content. Instead of modifying pixels or text, the biological version subtly influences the selection of amino acids, the building blocks of proteins.
A specialized verification system can detect the resulting pattern, allowing researchers to determine whether a specialized verification system can detect the resulting pattern, providing evidence that a sequence was generated through a SynthID Bio-enabled design process.
For 3D structures, DeepMind fine-tuned part of AlphaFold 3’s diffusion module, embedding the signature directly into the model’s weights. The company says this preserves prediction accuracy while keeping watermark detectability near-perfect, even after minor digital alterations.
Laboratory tests show promising results
DeepMind tested the technology using AlphaProteo and a modified version of ProteinMPNN to design protein binders, molecules engineered to attach to specific targets.
Experiments involving VEGF-A, the SARS-CoV-2 spike protein’s receptor-binding domain and PD-L1 showed that watermarked proteins matched their unmarked counterparts in hit rate, binding affinity and natural sequence diversity.
DeepMind has also integrated the technology into Evo 2 to watermark an AI-designed bacteriophage. Early laboratory tests confirmed that the modified virus remained functional, although further details are expected in a separate technical paper.
These results suggest that watermarking does not necessarily compromise protein performance, although the experiments cover a limited range of designs.
A new layer for biological security
The technology could help DNA synthesis companies distinguish AI-generated sequences from unfamiliar natural proteins when reviewing customer orders.
Currently, providers screen DNA sequences against databases of known biological threats. However, AI-generated proteins may bear little resemblance to previously identified hazards, making automated screening less reliable and forcing researchers to conduct time-consuming manual reviews.
SynthID Bio could provide an additional verification signal, helping providers identify designs originating from trusted models and concentrate scrutiny on sequences that lack a recognizable provenance. It could also help protect scientific databases such as the Protein Data Bank, UniProt and GenBank from mislabeled synthetic entries that could mislead future research.
The challenges DeepMind still faces
Despite the encouraging results, SynthID Bio is not a complete solution to biological security risks.
DeepMind acknowledges that making watermarks resistant to deliberate removal remains a major challenge. Short proteins may also lack enough amino acids to carry a reliably detectable signature. Researchers have raised additional concerns about whether the technique will work with more complex proteins, particularly enzymes, where even minor sequence changes can affect functionality.
Adoption presents another hurdle. Protein designers, synthesis providers and database managers will need compatible verification systems and agreed standards before watermarking can become a routine part of biological research.
DeepMind is exploring ways to strengthen the technology through provenance metadata and centralized repositories of AI-generated biological data.
More Google coverage
- New Google Search AI Mode is ‘Total Reimagining,’ Says CEO Sundar Pichai
- In Major Ruling, Judge Finds Google ‘Willfully Acquired and Maintained Monopoly Power’ Over Digital Ad Market
- Google’s Big Bet on Nuclear Energy: ‘The Race to Power AI-Driven Data Centers is Accelerating’
- Computer History Museum Releases Original AlexNet Code: Why It Matters
What this means for researchers and users
For scientists, SynthID Bio could eventually make it easier to establish the origin of experimental proteins, protect research integrity and avoid unnecessary delays when ordering synthetic biological materials.
Biotechnology companies could benefit from improved screening processes, while database operators may gain another way to identify and label AI-generated submissions.
The implications extend beyond laboratories. As AI becomes more capable of designing biological systems, the ability to trace those designs could become an important part of ensuring that scientific progress does not outpace safety measures.
The company is releasing its methods paper, code, laboratory data and model weights to the research community, inviting further testing and collaboration. For now, SynthID Bio is an early demonstration rather than an established industry standard.
As AI-generated biology becomes easier to create, knowing where a protein came from may become almost as important as knowing what it does.
Other news: OpenAI’s new $500-a-month ChatGPT Pro tier adds Astra Ultrafast for Work and Codex, promising up to eight times faster token generation for power users willing to pay $300 more than Pro 200.