One way to think about this is that when the system encounters a site where all chemically related amino acids function (like leucine/isoleucine/valine or serine/threonine), it will use one that matches the watermark if possible. Another way to look at the process was suggested by one of the people involved in developing the system: it searches the space occupied by functional proteins for the subset that happens to have a sufficient number of watermark amino acids.
This means that the watermark is randomly distributed along the entire length of the protein and detecting one is not a simple yes-or-no question. You need to scan the entire sequence, know the key, and measure how often the amino acids suggested by SynthIDBio actually appear in the final sequence. Google has also developed the necessary software for this.
It’s alive!
So the question is whether watermarked proteins are functional. The team used the system to design proteins that physically interact with key natural proteins previously targeted by AI designs. And the watermarked versions worked fine and bound the intended targets. This is not as rigorous a test as finding a catalyst, but it suggests that there is no reason to expect serious problems with more complicated design tasks.
So as long as a protein is long enough, the system can recognize a watermark. How could this be useful? Here too, biosecurity is important. When someone orders DNA sequences, the people making the DNA typically check the sequence for its ability to encode parts of viruses, toxic proteins and other similar threats. However, if they currently see a protein that resembles nothing we already know – which may be true of all AI-developed proteins – they cannot assess its threat.