MainAI is accelerating biological discovery and design, with wide-ranging applications in therapeutics, diagnostics and biomedical research5. Specialist models for protein design and structural predictions are becoming an integral part of biological workflows. Ensuring their safe and responsible use involves the ability to determine and track the original model of AI-generated sequences and structures. This holds promise to address potential challenges that stem from the widespread adoption of ever more powerful AI models in biology, including biosecurity and information veracity. Here, information on provenance may become a useful signal for DNA synthesis providers2,4,6,7,8 and institutions managing biological databases such as the Protein Data Bank (PDB)9, UniProt10 or GenBank11.Several means of establishing provenance for biological objects have been proposed, including through centralized databases or by associating metadata about the design process with the biological sequence2,12. Databases require central coordination, are prone to false positives at scale and may cause privacy challenges. Cryptographically signed metadata, by contrast, can preserve privacy, but is easy to remove. An alternative approach, watermarking, embeds signatures into the AI-generated objects themselves. This approach has successfully been applied to AI-generated text and multimedia and deployed in various products, including Google’s generative AI13,14. Watermarks promise high detectability and quality preservation by being imperceptible. They offer a privacy-preserving alternative where information on provenance is more challenging to remove and no central coordination is required. Moreover, watermarking can establish provenance for open model predictions. However, it is unclear whether existing watermarking methods transfer to biological modalities where preserving quality relates to the protein’s intended use: for protein design, this requires maintaining experimentally verifiable biological function; for structural prediction, it involves maintaining the accuracy of predicted structural features, ensuring a researcher’s interpretation and subsequent biological function hypotheses remain unchanged. Concurrent work for watermarking protein sequences is limited to in silico experiments with wild-type monomer structures from PDB3. Current methods result in a noticeable decrease in accuracy15 when watermarking protein structures. Combined, these show that function-preserving watermarking is not yet available.We introduce SynthIDBio, a family of methods for embedding highly detectable yet function-preserving watermarks directly into biological sequences and structures. SynthIDBio-sequence is a method for designing watermarked protein binders. Following Fig. 1a, SynthIDBio-sequence integrates watermarking into ProteinMPNN16, which is a commonly used sequence design model applied on top of a backbone structure before structural metrics are used to filter designs to predicted binders17,18. We show that detectability of the watermark can be ensured with minimal in silico performance reductions and explore the watermark’s robustness to intentional removal. Through in vitro validation, we obtained biologically watermarked, functional protein binders: low nanomolar binders for the SARS-CoV-2 receptor binding domain (SC2RBD) and subnanomolar binders for vascular endothelial growth factor A (VEGF-A) and programmed death ligand 1 (PD-L1). Our experimental design leveraged known backbones from AlphaProteo17, which were resequenced using ProteinMPNN with or without SynthIDBio-sequence. Across targets and backbones, we find that watermarking does not affect the binding affinity distribution or binding hit rates.Fig. 1: Graphical overviews of SynthIDBio-sequence and SynthIDBio-structure.Full size imagea, End-to-end protein design with SynthIDBio-sequence. During sampling, based on a generated backbone, we use ProteinMPNN with SynthID-text’s tournament sampling to sample watermarked sequences that are then filtered using metrics derived from AF3 and a further watermark (wm) detectability threshold that is calibrated to a target FPR. At test time, we use the secret watermarking key used during sampling for watermark detection. b, We fine-tune the diffusion and confidence modules of AF3 (blue) to obtain SynthIDBio-structure, which includes an imperceptible watermark within each structure prediction. We also train a watermark detector (blue) by adding an additional watermarking loss (orange) to AF3’s diffusion losses. The sampling process remains unchanged using the fine-tuned diffusion module. At test time, the watermark detector can be used in a standalone fashion to identify watermarked structures. The detector remains secret and is not shared with users of the fine-tuned AF3 model.We also present SynthIDBio-structure—a model fine-tuned from AlphaFold 3 (AF3)19 that predicts watermarked structures while preserving structural accuracy. Following Fig. 1b, we fine-tune AF3’s diffusion module, co-training it with a watermark detector that is guided by a watermarking objective integrated directly into AF3’s diffusion loss. SynthIDBio-structure offers near-perfect detectability with a negligible effect on global accuracy and a minimal effect on bond geometry, as validated on AF3’s evaluation set. The watermarks are robust to basic manipulations including rigid transformations and noise. Qualitative validation shows that SynthIDBio-structure successfully hides watermarking information in atom–atom distances, indistinguishable from the natural variance of AF3’s own predictions relative to the ground truth.SynthIDBio presents a technical proof-of-concept that function- and quality-preserving biological watermarking is possible. We also discuss potential use-cases, including for biosecurity and scientific integrity, where we expect the ability to reliably track provenance to be important in the future; however, operationalizing SynthIDBio for these applications will require further innovation in addition to industry-wide coordination and standardization.Tournament sampling for ProteinMPNNRecent approaches for protein binder design rely on target-conditional structure generation followed by sequence design using ProteinMPNN16,17,18,20 (Fig. 1a). ProteinMPNN, with left-to-right decoding, generates discrete sequences of amino acids autoregressively by sampling tokens (residues) xt at each step t, from a distribution p(⋅∣x