Search
Mobile menu Mobile menu
Data science & AI , Regulatory , AI in Life sciences Oct 04, 2026

AI-Generated Biology Has a Provenance Problem: What Biotech Engineering Leaders Need to Know About Synthetic Data Integrity

VECTOR Labs Team
VECTOR Labs Team
AI-Generated Biology Has a Provenance Problem: What Biotech Engineering Leaders Need to Know About Synthetic Data Integrity
Last updated on: Oct 04, 2026

The protein databases that underpin modern drug discovery and synthetic biology research were built on the assumption that their contents reflect observed reality. That assumption is increasingly strained. As generative AI tools produce millions of novel protein structures, sequences, and molecular designs at a pace no experimental validation pipeline can match, the question of whether a given biological record is computationally predicted, experimentally confirmed, or entirely synthetic has become a material engineering concern, not a philosophical one.

The Database Contamination Problem Is Already Here

Public repositories like the Protein Data Bank and UniProt have long served as ground truth for structural biology. When AI-generated structures enter these databases without clear provenance markers, downstream models trained on that data inherit the uncertainty. A model trained on a mix of experimentally resolved and AI-predicted structures will behave differently from one trained on verified data alone, and the difference may not surface until a compound reaches a wet-lab validation stage or, in worse cases, a regulatory submission.

The mechanism is straightforward: generative protein models can produce plausible-looking outputs that satisfy structural energy constraints without corresponding to anything that exists or could be synthesised safely. If those outputs are deposited into shared databases without adequate labelling, they become training data for the next generation of models. The contamination compounds silently.

This is not a theoretical risk. AlphaFold's release into the European Bioinformatics Institute's database added over 200 million predicted structures. The scientific value is undeniable. The provenance challenge is equally real, because predicted structures carry a fundamentally different epistemic status than crystallography or cryo-EM resolved entries, and that distinction is not always preserved as data moves between systems.

What Watermarking Approaches Actually Offer

Google DeepMind's SynthID Bio represents one of the more technically grounded attempts to address this problem. The approach embeds a statistical signal into AI-generated biological sequences, specifically protein sequences in the initial implementation, that persists through typical downstream processing and can be detected without access to the original model. The watermark is designed to be invisible to standard bioinformatic analysis while remaining recoverable by a detection algorithm.

The Signal Persistence Challenge

The core technical difficulty with biological watermarking is that biological sequences are not static artefacts. They get truncated, mutated, aligned, and recombined as researchers work with them. A watermarking scheme that degrades under standard sequence manipulation provides weaker guarantees than one that survives those operations. SynthID Bio's design targets this directly by embedding the signal in a way that is statistically recoverable even from partial sequences, though the robustness of that guarantee across all realistic manipulation scenarios remains an active area of evaluation.

What Watermarking Cannot Do

Watermarking addresses the detection problem for sequences generated by a specific model using a specific scheme. It does not address sequences generated by models that do not implement watermarking, which is currently the majority of deployed protein design tools. It also does not retroactively label the AI-generated content already present in public databases. Engineering teams should treat watermarking as one layer of a provenance architecture, not as a complete solution.

DNA Synthesis Screening Has a Structural Gap

The biosecurity dimension of this problem sits at the synthesis layer. DNA synthesis providers are expected to screen orders against databases of sequences of concern, including known pathogen genomes and toxin-coding sequences. The screening systems were designed for a world where sequences of concern are known and catalogued. Generative biology introduces a different threat model: novel sequences that do not match any existing entry but that encode functional properties that would be flagged if the screening system understood function rather than identity.

Current screening approaches are predominantly sequence-similarity based. A sufficiently divergent AI-designed sequence encoding a dangerous protein function may not trigger a match. The gap is not a failure of the screening vendors specifically; it reflects the fact that the threat model has changed faster than the regulatory and technical infrastructure designed to address it.

Engineering teams building synthesis workflows on top of AI-designed sequences should not assume that standard synthesis provider screening provides equivalent protection to what it offered five years ago. The risk surface has expanded, and the screening tools have not yet caught up.

What Data Lineage Actually Requires in a Biological AI Stack

Provenance in a biological AI context is not simply a matter of recording where data came from. It requires capturing what transformations were applied, which model version generated a given output, what confidence or uncertainty estimates accompanied that output, and whether the output has been experimentally validated at any stage. This is a data engineering problem as much as a biology problem.

Minimum Viable Provenance for Production Workflows

For engineering teams integrating biological AI outputs into research or commercial pipelines, a minimum viable provenance architecture should track the following at the record level:

  • Generation source: model name, version, and whether the output is predicted, designed, or experimentally derived
  • Validation status: whether any experimental confirmation exists and at what resolution or assay type
  • Transformation history: any downstream processing steps that may have altered the original output
  • Screening status: whether the sequence has been passed through biosafety screening and which screening version was used

Without this information at the record level, data governance becomes a post-hoc exercise rather than a structural property of the pipeline.

The Regulatory Exposure

Regulatory frameworks for AI-generated biological content are still forming, but the direction of travel is clear. Agencies are moving toward requiring explicit documentation of training data provenance for AI systems used in drug development. Engineering teams that build pipelines now without provenance infrastructure will face a retrofit problem later, and retrofitting data lineage into a production biological AI system is significantly more expensive than building it in from the start.

Where Vector Labs Fits

We build and certify AI systems for regulated life sciences environments where data provenance and validation standards are non-negotiable. In our cardiovascular certification work, we delivered a clinical-grade AI model on wearable ECG data that achieved Class 2A medical device certification, with validation structured from the outset to meet medical device software standards including prospective held-out test sets and full regulatory documentation. If you are building biological AI workflows that will need to meet regulatory scrutiny, contact us at vector-labs.ai/contacts.

FAQs

Does SynthID Bio protect us if we are using multiple protein design tools from different vendors?

Not comprehensively. SynthID Bio watermarks sequences generated by DeepMind's own models using its specific scheme. Sequences produced by other tools, including open-source protein design models, will not carry that watermark. A robust provenance architecture needs to operate at the pipeline level, recording generation source and model version for every record regardless of which tool produced it, rather than relying on any single vendor's watermarking implementation.

How do we assess whether AI-generated structures in public databases are affecting our model training?

Start by auditing the provenance metadata of your training corpus at the record level. Many public database entries now include method fields that distinguish experimental techniques from computational predictions, but the granularity varies and is not always machine-readable in a consistent format. For high-stakes applications, consider maintaining a curated internal dataset of experimentally validated structures with explicit provenance records, and treating public database entries as supplementary data with appropriate uncertainty weighting rather than as ground truth.

What should we require from DNA synthesis providers given the screening gaps described?

Request documentation of which screening database versions and algorithms are in use, and ask specifically whether their screening approach incorporates functional analysis or is purely sequence-similarity based. For novel AI-designed sequences, consider running an independent screening check using a separate provider or an internal biosafety review before submission, particularly for sequences with no experimental precedent. This adds process overhead, but it closes a gap that standard provider screening does not currently address.

Is there a regulatory requirement today to document whether biological AI outputs are AI-generated?

There is no single unified requirement today, but the regulatory direction is toward explicit AI documentation in drug development submissions, and several jurisdictions are developing guidance on AI-generated data in regulated research contexts. The more immediate risk is internal: if your validation studies or regulatory submissions rely on data whose provenance is unclear, that ambiguity becomes a finding during audit. Building provenance infrastructure now is a risk management decision, not just a compliance anticipation exercise.

How should we think about the cost-benefit of building provenance infrastructure versus moving faster with existing tools?

The cost calculation changes depending on where in the development pipeline a provenance failure surfaces. Early in research, a mislabelled structure costs time. In a regulatory submission or a manufacturing process, it can cost a programme. The infrastructure investment required to capture generation source, validation status, and transformation history at the record level is not large relative to the cost of a late-stage data quality finding. Teams that defer this work typically find that the retrofit cost, in engineering time and in re-validation of existing datasets, substantially exceeds what proactive implementation would have required.

A team that understands you
With 20+ years of experience in the world's leading consultancy companies, implementing AI and ML projects in industry-specific contexts, we are ready to hear your challenges.
Subscribe to our newsletter for insights and updates on AI and industry trends.
By clicking "Sign me up", you agree to our Privacy Policy.
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration