Methodology

We only claim what we can verify.

AI Watermark Center separates what a provider has published, what this project has reproduced, what has been seen only in limited testing, and what remains a claim without a verification path.

Evidence hierarchy

Officially documented Officially documented

Information supported directly by provider documentation, a published standard, or another primary source we can link. Official documentation is not the same as independent reproduction. We cite it as what the source says, with an access date.

Independently verified Independently verified

Information reproduced through AI Watermark Center testing with a method a reader can follow: inputs, software, dates, and what was actually inspected. “Verified” means we ran the check, not that the phenomenon is universal.

Observed Observed

Behaviour seen in limited testing but not yet shown broadly enough to call verified. Observed findings stay labelled as such and should not be rewritten as product capabilities.

Unverified Unverified

Claims circulating in search results, social posts, or third-party tools that we cannot currently validate. Unverified items may be described as claims. They are not reported as fact.

Provider tool verification

Results obtained from Google’s Gemini verification, Adobe Inspect, OpenAI’s public verifier, or any other provider tool are provider-verified. They are not independently generated by AI Watermark Center. Public detector availability and AWC integration remain separate fields.

Surface-specific claims

Product, model, surface, region, and subscription tier can change watermark and credential behaviour. A Gemini Apps setting is not automatically true of Google AI Studio, the Gemini API, Search, or Ads. Records should keep that scope explicit. Unknown stays unknown.

Invisible versus visible provenance

A visible logo, an imperceptible SynthID watermark, and C2PA Content Credentials are different mechanisms. Missing one is not evidence that the others are missing. Inspectors that read metadata do not detect SynthID.

Claim status versus evidence

Evidence authority and claim status are independent dimensions. “Official” means an official source supports the evidence, not that the proposed mechanism is confirmed. Official documentation can confirm a mechanism or rule one out. Product support — whether AI Watermark Center can inspect something — does not change the claim status. Unknown is stored as unknown; it is not coerced to false or “unsupported.”

Provider-side verification versus independent inspection

An OpenAI Verify or Content Provenance API result is provider-verified evidence. AI Watermark Center’s C2PA parser is independent embedded-metadata inspection for supported JPEG, PNG, and WebP. Those are different methods. Public detector availability (including a documented verification API) is not AWC integration. not-integrated does not mean the provider tool is unavailable.

OpenAI documents that a not_detected API or verifier outcome does not prove a file was not generated by OpenAI, is human-created, or has no AI provenance. Do not translate not_detected into clean, human, or not AI. Missing C2PA is likewise not proof a file is not OpenAI-generated.

AI Watermark Center does not independently detect SynthID. C2PA is not SynthID. A visible disclosure is not SynthID.

Record verification dates and the changelog

Each database row has its own last-verified date. The catalogue also shows a snapshot version and a last-database-review date. Factual changes to published rows should add a changelog entry on /database/changelog. Deploying the website does not move those dates.

Source selection

Provider-specific facts require first-party documentation, published standards, or peer-reviewed papers for underlying technology. Competitor blogs are not used as evidence for another company’s product behaviour.

Non-claims

AI Watermark Center does not claim that absence of a detectable signal proves content was not generated by AI.

AI Watermark Center does not claim successful removal of an invisible or statistical watermark unless that result can be independently verified.

Dates mean different things

Published is when a page first went public. Last updated is when the prose or structure changed in a way readers should know about. Last verified changes only when factual claims or experiments are rechecked. Deploy dates do not advance verification dates.

Provider marking versus inspectable Unicode

A provider may document a keyed statistical watermark that a third party cannot reproduce without that key. Hidden Unicode, EXIF, and embedded C2PA are inspectable without the provider’s secret. Mixing those categories is a claim error. Claude’s official text watermark is documented as statistical, not as hidden characters. Absence of a supported signal is not proof of absence of AI generation or provenance.

Reproducibility

Independent tests should record provider, model if known, surface, date, prompt class, language, and how the sample was exported. Retesting changes last-verified dates. Deploying the website does not.

Sample collection

Planned sample groups may include prose, creative writing, factual answers, translation, proofreading, code, and later Arabic, French, and Spanish. Counts are published only after collection. Long third-party outputs are not republished.

Corrections and retesting

If Anthropic or another provider revises documentation, factual pages are rechecked against primary sources before last-verified dates change. Protocols can be public while datasets remain pending.

Image inspection coverage

The image inspector reports embedded EXIF, GPS, XMP, and IPTC when the parser can read them, and embedded C2PA manifests when the CAI web SDK can read them. C2PA validation codes from that SDK are shown as validation information, not as authenticity or authorship scores.

Absence of a supported signal is not proof of absence of AI generation or provenance. External Content Credentials, soft bindings, SynthID, and other invisible pixel watermarks are outside Phase 3 coverage.

Sample provenance

Published datasets include only project-generated samples, user-supplied samples with permission, or appropriately licensed official fixtures. Unknown scraped social images and unclear-rights corpora are rejected. Missing provenance makes a sample ineligible for publication.

Dataset lifecycle

Datasets move through draft, collecting, published, and superseded. Draft and collecting datasets have no download CTA and no Dataset schema. Versions are explicit (for example 1.0, 1.1). Silent mutation of published counts is not allowed.

Provider verification versus AWC inspection

AWC observations come from the Phase 2 Unicode engine or Phase 3 embedded image inspection. A result from OpenAI Verify, Gemini’s verifier, or another official tool is provider-verified, not independent SynthID detection. Official provider documentation stays on the evidence axis “official.” AWC measurements, when they exist, are “observed” unless a separate independent protocol warrants “verified.” Small samples do not overwrite official ClaimStatus.

Aggregation and small samples

Totals are computed from the dataset, not typed by hand. Exploratory collections are reported as “in this sample” or “among the collected files.” They do not justify “Nano Banana always,” “Claude never,” or “all OpenAI images.”

Raw sample privacy

Raw text and images are private research material by default. Public records are metadata and analysis summaries. Private prompts, local paths, account details, and API keys are not published.

Tools inherit the same standard

A detector, cleaner, or C2PA reader must show the evidence it used. The text scanner reports inspectable Unicode only. The image inspector reports metadata and embedded credentials it can decode. Placeholder pages do not advertise SoftwareApplication schema, fake scan results, or guaranteed removal.