Research protocol

OpenAI Image Provenance Test: C2PA, Metadata & Provider Verification

A reproducible protocol for inspecting OpenAI- and ChatGPT-generated images for ordinary metadata, embedded Content Credentials, and optional provider verification. SynthID is not independently tested here. Dataset pending.

Published
Last verified
Review
Reviewed against primary sources

Dataset pending

Study ID
awc-openai-image-provenance-001
Independent dataset
Pending
Research question
What ordinary metadata and embedded C2PA can AI Watermark Center independently inspect in owned OpenAI-generated images, and how does that compare with OpenAI’s official verifier when a check is recorded as provider-verified?

Official provider documentation

What the company currently publishes. Not an AWC measurement.

AWC observations

Unicode or embedded-metadata inspection of owned/permissioned samples. Scoped to that sample.

Provider-verified results

Manual use of an official verifier. Not independent SynthID detection by AWC.

Limitations

Small samples are exploratory. not-detected is not human authorship.

Dataset

Status
draft
Version
None — not published
Public samples
0
Downloads
No download. Dataset pending.

Raw provider outputs stay private unless a sample is explicitly flagged for publication. Test fixtures are not research samples.

Research questions

  • What metadata is embedded in current OpenAI-generated image files?
  • Are Content Credentials readable by AI Watermark Center?
  • What OpenAI issuer or model information is exposed when a manifest is present?
  • Does C2PA differ by ChatGPT versus API generation, when both samples exist?
  • What happens after simple, ordinary file conversions (re-export, format change, platform upload/download)?
  • What does OpenAI’s official verifier report for original samples, if a check is performed?
  • How does provider verification compare to AWC embedded-C2PA inspection?

Sample surfaces

The protocol supports ChatGPT image generation, OpenAI API image generation, and Codex image generation only if genuine samples can be generated or obtained with permission. For each file we will record surface, model if known from the UI or API, generation date, generated versus edited, file format, dimensions, file size, and original export path. Surface and model are never inferred from pixels or a filename.

Provider verification versus AWC inspection

If OpenAI Verify is used manually, the record stores verification surface, date, result, signal types, and model or issuer only when returned. That row is classified Provider-verified. It is not an independent AWC SynthID detection. AI Watermark Center does not possess an independent SynthID detector.

Content Provenance API

API use is not required to complete this protocol. No public upload to OpenAI is implemented. Automated paid calls must not run in tests. If a later phase integrates the API, privacy, credentials, rate limits, and terms will be reviewed first.

Planned results model

Future tables will allow: sample, surface, model, format, embedded C2PA, C2PA issuer, OpenAI Verify C2PA, OpenAI Verify SynthID, AWC C2PA inspection, and notes. Unknown and Not tested are explicit. No blank ambiguous cells. No numeric results while the sample list is empty.

Benign transformation tests

If later performed, tests may include original download, re-export through a normal image editor, format conversion, and social or platform upload/download. The purpose is metadata durability and provenance preservation research. This study will not publish a “best way to remove C2PA,” run SynthID degradation hunts, or score watermark resistance.

Primary sources

Product facts follow current OpenAI Help Center provenance documentation, the advancing-content-provenance announcement, OpenAI Verify, and Content Provenance API documentation listed on this page.

Sources

  1. Provenance signals (Content Credentials, SynthID) in OpenAI-generated content — OpenAI Help Center

    Accessed August 17, 2026.

    Primary source for current OpenAI image C2PA + SynthID on supported ChatGPT, Codex, and API outputs; audio SynthID; prompt-requested visible disclosure; and openai.com/verify. Visible marks are separate from embedded signals.

  2. Advancing content provenance for a safer, more transparent AI ecosystem — OpenAI

    Published July 31, 2026. Accessed August 17, 2026.

    31 July 2026 update: supported audio SynthID, audio on the public verifier, and verification API access. Historical context for layered C2PA + SynthID; current product scope follows the Help Center.

  3. Verify — OpenAI

    Accessed August 17, 2026.

    Public verification surface referenced by OpenAI Help Center for supported image and audio provenance signals (C2PA and/or SynthID). External provider tool — not an AI Watermark Center feature.

  4. Content provenance — OpenAI API documentation

    Accessed August 17, 2026.

    Content Provenance API (also referenced as OPENAI-API-PROV-001 in Phase 7 notes). Images: C2PA + SynthID. Audio: SynthID. Documents not_detected semantics. Public/documented; not integrated by AI Watermark Center.