Hidden & Invisible Unicode Characters in Text
What hidden and invisible Unicode characters are, why they appear in copied text, and what listing them can — and cannot — prove about AI. Use the text scanner when you need to check a string.
- Published
- Last updated
Hidden Unicode characters are code points that often have no visible ink — zero-width spaces, joiners, bidirectional controls, tag characters, and other non-printing marks. They can ride along when you copy text from editors, websites, or messaging apps. This guide explains what they are and why they matter. It is not a live checker: to inspect a specific paste, use the AI Text Watermark Scanner.
AI Watermark Center’s Claude Unicode study tests whether owned Claude outputs contain observable hidden characters. That protocol is collecting genuine samples and has not published a dataset yet — missing cells are not fabricated. Official Anthropic documentation already states the statistical text watermark is not based on hidden characters; Unicode findings answer a different question.
Why hidden characters are worth understanding
If you paste text into a form, CMS, or compiler, invisible characters can change sorting, search, security filters, or visual layout. Knowing why they appear helps you interpret a scanner report: listing them is useful even when they have nothing to do with watermarks.
What a scanner should not claim
- Presence of a zero-width character is not proof that a named model produced the text.
- Absence of hidden characters is not proof that the text is human-written.
- Stripping hidden characters is not the same as removing a statistical watermark.
How the scanner works
The scanner is deterministic: it walks Unicode code points, classifies them, and shows them. Conservative Cleanup currently removes only a leading BOM candidate. Cleaning re-scans the result so the page can show before and after rather than asserting an unverifiable “undetectable” state.
Sources
The Unicode Standard — Unicode Consortium
Accessed August 15, 2026.
Related pages
- AI watermark detectorRoutes checker intent by media type — text vs image.
- Claude Unicode research protocolIndependent study of observable Unicode in owned Claude text — collecting; no fabricated samples.
- AI Text Watermark ScannerHidden Unicode and invisible character checker for pasted text — zero-width marks, unusual whitespace, and normalization differences. Runs locally in the browser.
- ClaudeAnthropic documents a statistical text watermark for supported new models and C2PA Content Credentials on supported generated files. AI Watermark Center can inspect Unicode and embedded credentials; it cannot currently detect Claude’s keyed text watermark.
- OpenAIOpenAI documents layered image provenance (C2PA and SynthID) on supported ChatGPT, Codex, and API outputs, SynthID on supported generated audio, a public verifier, and a Content Provenance API. A current OpenAI text watermark is not established. Research profile and image cluster available; AWC inspects embedded C2PA only.
- MethodologyHow this site labels documented, verified, observed, and unverified claims.