Remove Invisible Characters
This free tool helps you remove invisible characters from any text: zero-width spaces, bidirectional controls, tag characters, soft hyphens, non-breaking spaces and other hidden code points that you cannot see but that quietly break search, forms, URLs and code. Paste your text, review an itemised audit of every hidden character found, then strip them with a single click and copy or download the clean result.
It works entirely in your browser. Your text is never uploaded to any server, there is nothing to install, and nothing you paste ever leaves your device.
What to remove (all on by default)
Your text is processed entirely on your device with JavaScript — nothing is uploaded, logged or sent anywhere. Watch your browser’s Network tab: this page makes zero requests when you clean text.
How to use the invisible character remover
- Paste or type your text into the box. The tool scans it instantly on your own device.
- Read the audit panel. It lists every hidden character by name, Unicode code point and count, and highlights each one inline so you can see exactly where it sits.
- Use the per-category toggles to decide what to strip, then apply the clean-up.
- Copy the cleaned text or download it as a file. Your original stays in the box until you clear it.
What are invisible characters?
Invisible characters are real Unicode code points that occupy no visible space, or that look like an ordinary space or letter while behaving differently underneath. They are legitimate parts of the Unicode standard, used for line-breaking hints, text direction, emoji construction and multilingual typesetting. The trouble starts when they end up somewhere they were never meant to be, because your eye cannot spot them but a computer treats them as distinct characters. A password, a coupon code or a URL that looks perfectly correct can fail simply because a zero-width space is hiding inside it.
The most common culprits fall into a handful of families. The table below is a legend of the ones this tool detects, where they usually come from, and what they do.
| Character | Code point | Where it comes from |
|---|---|---|
| Zero-width space (ZWSP) | U+200B | Copy-paste from web pages, word processors, some AI output |
| Zero-width non-joiner (ZWNJ) | U+200C | Arabic, Persian and Indic scripts; also injected as a marker |
| Zero-width joiner (ZWJ) | U+200D | Emoji sequences and Indic scripts |
| Word joiner / BOM | U+2060, U+FEFF | Byte-order marks from text files and exports |
| Bidirectional controls | U+200E, U+200F, U+202A–U+202E, U+2066–U+2069 | Right-to-left text, and occasionally spoofing attacks |
| Tag characters | U+E0000–U+E007F | Hidden payloads in some AI and emoji exploits |
| Variation selectors | U+FE00–U+FE0F | Emoji styling; can be misused to hide data |
| Soft hyphen | U+00AD | Word processors adding optional line-break points |
| Non-breaking & special spaces | U+00A0, U+2000–U+200A, U+202F, U+205F, U+3000 | Layout spaces from PDFs, Word and web copy |
Why hidden characters cause problems
Because they are invisible, these characters break things in ways that are maddening to debug. A find-and-replace that should match refuses to, because a zero-width space splits the word in two. A form rejects an email address or a promo code that looks identical to the valid one. A URL you paste returns a 404 because a byte-order mark rode along at the front. Code fails to compile, a JSON key never matches, or a database lookup for a username comes back empty because one record carries a non-breaking space and the other does not. Word counts inflate, columns misalign in a spreadsheet, and two strings that look the same simply refuse to be equal. Stripping the hidden characters restores plain, predictable text that behaves the way it looks.
Cleaning text pasted from AI, PDFs and the web
Beyond stripping zero-width and control characters, the tool offers two optional passes. Invisible-space normalisation turns non-breaking and typographic spaces (U+00A0, U+2000–U+200A, U+202F, U+205F, U+3000) into ordinary spaces, which fixes most alignment and matching issues from PDF and Word copy. The “AI text tidy” pass converts smart quotes to straight quotes and em and en dashes to plain hyphens, handy when you need portable plain text for code or spreadsheets. Both are off by default, in case you want to keep the original punctuation.
Do ChatGPT, Claude and Gemini add invisible characters?
This is the fast-growing question, and it deserves an honest answer rather than a scare story. Text produced by AI assistants often does contain characters people class as “hidden”, but usually not the ones the rumours suggest. What you will genuinely find is smart typography (curly quotes, em dashes and non-breaking spaces), and zero-width joiners as a normal part of emoji. Some prompts, browser extensions and watermarking experiments do inject zero-width or tag characters as a fingerprint, so it is worth checking. What is not true is that every AI model stamps a reliable, countable invisible-character watermark you can detect by pasting text into a tool. There is no such universal marker. Google’s SynthID, for instance, is a statistical watermark applied across the pattern of tokens a model chooses; it is not a run of hidden characters, and clearing invisible characters will not reveal or remove it. Clean and normalise AI text with confidence, but treat any claim that a tool can “prove” text was AI-written with healthy scepticism.
Look-alike letters and homoglyph spoofing
Not every problem character is invisible; some are visible but wrong. Homoglyphs are letters from other alphabets that look identical to Latin ones. The Cyrillic “а” (U+0430) is indistinguishable from a Latin “a”, and Greek and Cyrillic offer look-alikes for many letters. Attackers use them to spoof domain names and brand names in phishing, and they sneak into copied code where they break identifiers that appear correct on screen. The optional homoglyph fixer maps common Cyrillic and Greek confusables back to their Latin equivalents, so “аpple” becomes “apple” again. Turn it on when you are auditing pasted usernames, URLs or source code; leave it off if your text is genuinely multilingual.
A language-safe warning before you strip
Zero-width joiners and non-joiners are not always noise. In Arabic, Persian, Hindi, Thai, Khmer and Myanmar text, among others, U+200C and U+200D are legitimate and change how letters connect and render. Stripping them from genuine multilingual content can corrupt words or alter their meaning. If your text contains any of these scripts, review the audit list carefully and untick the zero-width joiner and non-joiner categories before cleaning, so you remove only the characters that truly do not belong.
Developer snippets: strip zero-width characters in code
If you would rather handle this in your own pipeline, the core operation is a single regular expression targeting the zero-width range plus the byte-order mark. Here are copy-paste equivalents for three common languages.
JavaScript:
const clean = text.replace(/[-]/g, '');
Python:
import re
clean = re.sub(r'[]', '', s)
PHP:
$clean = preg_replace('/[\x{200B}-\x{200D}\x{FEFF}\x{2060}]/u', '', $text);
Extend the character class to cover bidirectional controls, tag characters or variation selectors if your data needs it, and remember the language-safe caveat above before you apply it to multilingual content.
Frequently asked questions
What exactly counts as an invisible character?
Any character that renders with no visible glyph, or that looks like a plain space or letter while behaving differently. That includes zero-width spaces and joiners, byte-order marks, bidirectional controls, tag characters, variation selectors, the soft hyphen, and the family of non-breaking and typographic spaces. The audit panel names each one it finds so you always know what you are removing.
Is my text uploaded anywhere when I use this cleaner?
No. The entire scan and clean-up runs locally in your browser, so nothing you paste is sent to a server or stored online. That matters because the text you are cleaning may be private, and uploading it to strip a few hidden characters would defeat the point. Your content stays on your device from start to finish.
Can this tool detect AI-generated text?
Not reliably, and neither can any other invisible-character checker. Some AI text carries hidden or smart characters, but there is no universal invisible watermark to count. Model watermarks such as SynthID work on token patterns, not hidden characters. Use this tool to clean text, not to prove where it came from.
Why does my text look the same before and after?
That is expected. Invisible characters have no visible glyph, so removing them usually leaves the readable text looking identical. The difference shows up when the text is searched, validated, compared or run as code. The audit count confirms how many hidden characters were present and removed.
Will this break emoji or other languages?
It can if you strip the wrong categories. Zero-width joiners build many emoji, and they are meaningful in Arabic, Indic and several South-East Asian scripts. Review the audit list and untick the zero-width categories when your text contains emoji sequences or those languages, so you only remove characters that are genuinely out of place.
What is the difference between a normal space and a non-breaking space?
They look identical but are different code points. A normal space is U+0020; a non-breaking space is U+00A0 and is common in text copied from PDFs and Word. Because the code points differ, string comparisons and searches can fail. The invisible-space normalisation option converts non-breaking and other special spaces back to ordinary spaces.
Does removing invisible characters change my word count?
It can. Some hidden characters split a single word into two as far as a word-count algorithm is concerned, inflating the total, while stray spaces can add phantom words. Cleaning the text gives you an accurate count. Pair this with our word counter if you need the exact figure afterwards.
How do I check for hidden characters without removing them?
Just paste your text and read the audit panel. The tool acts as an invisible character detector first: it lists and highlights every hidden character it finds, with names and code points, before you strip anything. You can review the report and only then decide which categories to remove.
Related design tools
- Whitespace Remover — trim extra spaces, tabs and blank lines from text.
- Case Converter — switch text between upper, lower, title and sentence case.
- Find and Replace — swap words or patterns across a whole block of text.
- Word Counter — get an accurate word, character and line count.
- Remove Duplicate Lines — strip repeated lines from lists and logs.
- Invisible Character Generator — the opposite tool, for copying a blank character.
- EXIF & Metadata Stripper — remove hidden metadata from images.