Character Encoding Converter
Conversion
1-4 bytes per code point, represents every Unicode character. The default for the web.
Result
Type or paste text above — the bytes appear as you type.
What this does and cannot do
Every codec is written out in this bundle: the Windows-1252 block, Latin-1, strict ASCII, UTF-8, UTF-16 and UTF-32. That matters because browser labels are not literal — a decoder asked for "ascii" or "latin1" silently returns Windows-1252. Here the label says what the bytes are.
Single-byte encodings cannot hold most of Unicode, so encoding reports each substitution instead of hiding it, and decoding flags bytes that are not valid with a count. Lossy conversion is a one-way door: the original characters cannot be recovered from the "?" bytes.
This is not a file inspector: it converts a string you paste. It does not detect an unknown file's encoding, and it does not attach a byte-order mark. Everything runs in the browser — nothing is uploaded.
This character encoding converter turns text into bytes and bytes back into text across eight encodings — UTF-8, UTF-16, UTF-32, Windows-1252, Latin-1 and strict ASCII — and reports every character a code page cannot carry.
What is Character Encoding Converter?
Most encoding tools hide their mapping behind the platform. This character encoding converter writes every codec out in the page bundle instead, so a byte you see is a byte you can verify: UTF-8, UTF-16 in both byte orders, UTF-32, Windows-1252, true ISO-8859-1 and strict US-ASCII. Nothing is delegated to a decoder label, because browser labels are not literal — a decoder asked for "latin1" or "ascii" actually returns Windows-1252. Here the label means what it says, and every encoding conversion the tool performs can be verified byte by byte.
The converter works in both directions. In text-to-bytes mode you paste a string, choose an encoding and read the result as hex pairs, Base64 or escaped \xNN sequences, with a hex-dump table showing offsets and an ASCII preview. In bytes-to-text mode you paste hex, Base64 or escaped bytes, choose the encoding they are assumed to be in, and read the decoded text together with a count of any replacement characters. The two modes share the same codec, so a round trip is a real check: encode café — €10 and decode the result, and you get back exactly what you typed as long as the encoding can hold it.
Lossy encoding is where most tools quietly damage text, so this one reports it. Windows-1252 cannot hold an emoji, ISO-8859-1 cannot hold €, and strict ASCII cannot hold é; each unrepresentable character is written as a question mark and listed with its code point and position. The substitution is deliberate and visible, which also means the original cannot be recovered from the output — the page says so rather than implying the conversion was harmless.
Decoding reports the other failure mode. Bytes that do not form valid text — a truncated UTF-8 sequence, an unpaired UTF-16 surrogate, an out-of-range UTF-32 value, a high byte in strict ASCII — become U+FFFD and are counted, so mojibake is explained rather than displayed as mystery characters. A UTF-8 byte order mark decodes as an ordinary character, and the converter never adds one, because a tool that silently inserts a BOM into generated bytes causes its own class of bugs.
The numbers are checkable by hand, which is the point of writing the mappings out. The UTF-8 bytes of é are C3 A9; the UTF-8 bytes of an emoji are four; a surrogate pair is four UTF-16 bytes; Windows-1252 puts € at 0x80 and the curly quotes at 0x91-0x94; Latin-1 maps byte to code point one to one. Those are standard facts, and the unit tests against them are part of the repository, so a regression in the mapping fails the build rather than shipping.
What this is not: a file inspector, an encoding detector, or a repair tool. It converts a string you paste and shows you the bytes; it cannot look inside an uploaded file, infer an unknown encoding, or recover characters that were already replaced by question marks. It is also not encryption — an encoding is a mapping, not a secret. Everything runs locally in the browser, and no data is uploaded.
Use it when you need to see exactly what your app is sending, to produce a byte sequence for a legacy system, to confirm which characters fit inside a fixed code page before a migration, or to work out why a string arrived as mojibake. The combination of an explicit codec, a visible byte table and honest lossy reporting makes the answer something you can paste into a bug report.
How to use Character Encoding Converter
- Choose a direction: text to bytes for encoding, bytes to text for decoding.
- Pick the encoding — UTF-8 is the default; Windows-1252 and Latin-1 behave differently and the page describes each.
- In encode mode, read the hex, Base64 or escaped output and check any replaced-character list.
- In decode mode, paste bytes in the matching format and read the text plus its replacement count.
- Copy or download the result; adjust the encoding if the byte count or the flag count surprises you.
When to use Character Encoding Converter vs related tools
Use this converter when you need to see or produce concrete bytes, and reach for the neighbouring tools when the question is about characters rather than codecs. The Unicode Text Normalizer handles composition and compatibility forms — the same visible text can have different code points — and the Language Detector identifies which language a sample is in when the bytes decode correctly but you are unsure what it says.
For counting what a string contains before encoding it, the character count tool is the simpler first stop; come here when the byte level is the problem.
Privacy & Security
This tool runs entirely in your browser — no data ever leaves your device. There is no server round-trip, no upload, no logging, and no account required. Your input is processed locally using client-side JavaScript and is never stored, transmitted, or accessible to anyone else. When you close the tab, everything disappears.