Text Encoding Converter
Decode a known text-file encoding, edit the text, and create UTF-8, UTF-16, UTF-32 or ISO-8859-1 bytes. Control endianness, byte-order marks and line endings without silently replacing invalid characters.
Source file or text
Auto recognizes Unicode BOMs; it does not guess unknown legacy encodings.
Up to 200,000 UTF-16 text units. Browser text fields may normalize entered CR/CRLF to LF; choose the output line-ending setting below to create CRLF explicitly. Invisible characters are visible in the hex preview.
See output settings ↓Output file & bytes
Paste text or try an example.
Encoded file information appears here.
The download contains the complete encoded bytes; hex preview is bounded. Empty text produces an empty file, or only a BOM if selected. Latin-1 supports U+0000–U+00FF and rejects other characters; it has no Unicode BOM.
Files are read and encoded locally. No processing upload or saved text/file history. Website ads and analytics follow our privacy information.
Using Text Encoding Converter
Create file bytes in an explicit encoding, with strict validation instead of silent character replacement.
- Paste text or choose a file, select its known encoding and click Decode. Auto recognizes Unicode BOMs; otherwise it tries strict UTF-8.
- Choose output encoding, optional Unicode BOM and line-ending normalization. Generate the encoded file.
- Review byte counts and the bounded hex preview, then download the full file or an encoding report.
Example
Decode a UTF-16LE file with a BOM, then create UTF-8 with CRLF line endings. Try multilingual text with Arabic, Chinese, Japanese and emoji; use a Unicode encoding to preserve it.
Questions & answers
Does Auto guess every encoding?
No. It recognizes UTF-8, UTF-16 and UTF-32 BOMs, checking four-byte BOMs before two-byte prefixes. Without a BOM it uses strict UTF-8. You must select a known legacy or endian encoding when appropriate.
What happens with invalid text or bytes?
Malformed UTF-8, unpaired UTF-16 surrogates, invalid UTF-32 scalar values, incomplete code units and unsupported Latin-1 characters cause an error. No replacement characters are silently inserted.
What is a BOM?
An initial Unicode signature/byte-order mark. You can remove a recognized matching initial BOM on decode and optionally add a new one during encoding. Retaining an existing leading U+FEFF and adding a BOM creates both; the tool reports that situation. A conflicting source BOM raises an error.
How is Latin-1 different from Windows-1252?
This tool implements ISO-8859-1 as the direct U+0000–U+00FF mapping, including C1 controls. It is not Windows-1252 and does not map bytes 80–9F to Windows punctuation. It rejects characters outside Latin-1 and does not add a Unicode BOM.
Are line endings kept exactly?
Unedited decoded source text retains its original CR, LF and CRLF sequences internally. Editing/pasting uses the browser text field, which can normalize line endings to LF. Select explicit LF or CRLF output normalization when needed.
Can I encode empty text?
Yes. It produces an empty file, or just the selected Unicode BOM. Files are limited to 1 MiB and text to 200,000 UTF-16 units; output is also limited to 1 MiB.
Does the hex field contain the entire file?
It shows only the first 512 bytes, in 16-byte rows. The file download contains every encoded byte. The JSON report includes counts and encoding settings without including the text itself.
Can I fix garbled text automatically?
Only if you know and select the correct original encoding. If incorrect decoding has already changed or replaced characters, this tool cannot reconstruct missing original bytes.
Are files uploaded?
Files are read and encoded in this browser with no processing upload or saved text/file history. Website ads and analytics follow the privacy information.
Help improve this tool
Report a problem or suggest an improvement
Describe the issue without pasting private tool input. Feedback goes to our admin inbox.
