| Char | Code Point | Decimal | Block | UTF-8 | UTF-16 | HTML Entity |
|---|
| Char | Code Point | Decimal | Block | UTF-8 | UTF-16 | HTML Entity |
|---|
Every character a computer can display — letters, digits, punctuation, CJK characters, emoji — is assigned a unique number called a code point in the Unicode standard, usually written as "U+" followed by hexadecimal digits (like U+0041 for "A"). This tool breaks down any text you type into its individual characters and shows each one's code point, decimal value, Unicode block, and byte representation in the two most common encodings, plus lets you look up a character from its code.
For text input, the tool iterates through each character, reads its code point using JavaScript's built-in Unicode-aware string methods, and looks up which named Unicode block that code point falls in (like "Basic Latin" or "CJK Unified Ideographs") from a reference table. It also computes the raw bytes for UTF-8 and UTF-16 encoding and the equivalent HTML numeric character entity. The reverse lookup accepts a code point in several formats (U+XXXX, 0x hex, plain decimal, or a single character) and looks up the same details.