ASCII & Unicode
Introduction to ASCII
ASCII, which stands for "American Standard Code for Information Interchange," is a character encoding standard that was developed in the early 1960s to represent text and control characters in computers and communication equipment.
It's one of the foundational building blocks of modern computing and is still widely used today, though it has been largely supplemented by more advanced character encoding standards like UTF-8.
Key points
- ASCII character encoding assigns a unique number to letters, numbers and other printable characters like spaces, commas and line breaks.
- Represented using 7-bit binary - 128 possible characters
- Only works for English language.
What does ASCII stand for?
What is the range of decimal values that the standard ASCII encoding can represent?
ASCII Table
Here is an ASCII table with all the associated charaters and their denary equivalent.
What is the ASCII value of 'A'?
What is the ASCII value of the uppercase letter 'Z'?
Limitations of ASCII
ASCII is limited to representing a relatively small set of characters. It primarily includes the characters used in the English language, such as letters (both uppercase and lowercase), numbers, punctuation marks, and some control characters.
Therefore:
- It does not support characters from other languages or writing systems.
- It doesn't have support for emojis or other special characters.
Which of the following is NOT a limitation of the ASCII character encoding?
Extended ASCII
One of the key limitations of ASCII is that it only worked with the English Alphabet. As computers were more widely available ASCII was extended to 8 bits, allowing another 128 characters. This was known as extended ASCII and included other European languages.
128 à 129 Ãŧ 130 Ê 131 Ãĸ 132 ä 133 à 134 ÃĨ 135 ç
136 ÃĒ 137 ÃĢ 138 è 139 ï 140 ÃŽ 141 ÃŦ 142 à 143 Ã
144 à 145 ÃĻ 146 à 147 ô 148 Ãļ 149 Ã˛ 150 Ãģ 151 Ú
152 Ãŋ 153 Ã 154 Ã 155 Âĸ 156 ÂŖ 157 ÂĨ 158 â§ 159 Æ
160 ÃĄ 161 Ã 162 Ãŗ 163 Ãē 164 Ãą 165 Ã 166 ÂĒ 167 Âē
168 Âŋ 169 â 170 ÂŦ 171 ÂŊ 172 Âŧ 173 ÂĄ 174 ÂĢ 175 Âģ
176 â 177 â 178 â 179 â 180 ⤠181 ⥠182 âĸ 183 â
184 â 185 âŖ 186 â 187 â 188 â 189 â 190 â 191 â
192 â 193 â´ 194 âŦ 195 â 196 â 197 âŧ 198 â 199 â
200 â 201 â 202 ⊠203 âĻ 204 â 205 â 206 âŦ 207 â§
208 ⨠209 ⤠210 âĨ 211 â 212 â 213 â 214 â 215 âĢ
216 âĒ 217 â 218 â 219 â 220 â 221 â 222 â 223 â
224 Îą 225 Ã 226 Î 227 Ī 228 ÎŖ 229 Ī 230 Âĩ 231 Ī
232 ÎĻ 233 Î 234 Ί 235 δ 236 â 237 Ī 238 Îĩ 239 âŠ
240 ⥠241 Âą 242 âĨ 243 ⤠244 â 245 ⥠246 Ãˇ 247 â
248 ° 249 â 250 ¡ 251 â 252 âŋ 253 ² 254 â 255
Unicode
As computers became avialable worldwide, Extended ASCII was no longer sufficient and so Unicode was introduced, with up to 21 bits in total.
Here are 10 examples of Unicode characters from different scripts, symbol sets, and emoji ranges:
Example Unicode Name / Description
A U+0041 Latin Capital Letter A
Ί U+03A9 Greek Capital Letter Omega
Đ U+0416 Cyrillic Capital Letter Zhe
× U+05D0 Hebrew Letter Alef
Ų U+0645 Arabic Letter Meem
⤠U+0915 Devanagari Letter Ka (used in Hindi)
æĨ U+65E5 CJK Ideograph âSun/Dayâ
â° U+23F0 Alarm Clock emoji
You can find many more examples here:
Unicode
Unicode is a character encoding standard designed to represent and handle text and symbols from virtually all writing systems in the world. It provides a unified and consistent way to encode characters and symbols, regardless of language, script, or platform.
Character Set
Unicode assigns a unique number (called a code point) to every character, symbol, and diacritic used in human writing systems, including Latin, Greek, Cyrillic, Arabic, Chinese, Japanese, and many others (including emojos!â¤ī¸ ). It aims to cover all written languages and scripts worldwide.
Code Points
Each character in Unicode is identified by a unique code point, typically represented in hexadecimal format (e.g., U+0041 for the Latin letter 'A'). Unicode currently defines over 143,000 code points, with room for expansion.
Unicode Bit Length
Unlike ASCII, Unicode can be represented using various bit sizes, depending on the encoding scheme chosen. The most common Unicode encoding schemes are :
UTF-8
Unicode characters in UTF-8 are encoded using 8, 16, 24, or 32 bits, depending on the specific character being encoded. This is the most common format as it is the most space efficient.
UTF-16
Unicode characters in UTF-16 are encoded using either 16 or 32 bits.
UTF-32
Unicode characters in UTF-32 are encoded using a fixed 32 bits (4 bytes) for each character. This encoding scheme provides a straightforward and fixed-length representation for all Unicode characters.
What is Unicode in the context of programming?
Review: Fill in the Blanks
The ASCII character encoding assigns a unique to letters, numbers, and other printable characters, allowing for efficient text representation. It is represented using , which allows for a total of , but it only supports the , limiting its use for global communication.
To address the limitations of ASCII, an extension called was introduced, which uses for encoding and supports an additional . This allowed for the representation of some characters from other , broadening its application beyond just English.
As the need for a more comprehensive character encoding arose globally, was developed, accommodating a vast range of writing systems and symbols from around the world. Unicode assigns a unique to every character and can be represented using various bit sizes, with UTF-8 being the most common encoding scheme due to its .
Complete! Ready to test your knowledge?
ASCII
- Introduction to ASCII
- ASCII Table
- Limitations of ASCII
- Extended ASCII
- Unicode
Unicode
- Unicode
- Unicode Bit Length