search
HomeCommon ProblemThe unicode character set uses several bytes to represent a character

The unicode character set uses several bytes to represent a character

May 07, 2021 pm 04:43 PM
unicodecharactercharacter setbyte

The unicode character set uses 2 bytes to represent a character. Unicode sets a unified and unique binary encoding for each character in each language to meet the requirements for cross-language and cross-platform text conversion and processing; it can unify all texts in the world using 2 bytes coding.

The unicode character set uses several bytes to represent a character

The operating environment of this tutorial: Windows 7 system, Dell G3 computer.

The unicode character set uses 2 bytes to represent a character.

Unicode (Unicode, Universal Code, Unicode) is a character encoding used on computers. It sets a unified and unique binary encoding for each character in each language to meet the requirements for cross-language and cross-platform text conversion and processing.

If various text encodings are described as dialects from various places, then Unicode is a language developed cooperatively by countries around the world.

In this language environment, there will be no more language encoding conflicts. Content in any language can be displayed on the same screen. This is the biggest benefit of Unicode. It means that all the text in the world is uniformly encoded using 2 bytes. In that way, with unified encoding like this, 2 bytes are enough to accommodate most text in all languages ​​​​in the world.

The scientific name of Unicode is "Universal Multiple-Octet Coded Character Set", referred to as UCS.

The early Unicode standards were called UCS-2 and UCS-4. UCS-2 is encoded with two bytes, and UCS-4 is encoded with 4 bytes. What is currently used is UCS-2, which is a 2-byte encoding, and UCS-4 was developed to prevent 2 bytes from being insufficient in the future.

UCS-4 is divided into 2^7=128 groups according to the highest byte with the highest bit being 0. Each group is divided into 256 planes according to the next highest byte. Each plane is divided into 256 rows according to the third byte, and each row has 256 code points (cells). Plane 0 of group 0 is called BMP (Basic Multilingual Plane). UCS-2 is obtained by removing the first two zero bytes of UCS-4's BMP.

For more related knowledge, please visit the FAQ column!

The above is the detailed content of The unicode character set uses several bytes to represent a character. For more information, please follow other related articles on the PHP Chinese website!

Statement
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Hot AI Tools

Undresser.AI Undress

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress AI Tool

Undress images for free

Clothoff.io

Clothoff.io

AI clothes remover

AI Hentai Generator

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)
4 weeks agoBy尊渡假赌尊渡假赌尊渡假赌
R.E.P.O. Best Graphic Settings
4 weeks agoBy尊渡假赌尊渡假赌尊渡假赌
R.E.P.O. How to Fix Audio if You Can't Hear Anyone
4 weeks agoBy尊渡假赌尊渡假赌尊渡假赌
WWE 2K25: How To Unlock Everything In MyRise
1 months agoBy尊渡假赌尊渡假赌尊渡假赌

Hot Tools

Zend Studio 13.0.1

Zend Studio 13.0.1

Powerful PHP integrated development environment

DVWA

DVWA

Damn Vulnerable Web App (DVWA) is a PHP/MySQL web application that is very vulnerable. Its main goals are to be an aid for security professionals to test their skills and tools in a legal environment, to help web developers better understand the process of securing web applications, and to help teachers/students teach/learn in a classroom environment Web application security. The goal of DVWA is to practice some of the most common web vulnerabilities through a simple and straightforward interface, with varying degrees of difficulty. Please note that this software

EditPlus Chinese cracked version

EditPlus Chinese cracked version

Small size, syntax highlighting, does not support code prompt function

SublimeText3 Mac version

SublimeText3 Mac version

God-level code editing software (SublimeText3)

Safe Exam Browser

Safe Exam Browser

Safe Exam Browser is a secure browser environment for taking online exams securely. This software turns any computer into a secure workstation. It controls access to any utility and prevents students from using unauthorized resources.