Skip to content
Categories:

I❤️U2 vs. 我也爱你

Post date:
Author:
Number of comments: no comments

This problem of efficiency of language conversion into unicode – and hence anything computer-related and especially AI related – isn’t going away. I’ll continue using the same Unicode Converter and pick up from I❤️U1 where I left it last week. It’s often the case that I❤️U2 is the response, so I added the newly-minted logogram 2 that means “too” to us, even though unicode is simply swapping the logogram with the unicode for the numeral “2” and not the meaning of “too” that’s encoded as the Chinese logogram . Below, all the underlined unicode for the English is false because it doesn’t code what we perceive the logogram to mean. Now I think about it, this also includes the logogram I because the single-letter logogram “I” doesn’t always mean the first person.

I❤️U
0049 2764 FE0F 0055 (50% false)
爱你
6211 7231 4F60
I❤️U2
0049 2764 FE0F 0055 0032 (60% false)
我也爱你
6211 4E5F 7231 4F60

Making a logographic writing system isn’t easy, let alone one that can be unambiguously converted into unicode. We know about emoji and the ❤️ logogram has a precise meaning that converts into unambiguous code. However, we wouldn’t use the emoji 💔 to convey the idea of heart disease. A logogram for that would look more like the one at left because we only understand 💔 as a noun describing a particular emotional state. We don’t use it as a verb in a sentence such as I💔U although one day it might acquire the meaning “I don’t love you” or “I unlove you” rather than the standalone “I am heartbroken” meaning it has now. Minefield. Let’s see what other logograms the English language might have.

ACRONYMS

NY is one candidate, as are all other acronyms. The trouble is, unicode converts letters as individual logograms and not their combined meaning as a new logogram. Just look how the unicode for LLM has two L’s and an M? All the underlined unicode below is false because it doesn’t convey what we understand the acronym/logogram to mean. 大语模 could well be a Chinese acronym for 大型语言模型 but, even if it isn’t, people’s first guess would still be that it has something to do with something called a large-language-model, whatever that is.

NY
004E 0059
纽约
7EBD 7EA6
LLM
004C 004C 004D
大型语言模型
5927 578B 8BED 8A00 6A21 578B
UN
0055 004E
联合国
8054 5408 56FD
USA
0055 0053 0041
美国
7F8E 56FD

CHEMICAL NOTATION

The previous post had the example of “Perchloroethylene is a volatile organic compound.” Substituting the chemical name for Perchloroethylene gives a string of false unicode, even if we (or at least a chemist) understood it as a logogram much as we do H2O. There’s no point using VOC as an acronym as it only adds false unicode. The third example below has unicode the same length as the Chinese sentence but thirteen fifteenths of it is false. Chemistry notation is a class of acronym that makes chemical statements that can be perceived as single logograms whose meaning never converts into unicode. There’s no need to even test for this in Chinese. Using logograms such as VOC for volatile organic compound makes it seem we wanted written English to be more concise even before the world was unicoded.

Perchloroethylene is a volatile organic compound.
0050 0065 0072 0063 0068 006C 006F 0072 006F 0065 0074 0068 0079 006C 0065 006E 0065 0020 0069 0073 0020 0061 0020 0076 006F 006C 0061 0074 0069 006C 0065 0020 006F 0072 0067 0061 006E 0069 0063 0020 0063 006F 006D 0070 006F 0075 006E 0064 002E
四氯乙烯是个挥发性有机化合物。
56DB 6C2F 4E59 70EF 662F 4E2A 6325 53D1 6027 6709 673A 5316 5408 7269 3002
C2Cl4 is a volatile organic compound.
0043 0032 0043 006C 0034 0020 0069 0073 0020 0061 0020 0076 006F 006C 0061 0074 0069 006C 0065 0020 006F 0072 0067 0061 006E 0069 0063 0020 0063 006F 006D 0070 006F 0075 006E 0064 002E
C2Cl4 is a VOC.
0043 0032 0043 006C 0034 0020 0069 0073 0020 0061 0020 0056 004F 0043 002E

We’re not getting anywhere trying to make written English more computationally efficient. Devising an unambiguous set of logograms for written English isn’t something that can happen overnight and, even if it were, there’s still the significant (as in, really significant) problem of teaching everyone, native speakers and non-native speakers alike, this new way of writing English. The point of this post – this exercise – is directed towards the more efficient transfer of information from the human world to, let’s call it, “AI”.

Many languages are transliterated into an English/Latin phonetic system to help non-native speakers learn to make approximations of the sounds. It’s imperfect, but better than nothing. Japanese can be written as Latin-alphabet “romaji” [Roman letters] – which becomes nihongo wa romaji de kaku koto ga kanoo desu. More or less. The second last word kanoo (meaning possible) has a long “o” which is often represented as ō as it’s a different sound, and not “o” repeated. The name Tokyo, technically, should be written Tōkyō. It was never pronounced as Tokeeo and never written as Tokio except for when the US occupation forces were there. But every language has sounds not present in other languages and so romanizations such as the Japanese one have to either lose information or add ways of expressing different but necessary information. Over the years, there have been many systems for rewriting Chinese2 but one dominant one using a “Latin” or “Roman” alphabet was Wade-Giles romanization that was standard until 1982 when the Pinyin system replaced it. Part of the reason for this was that Pinyin was more accurate. It’s the difference between Wade-Giles’ Peking and Pinyin’s Běijīng. The Pinyin system includes information on Mandarin’s four tones, respectively known as:

  • Tone 1: a flat tone, represented as e.g. flāt)
  • Tone 2: a rising tone, represented as e.g. rísíng
  • Tone 3: a falling followed by a rising tone, represented as fǎlling or rǐsing
  • Tone 4: a sharp falling tone, represented as e.g. fàllìng

It turns out that Beijing people had been saying Běijīng (北京, to Beijing speakers) all along. For computing purposes, there’s no point transliterating a language but, although it doesn’t happen very often, the system of logograms a language is written in, is not as fixed as we think. Adopting a new system of writing a language happens for various historical reasons usually to do with politics, commerce, conquest and identity, inasmuch as they can be separated. To start with, let’s consider languages that adopted the Arabic script. Dates are approximate.

  • Persian3 4(spoken as Farsi in Iran; as Dari in Afghanistan; and Tajik in Tajikistan) (600CE onwards)
    • Before Arabic, Persian was written using Sanskrit script (400–600CE)
    • Before Sanskrit, Persian was written in Aramaic script (300BCE–400CE)
    • Before Aramaic, Persian was written using Cuneiform which goes back to Mesopotamia and was most likely adopted to facilitate trade with the Gulf and eastern Mediterranean countries (1500–300BCE). FYI Aramaic is the shared root of both Hebrew and Arabic.

Below is the Persian alphabet written in Cuneiform (left), Aramaic script (middle) and Sanskrit script (right).

Here’s Persian writting using the Arabic script (left), and the cursive Perso-Arabic variant script (right).

Other languages that shifted to Arabic script are:

  • Pashto (spoken in Afghanistan and Pakistan)
  • Kurdish (spoken in parts of Iran, Iraq, Syria, and Turkey)
  • Urdu (spoken in Pakistan and parts of India)5

It’s not just about changing to Arabic.

  • Turkish used to be written in Arabic script but in 1928 the decision was made to change to a Latin alphabet because, I can imagine, the use of a Latin alphabet represented a break with the past and a symbol of modernization.
  • Vietnamese used to be written in Chinese logograms with one for each letter of its alphabet until the early 20th century when a Latin alphabet was adopted. I can see why. If Chinese logograms have no meaning other than as alphabet and syllables representing sounds (a.k.a. phonemes), there’s no point to using them.

And it’s not just about changing to “Roman” alphabetization.

  • For a start, written English uses the Roman alphabet that was used to write Latin and that was adopted and adapted from the Greek alphabet. The English alphabet consists of 26 upper-case letters descended from the Roman letters apart but J and U and W that were added in the mediaeval period . This is why “August” was written as AVGVST. Romans were always shouting at each other. The English language lower-case letters come to from the Carolingans circa 800CE give or take. Inventing them may have saved space when illuminating manuscripts 750–850CE. Lower case letters are a more compact system of notation and could mean the Carolingans were bumping up against the same speed-of-input problem that written English is bumping up against now. This lower case/upper case difference may be archaic and redundant, but it has already been “baked into” unicode where upper and lower case are treated equally.

Changing the system of writing of a language can’t be easy, but it can and does happen when circumstances make it seem the better option. Now might be one of those times for English as it is currently written. Going back to the unspoken problem of the English language’s inefficiency of conversion to unicode, let’s see what happens if we keep English as a distinct language yet write it in Chinese “script”. As a sample sentence, I’ll use It is possible to use Chinese characters to write English. If we keep the English language word order and grammar, we get 是可能使用汉字为写英语。Here’s how it goes. This sentence is not that different from Chinese langage word order.

ItispossibletouseChinese characters(in order) towriteEnglish.
可能使用汉子英语

It reads Is possible use Chinese characters to write English. This sounds familiar to us from movie stereotypes or perhaps from hearing how people in Chinatowns around the world speak English. Nothing’s lost except a few things like prepositions and articles that the Chinese language manages to do without, along with plurals, gendered nouns and verb conjugations and it’s the very absence of those that make English spoken by some Chinese people sound either pidgin or abrupt to people whose first language is English. In the table below, the top line is how English is now spoken, the middle line is how is how it could be written, and the bottom line is how it would be unicoded if it were. Information flow isn’t about typesetting and typewriters anymore but we can think of unicode as the new typewriting or typesetting because this problem of the speed of conveying information hasn’t gone away.

IspossibleuseChinese characters(in order) towriteEnglish.
可能使用汉子英语
662F3EF 80FD4F7F 75286C49 5B57 4E3A51994E3A 8BED3002

There’s still the problem of visually discriminating between the verb “use” the verb and the noun “use” but we can fix this by using ū for the long “u” in the same way as the same problem was dealt with when romanizing the Japanese language.

Adding spaces to separate words is a non-starter as it needlessly increases inefficiency . The meaning of the next logogram will tell you if it formed a new logogram with the one before.

是 可能 使用 汉字 为 写 英语
662F 0020 53EF 80FD 0020 4F7F 7528 0020 6C49 5B57 0020 4E3A 0020 5199 0020 82F1 8BED 3002
是可能使用汉字为写英语
662F 53EF 80FD 4F7F 7528 6C49 5B57 4E3A 5199 82F1 8BED 3002

Korean and Japanese mix Chinese logograms with local alphabets adding unnecessary conjugations, particles and other elements of grammar.

Only when the English language is rewritten as a logographic language will it have computational parity of efficiency for processes dependent upon language i.e. all of them. The logograms don’t have be the same ones the Chinese people use but, right now, they’re the only ones we have and roughly one sixth of the world’s population already use them.

Written English
Hex/UTF-32
Written Chinese
Hex/UTF-32
I
0049

6211
you
0079 006F 0075

4F60
encyclopaedia
0065 006E 0063 0079 0063 006C 006F 0070 0061 0065 0064 0069 0061
百科全书
767E 79D1 5168 4E66
antidisestablishmentarianism
0061 006E 0074 0069 0064 0069 0073 0065 0073 0074 0061 0062 006C 0069 0073 0068 006D 0065 006E 0074 0061 0072 0069 0061 006E 0069 0073 006D
反政教分离主义
53CD 653F 6559 5206 79BB 4E3B 4E49
It is possible to use Chinese characters to write English.
0049 0074 0020 0069 0073 0020 0070 006F 0073 0073 0069 0062 006C 0065 0020 0074 006F 0020 0075 0073 0065 0020 0043 0068 0069 006E 0065 0073 0065 0020 0063 0068 0061 0072 0061 0063 0074 0065 0072 0073 0020 0074 006F 0020 0077 0072 0069 0074 0065 0020 0045 006E 0067 006C 0069 0073 0068 002E
是可能使用汉字为写英语
662F 53EF 80FD 4F7F 7528 6C49 5B57 4E3A 5199 82F1 8BED 3002
是可能使用汉字为写英语。(English rewritten using Chinese characters)
662F 53EF 80FD 4F7F 7528 6C49 5B57 4E3A 5199 82F1 8BED 3002
是可能使用汉字为写英语
662F 53EF 80FD 4F7F 7528 6C49 5B57 4E3A 5199 4E3A 8BED 3002

• • •

Recently Revisited:

Further Information:

Notes:

  1. https://misfitsarchitecture.com/2025/12/07/i%e2%9d%a4%ef%b8%8fu-vs-%e6%88%91%e7%88%b1%e4%bd%a0/ ↩︎
  2. https://en.wikipedia.org/wiki/Transliteration_of_Chinese ↩︎
  3. https://en.wikipedia.org/wiki/Persian_language ↩︎
  4. https://persian.religion.ucsb.edu/home/history-of-persian/ ↩︎
  5. https://www.metmuseum.org/learn/educators/curriculum-resources/art-of-the-islamic-world/unit-two/the-arabic-alphabet-and-other-languages ↩︎

Leave a Reply