Skip to content
Categories:

I ♥️ LLM

Post date:
Author:
Number of comments: 6 comments

Just prior to Chinese New Year on January 26 came the announcement of the next Large Language Model AI assistant, Hangzhou Tencent’s DeepSeek. Attention immediately focussed on how it could have a performance comparable to, say, oldsters like ChatGPT using less powerful chips when we’d been led to believe that increasingly higher powered chips were crucial.

Immediately after the Deepseek announcment, a Google search would offer images of DeepSeek logos in Chinese in China (深度求素) as well as English. I used one of them for the header image of that post. Now, three weeks later, any sign of 深度求素 is gone – even on Wikipedia for someone in China and using the Hong Kong version of Google. This is an example of how statistically generated search results skew search results. Each of us has a different filter bubble – there’s no “probably” about this, so try asking “What is the Chinese for Deepseek?”” and see what’s returned. Google of course has no vested interest in promoting a competitor and we have no reason to assume Google or any other search engine is the neutral provider of search results that we relax into thinking it is.

Two posts ago I floated the idea that the reason for DeepSeek’s comparable performance was perhaps it had a more accurate model of the reality in which it is intended to operate. At the time, I’d just learned that DeepSeek had its beginnings as a financial trading algorithm and financial trading algorithms are nothing more or less than models of cause and effect in the financial trading industry. Whatever financial working model/algorithm the founders of DeepSeek produced must have been a fairly accurate one if it generated the capital to fund the training of DeepSeek. Initial media reports were skeptical of the stated mere US$6 billion, thus sustaining the fallacy that the more money you throw at something the better it must be. Rephrasing this, it suggests that any innovation of worth is proportional to R&D and capital expenditure (capex) and, accordingly, share price. Inevitably, Hangzhou Tencent’s claim of US$6 billion has been variously disputed and, in a case of the pot calling the kettle beige, San Altman has made allegations of theft of “training data”. All or none of this may be true but the financial markets that, in our neoliberal world, seem to be our ultimate arbiters of truth, gave DeepSeek’s developers the benefit of the doubt and share prices in the then AI-dominant companies duly crashed. My extremely simple model of how financial markets work goes something like “Hey, we developed a cure for cancer!” (Share price goes up.) “Oops, no we didn’t.” (Share price goes down.)

Modeling financial markets is something very goal-specific ($$$) but, before anything is coded, there has to be astute human observation of cause and effect in those financial markets. Coding is not magic. All it does is translate human observations and perceptions, whether accurate or not, into machine language and it is important those human observations and perceptions of causes and effects be recognized and code-able, if not necessarily understood. Everything is a working model. Skilled traders intuitively develop their own, and the success of a financial (or commodity) trader depends upon their perception of market conditions they see relevant, and the speed at which transactions can be placed. Algorithmic models and the automated transactions they initiate have the significant advantage of speed.

Despite all the noise of training investment and the shifting definition of what constitutes intellectual property and its ownership and/or theft, the question remains: “How did they (Hangzhou Tencent) do it?” Modeling the relationships between any conceivable question and an answer is a problem of a vastly different magnitude than modeling financial market causes and effects. A statistically siftable haystack is not a good model, let alone a sensible way to burn through humungous quantities of energy.

Consider the well-known statement I♥️NY. It consists of the three symbols: the “I” that is about as simple as a symbol can get, the heart that is instantly comprehended as “love” and “NY” that, although we understand it as an abbreviation, is instantly comprehended as a symbol referring to the city of New York. The correct name for these three symbols is “logogram” – which is “a sign or character representing a word or phrase, such as those used in shorthand and some writing systems” . Logograms are understood quickly. The Chinese writing system consists of nothing but logograms. It’s fiendishly difficult to learn and write and, for centuries, the ability to read and write it was an indicator of educational level. The writing system was a handicap to written communication and learning dependent upon it.

Another example. The seemingly pidgin-English sentence “Love you long time” is how Chinese is constructed. It’s a straight word-for-logogram transliteration of the statement 好久爱你 – long time love you. This lack of prepositions is why the Chinese language, in movies at least, often sounds over-direct, abrupt, almost rude but it’s just how the language is structured. 好久爱你 is four logograms four words, all of which are relatively unambiguous. Back in the early days of what was then called “machine translation”, pre-editors would edit source languages to remove homonyms and metaphors and other constructions that couldn’t be directly converted. After the machine “translation” (conversion?) bit, post-editors would then reassemble it into a form recognizable in the destination language. Opponents of machine translation would be smugly amused when “The spirit is willing but the flesh is weak” translated into Russian as The vodka is good but the meat is rotten. Even today’s DeepL translates it into Chinese as 圣灵有心,肉体无力 which is The Spirit has a heart, the flesh has no power. But suppose, to begin with, we input the same thing in Chinese as 心有余而力不足, we get heart/spirit has will but strength not enough which is accurate and understandable, even if it’s not English as a native would say it. This, incidentally, is where the joy in speaking to non-native speakers lies. They make language strange for us again, and compel us to listen to what they are really saying. That aside, we have a situation where English into Chinese leaves ambiguities intact and vulnerable to mistranslation, but Chinese into English doesn’t because the logograms leave little room for ambiguity. Basically, the sentence is already pre-edited for meaning before it is input. Of course, spoken Chinese has an enormous number of synonyms as well as the added complications of tones (four different ones for Mandarin and seven for Cantonese) but this is only a problem for voice input and translation. Chinese people seem to have a lot of fun with their language, freely using synonyms and ambiguities in (as far as my experience goes) classroom banter. Me, I’m a long way from using Chinese to voice-control my television, or less usefully, my refrigerator, washing machine or rice cooker. Although spoken Chinese has huge potential for misunderstanding, the written language has extremely little. Poetry, as ever, is another matter as it often deals with metaphor and allusion. “The moon was a ghostly galleon, tossed upon cloudy seas …” etc.

Any performance advantage DeepSeek might have might have little to do with training data, CN¥ thrown at it, or even algorithms. It makes sense that if a question is asked more clearly and unambiguously, then it’s more likely that an answer will be more quickly forthcoming, even if that answer is statistically arrived at in the conventional way. DeepSeek’s performance just may be the simple consequence of written Chinese being an efficient conveyor of information now that its previous handicap of writing it by hand has been taken out of the equation. This is one possible processing advantage the Chinese language may have and this advantage may be amplified by its absence of articles as the English language does, or gendered articles such as German has, verb conjugations such as any of the Romance languages, or the case, verb and plural complexities of Arabic or Russian. I heard that certain African languages basically conjugate everything with everything else, and there’s some Australian aboriginal languages where sentences are conjugated according to the direction the speaker is facing with respect to the object of the sentence. Just as the Inuit might have sixteen or however many different words for snow, languages evolve to communicate what needs to be communicated. Now, when we feel we need to convey our meaning accurately to a CPU interface, the Chinese language just might be having a moment. Of course, poetry still gets written in Chinese as it does in English but AI assistants aren’t about poetry.

I can’t log into LinkedIn at the moment, even with a VPN. It seems to have gone dark. I’m not the only person in China with this problem this weekend. It may be temporary, or may not be. I don’t particularly care. For me, LinkedIn, is a legacy application, like Skype that I learn is soon to be discontinued. So if you don’t see anything new from me on LinkedIn, it doesn’t mean I’m not continuing to post on Wordpess.com as usual. Graham.

• • •

Comments

  • Thanks. I got the books by Atelier Bow Wow that’s I’m interested in. Like MUJI style, the house can be economic for the poor. I will have a study with them. Ps: I just right now you are a famous professor. ha-ha nice to meet you.

  • says:

    You’re welcome. What kind of Japanese house are you most interested in learning about? There’s two books by Atelier Bow Wow that have sectional perspective drawings indicating construction details. These are nice to have anyway.

  • Mr Graham Mckay, Thank your promote reply. I just right now know the merit of Japanese architecture style and lovin’ it very soon. Regarding MUJI HOUSE’s drawings, I realize they are trade secret today so that it’s difficult to get them. Now I change my mind. First, I try to describe the house as you said, though I think it will spend me lots of time. Second, I searched the book regarding the Japanese style house in the Z-library, found there are so many books either. but I don’t know which is suitable to me? I know you read lots of book, to save time, could you recommend one book where there is a complete construction drawings case about Japanese house. so I can study and practise my revit software skills effectively. Thanks a lot.

  • says:

    Hell Jade! everything on the blog is from the MUJI website which is very good. It does keep changing over the years but whatever they have shouldn’t be a problem for describing the houses, especially if i9ts for a general audience. Actual construction drawings are probably something they wouldn’t make public. Good luck with it!

  • says:

    Hi, Graham, I have seen your post about “MUJI HOUSE” which is earlier in 2008. I wonder whether you have their drawings and if it is convenient to share with me. Currently I am preparing my dissertation about architecture and have a big interesting on “MUJI HOUSE”. Your attention to my request would be highly appreciate. Jade Wang

  • says:

    Very insightful and level-headed commentary on language and intelligence. Thank you.

Leave a Reply