Blog
Serious hobbies, rabbit holes, or just casual notes.
Article migration in progress
A Scientist Walks Into a Debate... What Exactly Makes Something Machine Learning?
2023-09-22
Having worked with people from both scientific research and tech, I’ve met quite a few ML engineers and data scientists who hold surprisingly strong views about what does and doesn't count as "machine learning", especially when applied to scientific research. One claim I hear all the time goes like this: > "Curve fitting is not machine learning because it doesn't make predictions." It sounds reasonable at first glance. But like many things in biology and complex systems, the boundary isn't always black and white. This prescriptive view of ML seems to stem from a few common misconceptions, or maybe just from not building a deep intuition for the topic when first starting out. To share a scientist's perspective on machine learning, I’m going to use an imaginary chat between a scientist and an ML engineer. This isn't an attempt to define machine learning. Instead, I want to use curve fitting as a comparison to help build a more intuitive way of thinking about ML and AI. Whether you're a beginner looking for a high-level intuition or an ML practitioner looking for a fresh perspective, I hope this gives you some satisfying food for thought! ## The Conversation **Engineer:** Curve fitting isn't machine learning. **Scientist:** Why not? **Engineer:** Curve fitting means fitting data to a predefined equation and estimating its coefficients. Machine learning means training a model on data to make predictions. **Scientist:** Sure. Suppose I fit $y=ax+b$ to data by minimizing the "errors": $$ \sum_i (y_i-(ax_i+b))^2 $$ What is that? **Engineer:** Linear regression. **Scientist:** And linear regression is machine learning? **Engineer:** Of course. **Scientist:** So at least some curve fitting is ML. ### "The equation is known in advance" **Engineer:** Fine. But that's just a special case. In curve fitting, you already know the equation. In ML, the model learns the function. **Scientist:** Does it? With a neural network, don't you choose the architecture first? **Engineer:** Yes. **Scientist:** And doesn't the architecture determine the structure of the function? **Engineer:** Essentially. **Scientist:** So you specify a model family first, then learn its parameters. That's also what I do when I choose, for example, $y=ae^{-bx}$, and estimate $a$ and $b$. **Engineer:** But a neural network is different. It is much more flexible. **Scientist:** Certainly. But that's a difference in *model capacity*, not necessarily in the underlying process: you first choose the model/function form, then learn the parameters from your data. ### "But neural networks *learn* their parameters" **Engineer:** Neural networks learn their parameters. That's what makes them machine "learning". **Scientist:** What does "learn" mean? **Engineer:** The training algorithm finds parameter values that optimize a loss. **Scientist:** Right. Let's say I have this optimization process: $$ \theta^*=\arg\min_\theta L(\theta) $$ If I fit $y=ae^{-bx}$ by finding $a$ and $b$ that minimize the error function, am I not also estimating parameters from data? **Engineer:** Yes, but I would just call it "parameter estimation". **Scientist:** That's perfectly reasonable, though that's just terminology and context, rather than a fundamental mathematical difference. ### "Scientists care About the parameters" **Engineer:** Wait... What about how the parameters are interpreted? Scientists often care about the parameters themselves. For example, when you fit your synaptic currents or whatever brain signals to the decay equation, $$ N(t)=N_0e^{-kt}, $$ You might actually want to know $k$ because it has scientific meaning. An ML engineer usually doesn't care what an individual parameter means, but whether the model *predicts* well. **Scientist:** That's an important distinction, but is it a distinction between curve fitting and ML, or between *scientific inference* and *predictive modelling*? **Engineer:** What's the difference? **Scientist:** Suppose I fit $y=a+bx+cx^2$ because it predicts future observations well. I don't care what the coefficients mean physically. Would you call that ML? **Engineer:** Probably. **Scientist:** And if I use logistic regression because I want to understand $\beta_1$ scientifically? **Engineer:** It can still be ML. **Scientist:** Then the interpretability of parameters isn't the boundary. ### "ML is about prediction" **Engineer:** Fine. The real distinction is prediction. ML is about *generalization to unseen data*. Curve fitting is about *explaining or describing the observed data*. **Scientist:** If I fit $y=ax+b$ on four observations and use it to predict an unseen data point, have I generalized beyond the fitting data? **Engineer:** Yes. But don't make me call every least-squares problem machine learning. **Scientist:** I'm not. I'm saying prediction alone doesn't separates them. ### "But ML is much more sophisticated" **Engineer:** ML is much more sophisticated. Neural networks have millions or billions of parameters. **Scientist:** Is there a minimum number of parameters before something becomes ML? **Engineer:** No. **Scientist:** Is a one-neuron neural network ML? **Engineer:** Yes... **Scientist:** So model complexity can't define the boundary either. ### "But curve fitting uses *equations*" **Engineer:** That's fair, but curve fitting explicitly uses *equations*. ML uses *models*. **Scientist:** What is a neural network mathematically? **Engineer:** A parameterized function. **Scientist:** Right. For example: $$ f(x;\theta)=W_2\sigma(W_1x+b_1)+b_2. $$ That's an equation. So the neural-network architecture is just like the functional form of a fitted equation. In other words, $$ \boxed{\text{architecture}+\text{parameters}} $$ is conceptually similar to: $$ \boxed{\text{functional form}+\text{parameters}} $$ **Engineer:** Makes sense, though the neural network can represent a much larger class of functions. **Scientist:** Absolutely. But again, that's model capacity, not a fundamental difference. ### Looking at a higher level **Scientist:** Let's try to think at a higher level. What makes something machine learning? **Engineer:** Learning useful patterns or a function from data. **Scientist:** And can fitting an equation involve learning a function from data? Can the resulting function make predictions? **Engineer:** Yes and yes. **Scientist:** Can its parameters be learned by optimization? Can its functional form be specified in advance? Can ML do the same? **Engineer:** Yes, yes, and yes... **Scientist:** Then perhaps the view "curve fitting is not ML" is a bit too strong **Engineer:** Fair. It's machine learning, just without an expen$ive cloud bill 😉
The Perception-Production Asymmetry of Cantonese Tone Mergers
2022-05-04
The phonological tone is an interesting linguistic phenomenon with characteristics distinguished from other phonological segments such as consonants and vowels. Tones have been shown to be autonomous to the segments they are associated with (i.e., the underlying representation and operation of tones are independent of the associated segments. See Goldsmith (1976)). This suggests that tone and other segments may have distinct structures of mental representation. While the emergence of tones in some languages is not completely understood, it has been shown that tonogenesis in some languages was likely induced by contacts with other phylogenetically unrelated tonal languages in neighboring areas (e.g. in some Southeast Asian languages; see Kirby and Brunelle (2007)), and that tonogenesis and tonoexodus (loss of tones) could happen in a relatively short period of time (e.g., Korean). Such an areal feature of tone and its relatively high fluidity compared to other segments raise the question of how the brain represents and processes tones and other segments differentially. The unique characteristics of tone discussed above make tonal languages promising targets for studying speech processing and gaining a more comprehensive understanding of how language works. Indeed, the study of tonal languages revolutionized the traditional approach to segment analysis, leading to the development of autosegmental phonology. Cantonese is a tonal language spoken in the Guangdong Province of China, Hong Kong, Macau, and many other Han Chinese communities around the world. Compared to other equally-resourced tonal languages, it is particularly well suited for the study of the cognitive processing of tones since most native Cantonese speakers have little metalinguistic awareness of Cantonese tones [[1](#footnotes)], and therefore, expectedly, less influenced by top-down processing in tone perception. It is widely accepted that standard Cantonese [[2](#footnotes)] has six canonical tones. However, it has been observed recently that some of the tones seem to be merged and are no longer distinguished among younger Cantonese speakers in their perception or production. In this paper, I will first summarize the current findings about tone merging in Cantonese. Using Cantonese tone merging as an example, I will also review some of the proposed mechanisms to shed light on an interesting linguistic phenomenon – the asymmetry between perception and production. Tone merging can happen in perception, production, or both language modalities. Different studies may use different wordings to describe their findings. *Continue reading [here](https://github.com/manhowong/manhowong/blob/main/cantonese_tone_mergers.pdf)...* #### Footnotes > [1]: Most native Cantonese speakers have difficulties in categorizing syllables based on lexical tones, regardless of their ability to distinguish tones. This is likely due to the lack of formal Cantonese education: Even in Hong Kong where Cantonese is the de facto official language, essentially no native Cantonese speakers have received formal education on the Cantonese language in school (except for students who study Cantonese as a linguistics subject in university). In China, including Cantonese-speaking regions, Mandarin Chinese is the only Chinese language taught in school. In Hong Kong, while Cantonese is the teaching medium of the Chinese language in most schools, students only study the official variety of written Chinese (i.e. written Chinese based on Mandarin syntax). As a result, most native Cantonese speakers do not possess the same level of linguistic knowledge about tones in their native language compared to other tonal language speakers who study their native language formally in school, such as Mandarin speakers (e.g. most Mandarin speakers can categorize tones in Mandarin without much effort). > [2]: Currently, two different dialects of Cantonese are both considered the de facto standard form of Cantonese: the Guangzhou dialect spoken in Guangzhou and the Hong Kong dialect spoken in Hong Kong. The two varieties are nearly identical phonologically, but the vocabulary of each of the varieties is influenced by Mandarin and English respectively. At the moment, Hong Kong is the only place where Cantonese is recognized as a working language and widely used in different settings such as education, literature, and courts, while Guangzhou Cantonese is mainly used between friends and family. This makes Hong Kong Cantonese a good study target for Cantonese as most of its speakers are able to express their ideas on different topics fully in Cantonese. ### References - **Goldsmith, J. A. (1976).** Autosegmental phonology (Doctoral dissertation, Massachusetts Institute of Technology). - **Kirby, J., and Brunelle, M. (2017).** Southeast Asian tone in areal perspective. *The Cambridge handbook of areal linguistics*, 703-731. DOI:10.1017/9781107279872.027
A (Very) Brief Comparison of Bengali and English Phonological Systems
2021-12-01
Bengali (বাংলা ভাষা, *Bangla-bhasha*) is an Indo-European language widely spoken in Bangladesh and neighboring areas in India. A language with a rich history of music and literature, Bengali is the mother language of many influential figures in South Asian arts and culture, such as Rabindranath Tagore and Kazi Nazrul Islam. The language also has a special place in the modern history of linguistic human rights: The Bengali Language Movement of the early 1950s, particularly the events of 21 February 1952 in East Bengal, advocated official recognition of Bengali alongside Urdu. The movement and the sacrifices of Bengali language activists were subsequently commemorated through UNESCO's establishment of International Mother Language Day (United Nations, 2021). Bengali can be considered a diglossic language. While the modern literary form, *Cholito-bhasha* (based on the Nadia dialect), is relatively consistent across Bengali communities, colloquial Bengali varies considerably in phonology and, according to Behrman et al. (2021), can be grouped into six major dialects. In this paper, I will describe the basic phonetics and phonology of the standard forms of Bengali (the Bengali dialect spoken in Dhaka, Bangladesh, and the West Bengali dialect spoken in the West Bengal region of India) and compare them with the phonological system of General American English. I will first analyze the consonant system, the vowel system, and the syllable structure. Bengali’s notable prosody patterns will be discussed briefly in the last section. *Continue reading [here](https://github.com/manhowong/manhowong/blob/main/Bengali_Man%20Ho%20Wong_2023.pdf) or on [LingBuzz](https://lingbuzz.net/lingbuzz/007140)...* ### References - Behrman, E., Santra, A., Sarkar, S., Roy, P., Yadav, R., Dutta, S., & Ghosal, A. (2021). Dialect Identification of the Bengali Language. In Tyagi A. K. (Ed.), *Data Science and Data Analytics* (1st ed., pp. 142–170). Chapman and Hall/CRC. - United Nations. (2021). *International Mother Language Day*. United Nations. Retrieved November 17, 2021, from https://www.un.org/en/observances/mother-language-day