ภาษาphasa
Join the waitlist

Colophon · How Phasa is made

Method & Sources

An account of how Phasa is built and why — the tradition, the research, the sources of its data, and the things it deliberately does not claim.

On method

Most apps for learning Thai promise that it is easy. Phasa is built on the opposite belief: that learning a language is one of the longer, harder, more rewarding things a person can undertake, and that the tools should respect the size of it rather than pretend it away. There is no five-minute version of reading a novel in Thai. There is only the reading, done often enough and close enough to your level that it becomes possible.

This page is an account of how Phasa is built and why — the tradition it comes from, the research it rests on, the sources of its data, and the things it deliberately does not claim. We write it because a serious learner deserves to know what is under the hood, and because a tool that asks for years of your attention should be willing to show its reasoning.

There is no five-minute version of reading a novel in Thai.

The tradition

The idea beneath everything here is comprehensible input — the hypothesis, associated with the linguist Stephen Krashen, that we acquire a language mainly by understanding messages slightly beyond our current level, not by studying rules and drilling them.Krashen, S. D. (1982) Principles and Practice in Second Language Acquisition. Pergamon. — the input hypothesis. Acquisition, in this view, is something that happens to you when you understand things; conscious study plays a supporting role.

From that root grew the modern internet immersion movement. AJATT — “All Japanese All The Time,” Khatzumoto, mid-2000s — took the idea to its limit: surround yourself with the language every waking hour. The Mass Immersion Approach refined it into a practice — heavy input plus sentence mining, the habit of harvesting the specific sentences just past your reach and reviewing them — and was later rebuilt as Refold, a language-agnostic, codified version of the same method.We name the methods and channels rather than individuals; some attributions are inconsistently reported.

Phasa’s disagreement with this lineage is small but real: the method has always been sound, and it has always lacked a tool built for Thai. The good tools were built for other languages and fitted to Thai afterward. Phasa is built the other way round.

One finding from this literature shapes the product directly: combining video, audio, and text — captioned video — has been shown to aid second-language listening and vocabulary learning.Montero Perez, M., Van Den Noortgate, W., & Desmet, P. (2013). “Captioned video for L2 listening and vocabulary learning: a meta-analysis.” System, 41(3), 720–739. It is the reason the reader pairs every text with a recording, and the reason the browser extension turns subtitles into a living, tappable text rather than a passive caption.

Why Thai, specifically

Thai is not a language you can learn without learning to read it. It is written without spaces between words, it is romanised inconsistently by everyone who tries, and its five tones carry meaning that transliteration routinely flattens. The polyglot and applied linguist Stuart Jay Raj argues that the writing system is not arbitrary but logical — its consonants organised in large part by where in the mouth each sound is made — and that learning to read it is foundational, not optional.Raj, S. J. (2015). Cracking Thai Fundamentals: A Thai Operating System for Your Mind. We agree, which is why Phasa teaches the script properly, in both its contemporary hands — the looped form Thai children meet first, and the loopless form on signs and screens — rather than asking you to skip it.

The rooms, and the reasoning behind them

The reader holds the dictionary inside it. Tap a word and its entry opens in place; the same popup follows you onto YouTube and Netflix subtitles and any Thai on the open web. Reading with instant, in-context lookup is the shortest path from “I can decode letters” to “I can read,” which is the exact gap most learners get stuck in.

Sentence mining, then review. The sentences you save become flashcards. Review runs on FSRS, the Free Spaced Repetition Scheduler — an open-source algorithm that models each card by its stability, difficulty, and retrievability, trained on hundreds of millions of real reviews, and needs materially fewer reviews than the older SM-2 algorithm for the same retention.Ye, J., Su, J., & Cao, Y. (2022). “A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition Scheduling.” KDD ’22. — FSRS (creator Jarrett Ye). With SM-2, Woźniak, P. A. / SuperMemo (1987). There are no streaks to feed and no badges. The record is for you, not for the app to perform back at you.

The tone trainer. Thai’s five tones are the thing learners most fear and most tools handle worst. Phasa treats them as three connected machines. The first teaches the derivation: a syllable’s tone follows, by rule, from its initial consonant’s class, its live or dead shape, its vowel length, and any tone mark — and the trainer lets you operate that rule, assembling a syllable and reading off its tone, rather than memorising a grid. The second is a drill that adapts to the specific tone pairs you confuse — rising against low, falling against high — narrowing from a two-way to a five-way choice as your ear sharpens, and telling you which pair still trips you; it is built for the particular shape of the tone problem, not generic flashcard scheduling. The third is a periodic “check your ear” with no feedback until the end, so a single honest number can move over months. Throughout, you hear each tone in more than one voice — Jiji, Ploy, May, Koko, and Beer, with more to come — because the research on high-variability phonetic training shows that hearing a sound from multiple talkers is what lets you recognise it in a voice you have never met.Wang, Y., Spence, M. M., Jongman, A., & Sereno, J. A. (1999). “Training American listeners to perceive Mandarin tones.” JASA, 106(6). With the founding HVPT work of Logan, Lively & Pisoni (1991–93). A single unnamed voice, the norm among competitors, teaches you only that voice.

The shadowing room. Shadowing — saying a line back the instant you hear it, matching its melody — is among the oldest and most effective ways to train prosody, and Phasa measures it honestly. When you record yourself, the app extracts the pitch contour of your voice in the browser using the YIN algorithm, and compares its shape against the native speaker’s — register-normalised and time-warped, because you speak at a different pitch and pace, and it is the rise and fall that matters, not the exact frequency.de Cheveigné, A., & Kawahara, H. (2002). “YIN, a fundamental frequency estimator for speech and music.” JASA, 111(4), 1917–1930. Your melody is drawn against the native’s, warming where you match and cooling where you drift — never a red alarm. There is no cloud service and no black box: the pitch detection runs on your device, and we built it ourselves. And it is careful about what it claims: if your take is too quiet or too short to read, it says so rather than invent a score; it grades the whole line, not individual syllables, because we do not yet have the timing data to mark syllables honestly; and the pitch score is advisory — it never punishes you and never decides what you review next. That the approach helps is not merely intuition: real-time visual pitch feedback has been shown to improve second-language intonation,Hincks, R., & Edlund, J. (2009). “Promoting Increased Pitch Variation…” LL&T, 13(3). With Sakai & Moorman (2018), on perception→production transfer. and training the ear transfers to the voice.

How we build it honestly

The voices are real people, and we say who. Every recording of a Thai word or sentence — in the reader, the dictionary, the tone drills, the shadowing room — is a named native speaker, more than a dozen of them, from across the country, photographed as themselves. Not text-to-speech. The reason is not sentiment: synthesised audio still mishandles Thai tone and, more importantly, cannot join words into the natural melody of a real sentence, which is the whole thing a learner needs to hear.

There is exactly one synthesised sound anywhere in Phasa, and it is not pretending to be a voice: in the tone trainer, when a tone’s contour is drawn on screen, you can also hear its shape — a pure pitch glide that traces the curve. That glide is the melody made audible, deliberately not a word and not a person. And where a native recording of an example word is not ready yet, we tell you it is coming; we do not dress the synthesised shape up as a word.

The dictionary, and where it comes from. Phasa’s dictionary is built on open Thai lexical data — principally Wiktionary’s Thai entries, alongside other open sources — hand-curated and growing. We are precise about coverage: the common words are well-served; the long tail is not yet, and it grows every week. We do not claim depth we do not have.

How we check the tones. Because our pronunciation data comes from one family of sources, we do not simply trust it. We cross-checked every word that carries a pronunciation against an independent engine — tltk, the Thai Language Toolkit, a rule-based orthographic analyser that derives tone from the spelling itself, with no shared ancestry with our data.tltk — Wirote Aroonmanakun, Chulalongkorn University; rule-based Thai NLP. Across the common vocabulary the two agree on the tone of essentially every native word; the only disagreements are loanwords and borrowings, where spelling legitimately does not predict tone and the recorded pronunciation is the better authority. This does not make us infallible — no cross-check proves a shared error absent — but it means the specific failure a native speaker would catch, a wrong tone on an ordinary Thai word, has been actively looked for and not found.

Segmentation you can correct. Thai’s missing word-boundaries are the reason automatic tools stumble; the well-known ones split ตัดสินใจ into three pieces and break the reading. Phasa segments with a trained tokeniser and then lets you fix the split with a tap when it is wrong — because on a language with no spaces, the ability to correct the machine is not a nicety, it is the feature.

We do not dress the synthesised shape up as a word.

What we do not claim

  • We do not promise fluency in a number of months. Nobody can, and the people who say so are selling something.
  • We do not claim the dictionary is complete. It is roughly a sixth of the way to covering every entry with a pronunciation, and honest about it.
  • We do not present computed or borrowed pronunciation as native audio, or a synthesised pitch-shape as a spoken word.
  • We do not claim our tones are provably perfect — only that they are sourced, cross-checked against an independent engine, and reviewed by a native speaker where it matters most.
  • We do not grade your pronunciation with a number that pretends to more certainty than the recording allows; the shadowing score is advisory and says when it cannot judge.

Sources & further reading

Krashen, S. D. (1982). Principles and Practice in Second Language Acquisition. Pergamon.

AJATT (Khatzumoto) → the Mass Immersion Approach → Refold. The modern immersion method and sentence mining.

Raj, S. J. (2015). Cracking Thai Fundamentals: A Thai Operating System for Your Mind.

Montero Perez, M., Van Den Noortgate, W., & Desmet, P. (2013). “Captioned video for L2 listening and vocabulary learning: a meta-analysis.” System, 41(3).

Ye, J., Su, J., & Cao, Y. (2022). “A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition Scheduling.” KDD ’22. With SM-2, Woźniak, P. A. / SuperMemo (1987).

Wang, Y., Spence, M. M., Jongman, A., & Sereno, J. A. (1999). “Training American listeners to perceive Mandarin tones.” JASA, 106(6). With Logan, Lively & Pisoni (1991–93).

de Cheveigné, A., & Kawahara, H. (2002). “YIN, a fundamental frequency estimator for speech and music.” JASA, 111(4), 1917–1930.

Hincks, R., & Edlund, J. (2009). “Promoting Increased Pitch Variation in Oral Presentations with Transient Visual Feedback.” LL&T, 13(3). With Sakai, M., & Moorman, C. (2018).

Aroonmanakun, W. tltk (Thai Language Toolkit) — rule-based Thai NLP, used for our independent tone cross-check.

Wiktionary (Thai) and other open Thai lexical sources — the dictionary’s data.

ภาษา · phasaBuilt in อยุธยาOpen betaSet in IBM Plex Serif & Sarabun

Phasa is made by a small team, built in Thailand, in open beta. Write to us at hello@phasa.app — including, especially, to tell us where we are wrong; a dictionary improves by being corrected.

The Thai language, taken seriously.