Making Sensei
Teaching My Phone to Play Mahjong
Written by Vu NguyenI started playing mahjong to meet people. Then I trained an AI model to play it on my phone. I still haven’t beaten it.
Which tile would you throw away to get ahead?
Mahjong tiles come in suits, and one suit is dots. The hand below holds four dots in a row: . Three in a row make a set. Throw away the 6 and you keep . Throw away the 3 and you keep .
Either way the hand is one tile from winning: a red dragon or a 1 of dots . What changes is what you give away. The tile you throw can help another player finish their hand. You can see what they’ve thrown away and the sets they’ve shown. You can’t see what they’re holding.
Which tile would you throw?
Player 1 across the table
Thrown away
Has laid down 7-8-9 of dots, and hasn’t thrown away a single dots tile all round.
Player 2 on its left
Thrown away
Player 4 on its right
Thrown away
37tiles left in the wall
The model player 3, its turn
Thrown away
The model’s hand is one tile from winning. Throw the 3 or the 6 and it still waits for a 1 of dots or a red dragon. Now look across the table: player 1 is collecting dots. Either tile could be the one they want.
Why the 6? There are four copies of every tile, and the two choices differ in how many copies were already showing. Two 6s were: the one in its hand, and the one in the it had laid down. Only one 3 was showing, its own. The fewer copies left out of sight, the less likely another player is holding a pair of them, ready to take yours.
To check that the model was doing this, I put one more 3 on the table, in each discard pile in turn. In three of the four, it switched and threw the 3 instead. It’s counting tiles, or something close to it.
It was a good reason, and it still went wrong. The player in seat 1 was holding the other two 6s and took it. On their next turn, they threw a red dragon, and the model won with it anyway.
This was the part I wanted to build: an AI that could make those choices. I could write the rules. I wasn’t good enough to write its strategy.
How do you teach a model a long game when you can’t tell which moves helped?
I couldn’t give it the right answers, so I gave it games. It plays millions of rounds against copies of itself and a few bots I wrote. Nobody marks each discard right or wrong. The only feedback comes at the end of a round: how many points it won or lost.
Points are how mahjong keeps score. When someone wins, they collect points: all of them from the player who threw the winning tile, or a share from everyone if they drew it themselves. Harder hands are worth more. The cheapest win pays 8.
The model never picks a tile outright. It gives every move it’s allowed to make a probability. In the hand above, it put 48% on the 6 and 41% on the 3. In practice games it picks by rolling dice weighted by those numbers, so it keeps trying things.
Learning means adjusting those probabilities. When a round goes better than the model expected, every move it made becomes a little more likely. When it goes worse, a little less. That’s reinforcement learning.
1. Play
Player 1 across the table
Thrown away
Player 2 on its left
Thrown away
Player 4 on its right
Thrown away
37tiles left in the wall
The model player 3, its turn
Thrown away
Every turn, the model gives each legal move a probability and one gets played. It sees its own hand and the table. Nothing else.
In this round the model won 24 points, all paid by player 1, who threw the red dragon. That doesn’t mean every move it made was good. Another player might rescue a bad decision; an unlucky draw might ruin a sensible one. One round is a noisy clue. Over millions of rounds the luck averages out, and the moves that keep showing up in good rounds are the ones that grow.
The method I used is called PPO, short for proximal policy optimization. The policy is the model’s set of probabilities. Proximal means nearby: each update stops pushing once a probability has moved about 2% from where it started, so one lucky batch can’t undo what the model already knows.
All of this happens during training. The phone gets the finished model. It doesn’t keep learning while you play.
Does more practice help?
Beating me was a low bar. I needed a comparison I could repeat.
As it trains, I save copies and name each one for how many decisions it has made. Every discard, claim or pass counts as one decision. So “400M” is the version that had made 400 million of them, and “850M” had made 850 million. They are the same model at different points in its practice. To compare two versions, I play them against the same opponents on the same deals, rotating seats, and count the points each takes home.
The 850M version earned about one point per round more than 400M. That’s an average over wins, losses and draws, not a point in every round. Both are the same size. More decisions means more practice, not a bigger model.
I almost stopped too soon.
At 711 million decisions the results looked flat, and I stopped training about a day before the run was due to finish.
Then I checked the test. It played 2,304 rounds per comparison, which can only tell two versions apart if they differ by more than about 0.86 points a round. The gains I was looking for were about half that. I’d read “no clear improvement” as “no improvement.”
| Checkpoint | 2,304-round test | 9,216-round test | 95% range, rebuilt test |
|---|---|---|---|
| 400M | 0.000 | 0.000 | exact by symmetry |
| 550M | −0.121 | +0.080 | [−0.348, +0.515] |
| 700M | +0.158 | +0.475 | [+0.029, +0.918] |
| 850M | not measured | +0.989 | [+0.546, +1.435] |
| 1B | not measured | +1.406 | [+0.951, +1.863] |
I rebuilt the test with four times as many rounds and resumed the run. The biggest gain of the whole run came in the stretch I had cut off. The 850M version went into the app in August.
I kept training, and this time it did stop: from 1 billion to 1.3 billion decisions, my test, now big enough to see a gain, found none.
The problem turned out to be flowers: bonus tiles like these , dealt at random and set aside. You don’t choose them, but under the rules I trained with, they change what a win pays. Here’s the winning hand from the top of this page, scored three ways.
Same hand. Different flowers.
Scored by the rules the model trained with: Hong Kong rules, where a win needs at least 3 faan and every extra faan pays more. The model is player 3, sitting West. Player 1 dealt, it’s the East round, and player 1 threw the red dragon that won it.
The tiles score 5 faan every time: one suit plus honour tiles 3; a set of red dragons 1; a set of its own wind, West, 1. Only the flowers change.
| Flowers dealt | Faan | Pays |
|---|---|---|
as dealt | 5+0: not its own | 24 |
—no flowers | 6+1: no flowers | 32 |
its own flower and season | 7+2: its own flower, its own season | 48 |
That swing has nothing to do with how the model played, but it still lands on the moves it made. I measured it: 37% of the variation in what winning hands paid came from flowers. The model was learning from a lot of luck.
The Hong Kong Mahjong Association’s rules don’t score flowers. So I went back to the 1-billion-decision model and trained it for another 300 million decisions with that scoring. Nothing else changed. It beat the run that kept the old scoring by half a point per round, even with both scored the old way, so it replaced 850M in the app.
What if the AI could see everyone’s tiles?
I found this idea in the paper behind Suphx, Microsoft’s AI for Japanese mahjong. It trained its model with every player’s tiles visible, then slowly took them away, so the model would learn what to look for. The paper calls this oracle guiding. I tried it. For 40 million decisions my model could see every other player’s hand. It won a little more often, and it threw other players their winning tile just as often as before. It could see the answer and never learned to use it.
So I asked a narrower question: how much is seeing everything worth? I let the model cheat while it plays: before each move, a search looks at every other hand and the order of the tiles still in the wall, and plays each option out to the end of the round.
Against a copy of itself, the cheating 400M won by about 8 points a round, roughly eight times what the model gained between 400M and 850M. But that’s the easy case. To play a move out to the end of the round, the search has to guess what the other players will do, and it assumes they’ll play like the model. Against copies of the model, that assumption is correct. Against anyone else, it isn’t.
So I put every version I’ve shipped against the same five bots I wrote, on the same 9,216 deals per bot, once playing fair and once cheating. v9.2 is the version before 400M, and h1 is the one I retrained without flower scoring, the model in the app. The chart scores them by HKMA rules, the ones h1 finished training on.
Plays fairCheats: sees every hand and the wallShort line above or below a dot: its 95% range
| Version | Rules | Plays fair | Cheats | Cheating adds | 95% range of the gain |
|---|---|---|---|---|---|
| v9.2 | HKMA | 13.13 | 15.34 | +2.21 | [+1.99, +2.44] |
| 400M | HKMA | 15.03 | 17.06 | +2.03 | [+1.80, +2.27] |
| 850M | HKMA | 15.94 | 18.02 | +2.08 | [+1.83, +2.31] |
| h1 | HKMA | 16.96 | 18.25 | +1.28 | [+1.03, +1.52] |
Cheating adds 1.3 to 2.2 points a round against the bots, much less than against itself, because the bots don’t play like the model. The part I keep coming back to is the dashed line. Playing fair, the model in the app scores 17.0 a round. 400M, cheating, scores 17.1. In other words, all the training after 400M made the model about as good as seeing everyone’s tiles made the old one.
A fair player can’t see those tiles, but can guess them from what’s on the table, so I also tried a search that does that. It gained nothing reliable. Most of what the cheat gains only appears when it knows the other hands and the wall at the same time, and that isn’t something you can guess your way to.
An AI that fits in my pocket. Literally.
The model in the iOS app is 8.9 MB, about the size of two songs, with 2.3 million parameters. It runs entirely on the phone, so you can play without internet. It plays the other three seats at the table, and it powers the hint button.
It’s also fast. A decision takes 0.4 milliseconds on a single processor core. When I had Claude Opus 5 play against it, each move took 36 seconds and cost about 9 cents. The app runs smoothly on phones as old as the iPhone 15, from 2023. The bots pause before each move on purpose, so you can follow the game.
At a mahjong event, I finally put it in front of someone who had played for more than ten years. He lost twice, then said it made the right moves.
Two games aren’t a benchmark. But I’d spent so long looking at training logs that hearing an experienced player talk about its actual moves meant a lot. I want more of that feedback.
Play a round with it.
The 850M model, in your browser. No download or account.
Evidence record
Measurements, methods, and limits
This is the detail behind the story: what I measured in August and September 2026, how, and where each result stops. A range is a 95% confidence interval: the true difference is very likely inside it. If a range includes zero, the test couldn’t tell the two versions apart. The app runs h1 and the browser runs 850M. The recorded games on this page were played by 400M.
How it trains, and the curve I ignoredAugust 20, 2026
Training signal — not a strength score
| Choice range | Average payment | Win rate | Deal-in rate |
|---|---|---|---|
| 96.1M–149.4M | 2.9236 | 27.20% | 14.59% |
| 149.4M–202.7M | 3.2425 | 27.97% | 14.40% |
| 202.7M–256.0M | 3.3322 | 27.85% | 14.00% |
| 256.0M–309.3M | 3.4259 | 27.93% | 13.88% |
| 309.4M–362.7M | 3.5427 | 27.79% | 13.76% |
| 362.7M–416.0M | 3.6139 | 27.67% | 13.82% |
| 416.0M–469.3M | 3.6367 | 27.58% | 13.74% |
| 469.3M–522.6M | 3.6722 | 27.47% | 13.57% |
| 522.6M–575.9M | 3.7362 | 27.22% | 13.38% |
| 575.9M–629.2M | 3.7172 | 27.21% | 13.46% |
| 629.3M–682.6M | 3.7535 | 27.14% | 13.49% |
| 682.6M–735.8M | 3.8655 | 27.13% | 13.30% |
The model has 2,307,140 parameters. On each turn it chooses from 127 possible moves: every tile it could throw, every call it could make, and declaring a win. The game removes the moves that aren’t allowed before it chooses, so it can never make an illegal one. Every version in this article is the same network, the same size and shape (3 residual blocks, width 96). They differ only in how much they’ve played.
I train it with PPO, the method described in the story. The only reward it ever gets is the points paid at the end of a round. Nothing scores a single discard.
Those points count in full for every move in the round, the first discard as much as the last. A lot of reinforcement learning shrinks rewards that arrive later, which is called discounting. That’s needed when a task never ends. A mahjong round ends after about 19.4 decisions, then points change hands, and those points are all that matter. Discounting would teach the model to prefer a small win now over a bigger win four turns later, and the game doesn’t pay that way. AlphaZero, which learned chess and Go by playing itself, scores its games the same way. There is a cost. With nothing to reward along the way, I tried rewarding good defence sixteen different ways, and none of them made a measurable difference.
The chart comes from the training log: 96,096,042 to 735,775,142 decisions, over 38,721 logged iterations and 32,826,560 player episodes. The log ends before the run did, and I haven’t drawn the line past it. The log also records how much the score jumped around between iterations. I left that off: a shaded band would look like uncertainty about strength, when it’s only noise. And 45% of its practice opponents change as it learns, so a rising line can also mean the opponents got easier to beat. This curve can’t show strength at all.
One of the practice bots was broken for the whole project. safety_aware was meant to avoid dangerous tiles, but a flipped sign made it throw the dangerous ones and keep the safe ones. Tested on its own it won 1.9% of rounds and averaged −2.17 points, the weakest player I measured, and it filled 10% of training seats from v8 on. I left it in, so older runs can still be reproduced against the opponents they trained with. Tests use a working defensive bot instead.
The 400M promotionAugust 10–17, 2026
Against the version before it, 400M scored +2.433 points a round, range [+1.989, +2.876], over 40,960 rounds where both played the same deals. Against four opponents it had never trained with, it scored +2.476, range [+1.510, +3.466]. That’s the harder test, because a model can learn the habits of practice partners it sees every day. These numbers came from a different test than the one that later failed its audit.
It ran in the browser from 2026-08-10 until 850M replaced it on 2026-09-22, and it played every recorded game on this page. The strongest model goes to the app, and the browser gets the one before it.
What a point isAugust 2026
Faan is how the rules value a hand: whether it’s allowed to win, and how much it’s worth. Points are what that value turns into. They’re what changes hands at the end of a round, and every number in this article is points, averaged over thousands of simulated rounds. Under the rules I trained with, the cheapest winning hand off someone’s discard moves 8 points from them to the winner.
What the audit found, and what it didn’t touchAugust 19–21, 2026
After 11 days of trusting it, I went back and checked the test. It had two bugs.
The small one: restarting a job should have drawn a fresh random seed, and didn’t. Repeat runs I had counted as separate evidence were partly the same run again.
The expensive one was the margin of error. The test treated every round as independent. But rounds came in blocks that shared the same opponents and seating, so results inside a block moved together. The margin has to be computed over blocks, not rounds. Counting rounds made every range I published 1.6× to 10.8× narrower than it should have been.
In points: at 2,304 rounds, the test could only tell two versions apart if they differed by about 0.86 points a round. Every gap I was asking it about was roughly half that. It answered anyway, and it answered “no difference” every time, which is the only answer it could have given.
Here is the same question put to both tests: did the model improve between 700M and 850M? The old test reads +0.256 [−0.694, +1.243]. The rebuilt one, at 9,216 rounds, reads +0.514 [+0.074, +0.959]. The first range includes zero and settles nothing. The second clears it: the later model is better. Neither model changed between those readings. Only the test did.
The audit also found that a model made by averaging the 400M, 550M and 700M versions had never been tested at all.
It retracts one conclusion: that training past 400M had stopped helping. It does not reach the 400M promotion above or the oracle result below. I measured both on other tests, and both found gaps far too large to be noise.
The 850M promotionAugust 21, 2026In the browser
After 850,004,713 decisions over 51,363 iterations, 850M played 400M over 9,216 rounds on the rebuilt test, against its six opponents (the five bots from the cheating comparison, plus v9.2, the version before 400M), and gained +0.989 points a round, range [+0.546, +1.435]. If I allow for a different set of opponents, the range widens to [+0.308, +1.694], which still clears zero. The previous candidate failed this step.
Then three checks that the gain didn’t come from the particular opponents I picked. Head to head, with only the two versions at the table, it gained +1.117 ± 0.836. One opponent at a time, it came out ahead against 6 of the 6. And leaving any one opponent out of the average keeps the gain between +0.876 and +1.175, so no single opponent carries the result.
Last, the dealer. Training always puts the dealer in the same seat, so a model could learn something that only works from there. This check moves the dealer and re-runs the comparison on a fresh seed, over 40,960 rounds: +1.503 [+1.075, +1.926]. Even the worst dealer seat gains at least +0.511, and the four seats barely differ from each other (χ² 0.56 on 3 degrees of freedom).
I wrote all four checks down before I ran them. Otherwise I could have tried everything and kept whichever comparison looked best.
| Candidate | Points per round | 95% interval | Outcome |
|---|---|---|---|
| 700M | +0.475 | [+0.029, +0.918] | Not promoted |
| Weight average | +0.491 | [+0.063, +0.917] | Not promoted |
| 850M | +0.989 | [+0.546, +1.435] | Promoted |
The two that didn’t make it are why those checks are required. 700M gained +0.475 [+0.029, +0.918] against the same six, which looks like a result until you allow for different opponents, and the range opens to [−0.342, +1.310], including zero. The averaged model gained +0.491 [+0.063, +0.917] against the six, then nothing head to head: +0.593 ± 0.809. Each passed one check and failed the rest. On the chart above, the two look the same.
I stopped this run on 2026-08-13 at 711,154,042 decisions, because the test said it had stopped improving, and restarted it on 2026-08-19. If the restart had quietly started from scratch, the whole line would mean nothing, so it had to prove it was the same run. The stopped run had logged four iterations (42961–42964) after its last save, and the restarted one had to reproduce all four exactly before anything after them counted. It passed 850,004,713 decisions on 2026-08-21, 25.6 hours later. That version ran in the app from 2026-08-21 until h1 replaced it, and in the browser since 2026-09-22. The run kept going to 1.3B decisions; the next section has how that ended.
On identical deals, the average payment across the four versions went 13.375 → 13.456 → 13.851 → 14.365 at 400M, 550M, 700M and 850M. The last step, +0.514, is the biggest, and it falls inside the stretch I had cut off.
I’m not claiming this is the best model the game allows. It was the last version to pass every check at the time, and the curve hadn’t flattened yet. It flattened after a billion. The gain is real but modest: 450 million more decisions bought about one point a round. Of the measurements behind the old “nothing more to gain” argument, cheating against the bots has since been run on every version (last section). Cheating against a copy of itself and the fair search were only run on 400M.
The plateau, the flowers, and h1August 24 – September 15, 2026In the app
Before training from 1B to 1.3B decisions, I wrote down what would count as progress: a gain over 1B whose range stayed above zero. I expected about +0.74. Over 18,432 paired rounds it read −0.067 [−0.400, +0.265]. The learning rate stayed the same the whole run, so this wasn’t a schedule running out. More training with the same recipe wasn’t going to help.
Next, the flowers. In 10,000 rounds of 850M playing itself, flowers explained 37.4% of the variation in what winning hands paid, and 53.3% of wins scored flower faan. When I hid every flower from the model, its top choice changed in only 4.2% of 12,000 positions. So flowers add a lot of luck to the score and almost nothing to the decisions, and a model trained without them gives up little. I set both thresholds before running either test.
h1 starts from the 1B version and trains for 300,016,656 more decisions with HKMA Classical scoring, which ignores flowers. Nothing else changed. The fair comparison is the 1B → 1.3B stretch: same start, same number of decisions, old scoring. Scored with flowers, the way the app plays, h1 beat it by +0.512 [+0.170, +0.852], and beat 850M by +0.693 [+0.349, +1.027], each over 18,432 paired rounds. The dealer check, on a fresh seed over 40,960 rounds, passed all eight of its conditions: no dealer seat or round wind doing worse than −2.5 points, no illegal moves or scoring failures, no dropped rounds, and no gain that vanishes against opponents it never trained with. Overall h1 gained +2.454 [+2.005, +2.907] against 400M.
This is one training run, with one random seed. The gain beat the bar I set, +0.50, by only 0.012, so all I can say is that the true gain is somewhere between 0.17 and 0.85. Two more stretches of 40 million decisions each read −0.012 at 340M and −0.234 at 380M against h1, so the gain was a one-time step, not a new climb. The flower example in the story is one recorded hand, rescored; the 37% is the measurement.
Cheating, a fair search, and the language modelsAugust – September 2026
Outside calibration · not a ranking
| Opponent | Games | Points from opponent side | 95% interval | Calibration |
|---|---|---|---|---|
| Dedicated adversary | 4,608 | +1.19 | [+0.96, +1.41] | Calibrated ladder |
| Previous champion | 4,608 | −3.15 | [−4.24, −2.06] | Calibrated ladder |
| wait_aware | 4,608 | −11.90 | [−12.85, −10.94] | Calibrated ladder |
| faan_aware | 4,608 | −13.22 | [−14.24, −12.20] | Calibrated ladder |
| shallow_search | 4,608 | −17.11 | [−18.06, −16.15] | Calibrated ladder |
| Claude Opus 5 | 96 | −0.08 | [−5.23, +5.06] | Outside calibration |
| GPT-5.6 | 96 | −0.92 | [−5.62, +3.79] | Outside calibration |
Impossible information
| Technique | Paired rounds | Settlement points per round | 95% interval | Legal at play time |
|---|---|---|---|---|
| Clairvoyant oracle | 4,608 | +7.93 | [+7.19, +8.66] | No |
| Legal belief search | 4,608 | −0.382 | [−0.920, +0.157] | Yes |
The first table gives each opponent’s points against 400M, so a negative number means my model won. One opponent beats it: the dedicated adversary, a separate model I trained to beat the frozen 400M and nothing else. It wins by about a point a round, but it is worse than 400M against every other opponent I tried, giving back about as much as it takes. That’s a weakness in 400M, not a stronger player, and I haven’t tested whether later versions share it.
For oracle guiding, following Suphx, I widened the model’s input with the other players’ hidden tiles, trained with them visible for 40 million decisions, and then played the same model on the same deals with the hidden tiles switched on and off. With them on, it won 27.9% of rounds instead of 27.3%, and threw someone their winning tile in 12.36% instead of 12.39%. It used what it saw to attack a little, and to defend not at all. Suphx’s next step fades the hidden tiles out; I didn’t run it, because there was no habit to keep.
The cheating search sees the other three hands and the order of the wall. This test used 400M, not the newer app model, playing a copy of itself. It won 2,170 rounds against 1,722 for the normal model, and the rounds where it threw someone their winning tile fell from 2,241 to 690. That’s what this particular search gains from seeing hidden tiles. It isn’t the best possible score, or what a fair player could gain. When I split the gain by what the search could see, 61.6% of it only appeared when it knew the hands and the wall together.
The comparison in the story is a separate test. Each version I’ve shipped played the same five bots on the same 9,216 deals per bot, once fair and once cheating, and every deal was played twice more, once under each set of rules: the training rules, which score flowers, and HKMA’s, which don’t. The ranges count blocks of six games that share a shuffle, as the audit above requires. Cheating is worth less here, 1.3 to 2.6 points, because the search assumes everyone plays like the model, and the bots don’t. Each version beats the one before it by a range that clears zero under both scorings; 850M over 400M, for example, is +0.96 [+0.46, +1.45] with flowers. h1 playing fair minus 400M cheating is −0.10 [−0.46, +0.28] under HKMA rules, and −0.74 [−1.24, −0.22] with flowers. At half as many deals, that last range was wide enough to include zero. The full test showed the gap.
Plays fairCheats: sees every hand and the wallShort line above or below a dot: its 95% range
| Version | Rules | Plays fair | Cheats | Cheating adds | 95% range of the gain |
|---|---|---|---|---|---|
| v9.2 | flowers scored | 13.26 | 15.87 | +2.61 | [+2.32, +2.89] |
| 400M | flowers scored | 15.46 | 18.06 | +2.60 | [+2.30, +2.91] |
| 850M | flowers scored | 16.42 | 18.80 | +2.37 | [+2.06, +2.69] |
| h1 | flowers scored | 17.32 | 19.09 | +1.76 | [+1.45, +2.09] |
I also tried the fair version of the same idea: guess what the others might be holding, play each guess out, and throw whatever does best on average. It only uses what a player at the table could know. It scored −0.382 [−0.920, +0.157]. The range includes zero, so it found nothing. That’s one attempt coming back empty, not proof the idea can’t work.
I also had two language models play against 400M, 96 rounds each. Claude Opus 5 scored −0.08 with interval [−5.23, +5.06]; GPT-5.6 scored −0.92 with interval [−5.62, +3.79]. Both ranges are far too wide to tell either one apart from my model or from each other. The Opus run cost $162.30 for 1,742 valid calls, about $0.093 a decision. This run trained at about 130 million decisions a day. Buying a single training day at that rate would cost about $12,090,000.
They were also slow. Claude Opus 5 averaged 36.2 seconds a decision over 3,176 calls, and GPT-5.6 25.0 seconds over 3,708, both including the trip over the network. My model takes 0.41 milliseconds on a single processor core. I haven’t measured it on a phone yet. Neither lab publishes a parameter count for these models, so I can’t compare sizes. I can only compare the time.