Making Sensei

Making Sensei

Teaching My Phone to Play Mahjong

Written by Vu Nguyen

September 2026

I started playing mahjong to meet people. Then I trained an AI model to play it on my phone. I still haven’t beaten it.

Which tile would you throw away to get ahead?

Mahjong tiles come in suits, and one suit is dots. The hand below holds four dots in a row: . Three in a row make a set. Throw away the 6  and you keep . Throw away the 3  and you keep .

Either way the hand is one tile from winning: a red dragon  or a 1 of dots . What changes is what you give away. The tile you throw can help another player finish their hand. You can see what they’ve thrown away and the sets they’ve shown. You can’t see what they’re holding.

A real hand, move 65

Which tile would you throw?

Player 1 across the table

Thrown away

Has laid down 7-8-9 of dots, and hasn’t thrown away a single dots tile all round.

Player 2 on its left

Thrown away

Player 4 on its right

Thrown away

37tiles left in the wall

The model player 3, its turn

Thrown away

The model’s hand is one tile from winning. Throw the 3 or the 6 and it still waits for a 1 of dots or a red dragon. Now look across the table: player 1 is collecting dots. Either tile could be the one they want.

A real hand, played by an earlier version of my AI model. Reveal its choice.From a game I recorded after training, against Claude Opus 5 in two seats. The percentages are the model’s own output for this position.

Why the 6? There are four copies of every tile, and the two choices differ in how many copies were already showing. Two 6s were: the one in its hand, and the one in the it had laid down. Only one 3 was showing, its own. The fewer copies left out of sight, the less likely another player is holding a pair of them, ready to take yours.

To check that the model was doing this, I put one more 3 on the table, in each discard pile in turn. In three of the four, it switched and threw the 3 instead. It’s counting tiles, or something close to it.

It was a good reason, and it still went wrong. The player in seat 1 was holding the other two 6s and took it. On their next turn, they threw a red dragon, and the model won with it anyway.

This was the part I wanted to build: an AI that could make those choices. I could write the rules. I wasn’t good enough to write its strategy.

How do you teach a model a long game when you can’t tell which moves helped?

I couldn’t give it the right answers, so I gave it games. It plays millions of rounds against copies of itself and a few bots I wrote. Nobody marks each discard right or wrong. The only feedback comes at the end of a round: how many points it won or lost.

Points are how mahjong keeps score. When someone wins, they collect points: all of them from the player who threw the winning tile, or a share from everyone if they drew it themselves. Harder hands are worth more. The cheapest win pays 8.

The model never picks a tile outright. It gives every move it’s allowed to make a probability. In the hand above, it put 48% on the 6 and 41% on the 3. In practice games it picks by rolling dice weighted by those numbers, so it keeps trying things.

Learning means adjusting those probabilities. When a round goes better than the model expected, every move it made becomes a little more likely. When it goes worse, a little less. That’s reinforcement learning.

One round of practice

1. Play

Player 1 across the table

Thrown away

Player 2 on its left

Thrown away

Player 4 on its right

Thrown away

37tiles left in the wall

The model player 3, its turn

Thrown away

Every turn, the model gives each legal move a probability and one gets played. It sees its own hand and the table. Nothing else.

1 / 3
The same hand, followed to the end of the round.A recorded game, used to show the loop. Training rounds work the same way, millions of times over.

In this round the model won 24 points, all paid by player 1, who threw the red dragon. That doesn’t mean every move it made was good. Another player might rescue a bad decision; an unlucky draw might ruin a sensible one. One round is a noisy clue. Over millions of rounds the luck averages out, and the moves that keep showing up in good rounds are the ones that grow.

The method I used is called PPO, short for proximal policy optimization. The policy is the model’s set of probabilities. Proximal means nearby: each update stops pushing once a probability has moved about 2% from where it started, so one lucky batch can’t undo what the model already knows.

All of this happens during training. The phone gets the finished model. It doesn’t keep learning while you play.

Does more practice help?

Beating me was a low bar. I needed a comparison I could repeat.

As it trains, I save copies and name each one for how many decisions it has made. Every discard, claim or pass counts as one decision. So “400M” is the version that had made 400 million of them, and “850M” had made 850 million. They are the same model at different points in its practice. To compare two versions, I play them against the same opponents on the same deals, rotating seats, and count the points each takes home.

The 850M version earned about one point per round more than 400M. That’s an average over wins, losses and draws, not a point in every round. Both are the same size. More decisions means more practice, not a bigger model.

I almost stopped too soon.

At 711 million decisions the results looked flat, and I stopped training about a day before the run was due to finish.

Then I checked the test. It played 2,304 rounds per comparison, which can only tell two versions apart if they differ by more than about 0.86 points a round. The gains I was looking for were about half that. I’d read “no clear improvement” as “no improvement.”

Points gained per round against the 400-million-decision model, plotted across the run and scored by two tests. The 2,304-round test reads minus 0.121 at 550 million and plus 0.158 at 700 million, and its resolution of plus or minus 0.85 covers zero at every point, so these readings did not establish an improvement. The 9,216-round test reads plus 0.080, then plus 0.475, then plus 0.989 at 850 million, and its range clears zero from 700 million onward. The first test has no reading at 850 million, because the run was stopped at 711 million decisions on its say-so, before 850 million existed. The hollow 1B point is a later measurement, never shipped.Did more training help? The same run, scored by two tests.Points gained per round against the 400M modelThe test I had · 2,304 rounds · never cleared zeroThe test I rebuilt · 9,216 rounds · saw the climb-10+1the 400M modelI stopped the run here+0.16+0.99+1.41400M550M700M850M1Bdecisions of self-playThe first test has no reading at 850M. I had already stopped the run.Hollow: measured on this test, never shipped.Points gained per round against the 400-million-decision model, plotted across the run and scored by two tests. The 2,304-round test reads minus 0.121 at 550 million and plus 0.158 at 700 million, and its resolution of plus or minus 0.85 covers zero at every point, so these readings did not establish an improvement. The 9,216-round test reads plus 0.080, then plus 0.475, then plus 0.989 at 850 million, and its range clears zero from 700 million onward. The first test has no reading at 850 million, because the run was stopped at 711 million decisions on its say-so, before 850 million existed. The hollow 1B point is a later measurement, never shipped.Did more training help?The same run, scored by two tests.Points gained per round against the 400M modelThe test I had · 2,304 rounds · never cleared zeroThe test I rebuilt · 9,216 rounds · saw the climb-10+1the 400M modelI stopped here+0.16+0.99+1.41400M550M700M850M1Bdecisions of self-playThe first test has no reading at 850M.I had already stopped the run.Hollow: measured, never shipped.
Points gained per round against the 400M model, scored by both tests
Checkpoint2,304-round test9,216-round test95% range, rebuilt test
400M0.0000.000exact by symmetry
550M−0.121+0.080[−0.348, +0.515]
700M+0.158+0.475[+0.029, +0.918]
850Mnot measured+0.989[+0.546, +1.435]
1Bnot measured+1.406[+0.951, +1.863]
The first test couldn’t separate the versions. The rebuilt one shows the climb.Both tests compare each version with the 400M model. The first has no reading at 850M: I had already stopped the run.

I rebuilt the test with four times as many rounds and resumed the run. The biggest gain of the whole run came in the stretch I had cut off. The 850M version went into the app in August.

I kept training, and this time it did stop: from 1 billion to 1.3 billion decisions, my test, now big enough to see a gain, found none.

The problem turned out to be flowers: bonus tiles like these , dealt at random and set aside. You don’t choose them, but under the rules I trained with, they change what a win pays. Here’s the winning hand from the top of this page, scored three ways.

The winning hand from above

Same hand. Different flowers.

Scored by the rules the model trained with: Hong Kong rules, where a win needs at least 3 faan and every extra faan pays more. The model is player 3, sitting West. Player 1 dealt, it’s the East round, and player 1 threw the red dragon that won it.

The tiles score 5 faan every time: one suit plus honour tiles 3; a set of red dragons 1; a set of its own wind, West, 1. Only the flowers change.

Flowers dealtFaanPays
as dealt
5+0: not its own24
—no flowers
6+1: no flowers32
its own flower and season
7+2: its own flower, its own season48
Same tiles, same decisions. The flowers alone move the payout from 24 to 48.The recorded win, rescored in the simulator under the rules the model trained with.

That swing has nothing to do with how the model played, but it still lands on the moves it made. I measured it: 37% of the variation in what winning hands paid came from flowers. The model was learning from a lot of luck.

The Hong Kong Mahjong Association’s rules don’t score flowers. So I went back to the 1-billion-decision model and trained it for another 300 million decisions with that scoring. Nothing else changed. It beat the run that kept the old scoring by half a point per round, even with both scored the old way, so it replaced 850M in the app.

What if the AI could see everyone’s tiles?

I found this idea in the paper behind Suphx, Microsoft’s AI for Japanese mahjong. It trained its model with every player’s tiles visible, then slowly took them away, so the model would learn what to look for. The paper calls this oracle guiding. I tried it. For 40 million decisions my model could see every other player’s hand. It won a little more often, and it threw other players their winning tile just as often as before. It could see the answer and never learned to use it.

So I asked a narrower question: how much is seeing everything worth? I let the model cheat while it plays: before each move, a search looks at every other hand and the order of the tiles still in the wall, and plays each option out to the end of the round.

Against a copy of itself, the cheating 400M won by about 8 points a round, roughly eight times what the model gained between 400M and 850M. But that’s the easy case. To play a move out to the end of the round, the search has to guess what the other players will do, and it assumes they’ll play like the model. Against copies of the model, that assumption is correct. Against anyone else, it isn’t.

So I put every version I’ve shipped against the same five bots I wrote, on the same 9,216 deals per bot, once playing fair and once cheating. v9.2 is the version before 400M, and h1 is the one I retrained without flower scoring, the model in the app. The chart scores them by HKMA rules, the ones h1 finished training on.

Plays fairCheats: sees every hand and the wallShort line above or below a dot: its 95% range

HKMA rulesno flowers1214161820points per round against the botsh1, playing fairv9.2+2.2400M+2.0850M+2.1h1+1.3
Points per round against five scripted bots, 9,216 games per bot, identical deals for every version
VersionRulesPlays fairCheatsCheating adds95% range of the gain
v9.2HKMA13.1315.34+2.21[+1.99, +2.44]
400MHKMA15.0317.06+2.03[+1.80, +2.27]
850MHKMA15.9418.02+2.08[+1.83, +2.31]
h1HKMA16.9618.25+1.28[+1.03, +1.52]
Each version beats the bots by more than the last. Playing fair, the model in the app (h1) scores what 400M scores when it cheats.9,216 games against each of five bots per version, identical deals throughout, HKMA rules. Lines show 95% ranges. With flowers scored the order is the same; the evidence record has that chart.

Cheating adds 1.3 to 2.2 points a round against the bots, much less than against itself, because the bots don’t play like the model. The part I keep coming back to is the dashed line. Playing fair, the model in the app scores 17.0 a round. 400M, cheating, scores 17.1. In other words, all the training after 400M made the model about as good as seeing everyone’s tiles made the old one.

A fair player can’t see those tiles, but can guess them from what’s on the table, so I also tried a search that does that. It gained nothing reliable. Most of what the cheat gains only appears when it knows the other hands and the wall at the same time, and that isn’t something you can guess your way to.

An AI that fits in my pocket. Literally.

The model in the iOS app is 8.9 MB, about the size of two songs, with 2.3 million parameters. It runs entirely on the phone, so you can play without internet. It plays the other three seats at the table, and it powers the hint button.

It’s also fast. A decision takes 0.4 milliseconds on a single processor core. When I had Claude Opus 5 play against it, each move took 36 seconds and cost about 9 cents. The app runs smoothly on phones as old as the iPhone 15, from 2023. The bots pause before each move on purpose, so you can follow the game.

At a mahjong event, I finally put it in front of someone who had played for more than ten years. He lost twice, then said it made the right moves.

Two games aren’t a benchmark. But I’d spent so long looking at training logs that hearing an experienced player talk about its actual moves meant a lot. I want more of that feedback.

Play a round with it.

The 850M model, in your browser. No download or account.

Evidence record

Measurements, methods, and limits

This is the detail behind the story: what I measured in August and September 2026, how, and where each result stops. A range is a 95% confidence interval: the true difference is very likely inside it. If a range includes zero, the test couldn’t tell the two versions apart. The app runs h1 and the browser runs 850M. The recorded games on this page were played by 400M.

How it trains, and the curve I ignoredAugust 20, 2026

Training signal — not a strength score

Average training payment rose from 2.92 in the first 53-million-decision bin to 3.87 in the last exported bin, ending at 735.8 million decisions of a run that finished at 850 million. The line rises across the whole stretch in which the evaluation ladder was reporting that further training had stopped helping. Across the same snapshot, deal-ins fell while wins stayed roughly flat. This is a training-league health signal, not a measure of absolute playing strength. 3.03.54.0100M400M700Mdecisions of self-playfirst bin · 2.92400M · 3.61export ends · 3.87Fewer deal-ins: 14.6% → 13.3%Wins roughly flat: 27.2% → 27.1%
Equal-width decision bins from the live policy training snapshot
Choice rangeAverage paymentWin rateDeal-in rate
96.1M–149.4M2.923627.20%14.59%
149.4M–202.7M3.242527.97%14.40%
202.7M–256.0M3.332227.85%14.00%
256.0M–309.3M3.425927.93%13.88%
309.4M–362.7M3.542727.79%13.76%
362.7M–416.0M3.613927.67%13.82%
416.0M–469.3M3.636727.58%13.74%
469.3M–522.6M3.672227.47%13.57%
522.6M–575.9M3.736227.22%13.38%
575.9M–629.2M3.717227.21%13.46%
629.3M–682.6M3.753527.14%13.49%
682.6M–735.8M3.865527.13%13.30%
The training signal kept climbing through the whole stretch the test was calling flat.Training-log export captured 2026-08-20, ending at 735,775,142 decisions of a run that finished at 850,004,713.

The model has 2,307,140 parameters. On each turn it chooses from 127 possible moves: every tile it could throw, every call it could make, and declaring a win. The game removes the moves that aren’t allowed before it chooses, so it can never make an illegal one. Every version in this article is the same network, the same size and shape (3 residual blocks, width 96). They differ only in how much they’ve played.

I train it with PPO, the method described in the story. The only reward it ever gets is the points paid at the end of a round. Nothing scores a single discard.

Those points count in full for every move in the round, the first discard as much as the last. A lot of reinforcement learning shrinks rewards that arrive later, which is called discounting. That’s needed when a task never ends. A mahjong round ends after about 19.4 decisions, then points change hands, and those points are all that matter. Discounting would teach the model to prefer a small win now over a bigger win four turns later, and the game doesn’t pay that way. AlphaZero, which learned chess and Go by playing itself, scores its games the same way. There is a cost. With nothing to reward along the way, I tried rewarding good defence sixteen different ways, and none of them made a measurable difference.

The chart comes from the training log: 96,096,042 to 735,775,142 decisions, over 38,721 logged iterations and 32,826,560 player episodes. The log ends before the run did, and I haven’t drawn the line past it. The log also records how much the score jumped around between iterations. I left that off: a shaded band would look like uncertainty about strength, when it’s only noise. And 45% of its practice opponents change as it learns, so a rising line can also mean the opponents got easier to beat. This curve can’t show strength at all.

One of the practice bots was broken for the whole project. safety_aware was meant to avoid dangerous tiles, but a flipped sign made it throw the dangerous ones and keep the safe ones. Tested on its own it won 1.9% of rounds and averaged −2.17 points, the weakest player I measured, and it filled 10% of training seats from v8 on. I left it in, so older runs can still be reproduced against the opponents they trained with. Tests use a working defensive bot instead.

The 400M promotionAugust 10–17, 2026

Against the version before it, 400M scored +2.433 points a round, range [+1.989, +2.876], over 40,960 rounds where both played the same deals. Against four opponents it had never trained with, it scored +2.476, range [+1.510, +3.466]. That’s the harder test, because a model can learn the habits of practice partners it sees every day. These numbers came from a different test than the one that later failed its audit.

It ran in the browser from 2026-08-10 until 850M replaced it on 2026-09-22, and it played every recorded game on this page. The strongest model goes to the app, and the browser gets the one before it.

What a point isAugust 2026

Faan is how the rules value a hand: whether it’s allowed to win, and how much it’s worth. Points are what that value turns into. They’re what changes hands at the end of a round, and every number in this article is points, averaged over thousands of simulated rounds. Under the rules I trained with, the cheapest winning hand off someone’s discard moves 8 points from them to the winner.

What the audit found, and what it didn’t touchAugust 19–21, 2026

After 11 days of trusting it, I went back and checked the test. It had two bugs.

The small one: restarting a job should have drawn a fresh random seed, and didn’t. Repeat runs I had counted as separate evidence were partly the same run again.

The expensive one was the margin of error. The test treated every round as independent. But rounds came in blocks that shared the same opponents and seating, so results inside a block moved together. The margin has to be computed over blocks, not rounds. Counting rounds made every range I published 1.6× to 10.8× narrower than it should have been.

In points: at 2,304 rounds, the test could only tell two versions apart if they differed by about 0.86 points a round. Every gap I was asking it about was roughly half that. It answered anyway, and it answered “no difference” every time, which is the only answer it could have given.

Here is the same question put to both tests: did the model improve between 700M and 850M? The old test reads +0.256 [−0.694, +1.243]. The rebuilt one, at 9,216 rounds, reads +0.514 [+0.074, +0.959]. The first range includes zero and settles nothing. The second clears it: the later model is better. Neither model changed between those readings. Only the test did.

The audit also found that a model made by averaging the 400M, 550M and 700M versions had never been tested at all.

It retracts one conclusion: that training past 400M had stopped helping. It does not reach the 400M promotion above or the oracle result below. I measured both on other tests, and both found gaps far too large to be noise.

The 850M promotionAugust 21, 2026In the browser

After 850,004,713 decisions over 51,363 iterations, 850M played 400M over 9,216 rounds on the rebuilt test, against its six opponents (the five bots from the cheating comparison, plus v9.2, the version before 400M), and gained +0.989 points a round, range [+0.546, +1.435]. If I allow for a different set of opponents, the range widens to [+0.308, +1.694], which still clears zero. The previous candidate failed this step.

Then three checks that the gain didn’t come from the particular opponents I picked. Head to head, with only the two versions at the table, it gained +1.117 ± 0.836. One opponent at a time, it came out ahead against 6 of the 6. And leaving any one opponent out of the average keeps the gain between +0.876 and +1.175, so no single opponent carries the result.

Last, the dealer. Training always puts the dealer in the same seat, so a model could learn something that only works from there. This check moves the dealer and re-runs the comparison on a fresh seed, over 40,960 rounds: +1.503 [+1.075, +1.926]. Even the worst dealer seat gains at least +0.511, and the four seats barely differ from each other (χ² 0.56 on 3 degrees of freedom).

I wrote all four checks down before I ran them. Otherwise I could have tried everything and kept whichever comparison looked best.

Three candidates measured against the 400-million-decision policy on the rebuilt 9,216-round exam. The 700-million checkpoint gained 0.475 points a round and a weight average of three checkpoints gained 0.491, and both look alike on this test; each failed the harder comparisons, so neither was promoted. The 850-million checkpoint gained 0.989 and was the only one to also beat the 400M policy head to head, survive resampling the opponents, and clear the randomised-dealer gate. Points gained per round against the model in the browser9,216 rounds each, identical walls for every candidate-0.50+0.5+1+1.5700Mfurther along the same run+0.475 · better on this test aloneWeight average400M, 550M and 700M blended+0.491 · better on this test alone850Mthe end of the run+0.989 · better on every testThe other tests: beating it in a direct match, resampling the opponents, and moving the dealer.
Post-400M candidates on the rebuilt 9,216-round exam
CandidatePoints per round95% intervalOutcome
700M+0.475[+0.029, +0.918]Not promoted
Weight average+0.491[+0.063, +0.917]Not promoted
850M+0.989[+0.546, +1.435]Promoted
Three candidates on the rebuilt test. Only one cleared the harder checks too.Rebuilt test: 9,216 rounds against the 400M policy, identical deals for every candidate.

The two that didn’t make it are why those checks are required. 700M gained +0.475 [+0.029, +0.918] against the same six, which looks like a result until you allow for different opponents, and the range opens to [−0.342, +1.310], including zero. The averaged model gained +0.491 [+0.063, +0.917] against the six, then nothing head to head: +0.593 ± 0.809. Each passed one check and failed the rest. On the chart above, the two look the same.

I stopped this run on 2026-08-13 at 711,154,042 decisions, because the test said it had stopped improving, and restarted it on 2026-08-19. If the restart had quietly started from scratch, the whole line would mean nothing, so it had to prove it was the same run. The stopped run had logged four iterations (42961–42964) after its last save, and the restarted one had to reproduce all four exactly before anything after them counted. It passed 850,004,713 decisions on 2026-08-21, 25.6 hours later. That version ran in the app from 2026-08-21 until h1 replaced it, and in the browser since 2026-09-22. The run kept going to 1.3B decisions; the next section has how that ended.

On identical deals, the average payment across the four versions went 13.375 → 13.456 → 13.851 → 14.365 at 400M, 550M, 700M and 850M. The last step, +0.514, is the biggest, and it falls inside the stretch I had cut off.

I’m not claiming this is the best model the game allows. It was the last version to pass every check at the time, and the curve hadn’t flattened yet. It flattened after a billion. The gain is real but modest: 450 million more decisions bought about one point a round. Of the measurements behind the old “nothing more to gain” argument, cheating against the bots has since been run on every version (last section). Cheating against a copy of itself and the fair search were only run on 400M.

The plateau, the flowers, and h1August 24 – September 15, 2026In the app

Before training from 1B to 1.3B decisions, I wrote down what would count as progress: a gain over 1B whose range stayed above zero. I expected about +0.74. Over 18,432 paired rounds it read −0.067 [−0.400, +0.265]. The learning rate stayed the same the whole run, so this wasn’t a schedule running out. More training with the same recipe wasn’t going to help.

Next, the flowers. In 10,000 rounds of 850M playing itself, flowers explained 37.4% of the variation in what winning hands paid, and 53.3% of wins scored flower faan. When I hid every flower from the model, its top choice changed in only 4.2% of 12,000 positions. So flowers add a lot of luck to the score and almost nothing to the decisions, and a model trained without them gives up little. I set both thresholds before running either test.

h1 starts from the 1B version and trains for 300,016,656 more decisions with HKMA Classical scoring, which ignores flowers. Nothing else changed. The fair comparison is the 1B → 1.3B stretch: same start, same number of decisions, old scoring. Scored with flowers, the way the app plays, h1 beat it by +0.512 [+0.170, +0.852], and beat 850M by +0.693 [+0.349, +1.027], each over 18,432 paired rounds. The dealer check, on a fresh seed over 40,960 rounds, passed all eight of its conditions: no dealer seat or round wind doing worse than −2.5 points, no illegal moves or scoring failures, no dropped rounds, and no gain that vanishes against opponents it never trained with. Overall h1 gained +2.454 [+2.005, +2.907] against 400M.

This is one training run, with one random seed. The gain beat the bar I set, +0.50, by only 0.012, so all I can say is that the true gain is somewhere between 0.17 and 0.85. Two more stretches of 40 million decisions each read −0.012 at 340M and −0.234 at 380M against h1, so the gain was a one-time step, not a new climb. The flower example in the story is one recorded hand, rescored; the 37% is the measurement.

Cheating, a fair search, and the language modelsAugust – September 2026

Outside calibration · not a ranking

On a calibrated 4,608-round ladder, the 400M policy beats the version before it and scripted opponents, while a purpose-built adversary is narrowly ahead. Separate 96-round reads put Claude Opus 5 at minus 0.08 and GPT-5.6 at minus 0.92 from the challenger side, with wide intervals crossing zero. Those two reads cannot rank the models. Calibrated opponent ladder4,608 rounds per opponentExpensive outside opinion96 rounds each · outside calibration-15-10-50Dedicated adversarytrained only to beat it, 300M decisions+1.19Previous championthe policy this replaced−3.15wait_awarescripted−11.90faan_awarescripted−13.22shallow_searchscripted−17.11-50+5Claude Opus 5−0.08 · [−5.23, +5.06]GPT-5.6−0.92 · [−5.62, +3.79]Not a rankingAll values are the opponent's net settlement points per round; negative means the 400M policy finished ahead.
Calibrated opponent ladder and separate language-model reads
OpponentGamesPoints from opponent side95% intervalCalibration
Dedicated adversary4,608+1.19[+0.96, +1.41]Calibrated ladder
Previous champion4,608−3.15[−4.24, −2.06]Calibrated ladder
wait_aware4,608−11.90[−12.85, −10.94]Calibrated ladder
faan_aware4,608−13.22[−14.24, −12.20]Calibrated ladder
shallow_search4,608−17.11[−18.06, −16.15]Calibrated ladder
Claude Opus 596−0.08[−5.23, +5.06]Outside calibration
GPT-5.696−0.92[−5.62, +3.79]Outside calibration
The 400M policy is ahead of the scripted opponents. The language-model comparisons remain uncertain.Scripted and neural opponents: 4,608 rounds each. Claude Opus 5 and GPT-5.6: 96 rounds each, with the two sides swapping seats so neither keeps a favourable position.

Impossible information

Over 4,608 paired rounds, a clairvoyant search that could see every concealed hand and the future wall gained 7.93 settlement points per round, with a 95 percent interval from 7.19 to 8.66. A legal belief search scored minus 0.382, with an interval from minus 0.920 to plus 0.157, so it could not be distinguished from no improvement. -10+2+4+6+8Clairvoyant oraclesees all concealed hands and the future wall+7.93Legal belief searchsamples plausible hidden states−0.382Headroom remains. The oracle is impossible information, not a deployable global maximum.
Clairvoyant oracle and legal belief search against the 400M policy
TechniquePaired roundsSettlement points per round95% intervalLegal at play time
Clairvoyant oracle4,608+7.93[+7.19, +8.66]No
Legal belief search4,608−0.382[−0.920, +0.157]Yes
Cheating gained about eight points per round against the 400M policy.4,608 paired rounds against the 400M policy. Measures what seeing hidden tiles is worth to this search, not the best possible score.

The first table gives each opponent’s points against 400M, so a negative number means my model won. One opponent beats it: the dedicated adversary, a separate model I trained to beat the frozen 400M and nothing else. It wins by about a point a round, but it is worse than 400M against every other opponent I tried, giving back about as much as it takes. That’s a weakness in 400M, not a stronger player, and I haven’t tested whether later versions share it.

For oracle guiding, following Suphx, I widened the model’s input with the other players’ hidden tiles, trained with them visible for 40 million decisions, and then played the same model on the same deals with the hidden tiles switched on and off. With them on, it won 27.9% of rounds instead of 27.3%, and threw someone their winning tile in 12.36% instead of 12.39%. It used what it saw to attack a little, and to defend not at all. Suphx’s next step fades the hidden tiles out; I didn’t run it, because there was no habit to keep.

The cheating search sees the other three hands and the order of the wall. This test used 400M, not the newer app model, playing a copy of itself. It won 2,170 rounds against 1,722 for the normal model, and the rounds where it threw someone their winning tile fell from 2,241 to 690. That’s what this particular search gains from seeing hidden tiles. It isn’t the best possible score, or what a fair player could gain. When I split the gain by what the search could see, 61.6% of it only appeared when it knew the hands and the wall together.

The comparison in the story is a separate test. Each version I’ve shipped played the same five bots on the same 9,216 deals per bot, once fair and once cheating, and every deal was played twice more, once under each set of rules: the training rules, which score flowers, and HKMA’s, which don’t. The ranges count blocks of six games that share a shuffle, as the audit above requires. Cheating is worth less here, 1.3 to 2.6 points, because the search assumes everyone plays like the model, and the bots don’t. Each version beats the one before it by a range that clears zero under both scorings; 850M over 400M, for example, is +0.96 [+0.46, +1.45] with flowers. h1 playing fair minus 400M cheating is −0.10 [−0.46, +0.28] under HKMA rules, and −0.74 [−1.24, −0.22] with flowers. At half as many deals, that last range was wide enough to include zero. The full test showed the gap.

Plays fairCheats: sees every hand and the wallShort line above or below a dot: its 95% range

Flowers scoredthe rules it trained with1214161820points per round against the botsh1, playing fairv9.2+2.6400M+2.6850M+2.4h1+1.8
Points per round against five scripted bots, 9,216 games per bot, identical deals for every version
VersionRulesPlays fairCheatsCheating adds95% range of the gain
v9.2flowers scored13.2615.87+2.61[+2.32, +2.89]
400Mflowers scored15.4618.06+2.60[+2.30, +2.91]
850Mflowers scored16.4218.80+2.37[+2.06, +2.69]
h1flowers scored17.3219.09+1.76[+1.45, +2.09]
With flowers scored, the order is the same. Playing fair, h1 falls a little short of 400M cheating.9,216 games against each of five bots per version, identical deals throughout, training rules (flowers scored). Lines show 95% ranges.

I also tried the fair version of the same idea: guess what the others might be holding, play each guess out, and throw whatever does best on average. It only uses what a player at the table could know. It scored −0.382 [−0.920, +0.157]. The range includes zero, so it found nothing. That’s one attempt coming back empty, not proof the idea can’t work.

I also had two language models play against 400M, 96 rounds each. Claude Opus 5 scored −0.08 with interval [−5.23, +5.06]; GPT-5.6 scored −0.92 with interval [−5.62, +3.79]. Both ranges are far too wide to tell either one apart from my model or from each other. The Opus run cost $162.30 for 1,742 valid calls, about $0.093 a decision. This run trained at about 130 million decisions a day. Buying a single training day at that rate would cost about $12,090,000.

They were also slow. Claude Opus 5 averaged 36.2 seconds a decision over 3,176 calls, and GPT-5.6 25.0 seconds over 3,708, both including the trip over the network. My model takes 0.41 milliseconds on a single processor core. I haven’t measured it on a phone yet. Neither lab publishes a parameter count for these models, so I can’t compare sizes. I can only compare the time.