Making Sensei

Making Sensei

I Couldn't Read the Tiles, So I Taught My Phone

Written by Vu Nguyen

September 2026

I started playing mahjong to meet people, and I still can’t read every tile. So I taught my phone to read them for me.

Can you name these four?

A mahjong set has 144 tiles. Some are easy: dots are dots . Others are Chinese characters, like the 9 of characters  or the West wind , and the flowers and seasons each have their own picture . I still get some of them wrong. That’s why the scanner exists: I wanted to point my phone at a table and know what I was looking at.

A real photo, the shipped model

Try the four numbered tiles.

A photo taken at night of face-up mahjong tiles scattered on a wooden floor, with face-down green-backed tiles to the right.

A photo I took at night, cropped to the tiles. Try the four numbered ones before you look.

A real photo and the shipped model’s real answers. Try the four first.Photographed at night on my floor, July 2026. Boxes from the shipped detector, run offline on this photo.

The last step is close to what happened the first time I took the app out. At a mahjong event after dark, I scanned over the shoulder of a friend. Tiles went missing, and moving the phone changed the answer. The new players were impressed anyway. The regulars were not. They could already read the table at a glance, which is what I had built the app to do for me.

I blamed the light. I didn’t yet know what the model was looking at.

What does the phone see?

The model is a detector: in one pass it finds every tile in a picture, draws a box around each, and names it. It doesn’t look at a tile the way you do. It gets a small, square version of the frame and guesses everywhere at once.

What happens to one frame

25,200 guesses

A photo taken at night of face-up mahjong tiles scattered on a wooden floor, with face-down green-backed tiles to the right. Stretched into a square.

The photo is squashed into a square of 640 by 640 pixels. The model lays grids over it, from fine to coarse, and every cell makes three guesses, each one a box, a tile name, and how sure it is. That’s 25,200 guesses for every frame, almost all of them about empty floor.

1 / 3
Every frame starts as tens of thousands of guesses and ends as one box per tile.Real candidate boxes from the shipped detector on the same photo.

The model gives every guess a score for how sure it is. Below 30%, the app ignores the guess. At 72% or more, it accepts the tile without asking. Anything in between gets a dashed box, and the app asks you to confirm it.

Was it the light, or the distance?

Now I could test what I had only guessed at the event. I took the close-up and made it darker, one step at a time. Then I made the tiles smaller, the way they look from farther away. I ran the shipped model on every version.

Test 1 · darker

100% of the light

A photo taken at night of face-up mahjong tiles scattered on a wooden floor, with face-down green-backed tiles to the right.

It found 44 of 44, and every one was sure enough to accept.

Test 2 · smaller

Tiles 54 pixels tall

A photo taken at night of face-up mahjong tiles scattered on a wooden floor, with face-down green-backed tiles to the right.

It found 44 of 44, and every one was sure enough to accept.

Darkness barely mattered. Tile size did.Each step is the shipped model run on the altered photo. Darkening is digital: no grain or blur.

At a tenth of the light it still found all 44. Shrinking was different: with the tiles 27 pixels tall it found 40, at 19 it found 23, and at 13 it found none. In the photo I took standing up, the tiles were about 18 pixels tall.

Pixels here means the model’s pixels. However sharp the camera is, every frame is squeezed to 640 by 640 before the model sees it, so a tile across the room is a smudge. That’s why the app asks you to hold the phone over the table. It’s also why I didn’t take the easy speed-up. A version with a smaller input ran faster, and the tiles it lost were the smallest ones.

Making a photo darker in software isn’t the same as a dim room, where the camera adds grain and blur. So this doesn’t prove light never matters. It does show I had been blaming the wrong thing first.

Why did it call a Winter tile “Autumn”?

One mistake the distance didn’t explain. The tile was Winter  and the first version of the detector said Autumn , with the same confidence it showed when it was right. A minute later, the model I shipped said Winter.

10:34 PM · first version
Mahjong Sensei capture of one tile. The physical tile is Winter, season number 4, but the app result says Autumn, season number 3.
Actual tile
Winter 4
App said
Autumn 3
10:35 PM · shipped model
The same Winter season tile a minute later, read by the shipped model, in darker framing. It identifies it correctly as Winter, season number 4.
Actual tile
Winter 4
App said
Winter 4
Same tile, one minute apart. The first version called it Autumn; the shipped model got it right.Author captures, August 15, 2026.

Seasons are still its weak spot. Of every tile it knows, it scores lowest on Autumn, and all four seasons are below its average. So the app never trusts one frame. A tile is only named once two separate frames, at least half a second apart, agree on it at 80% or more. If they later agree on something else, it doesn’t quietly switch. It asks you.

The dataset looked big until I checked it

A detector learns from examples: photos where someone has drawn a box around every tile and written its name. It guesses, compares its guess with the label, and adjusts, over and over. It can only be as good as those boxes.

I had assumed downloading a large dataset would save me time. The first one merged three downloads, and the same photos turned up more than once, in one source and across two. A photo that sits in both the training set and the test set turns the test into a memory test. The rarest tiles were the ones I needed help with: 5,097 examples of the most common tile, 891 of the bamboo flower.

So I combined six sources on one set of labels, removed the duplicates, and pasted 3,198 real flower and season crops onto real table photos. I froze the test set before adding them, so they couldn’t flatter the score.

  1. V1 · downloaded · 3 sources

    5.7× between the most and least common tile

    duplicates inside and across sources

  2. V2 · deduplicated · 4 sources

    4.6× between the most and least common tile

    exact and near-duplicate photos removed; a fourth source added

  3. V3 · combined · 6 sources

    1.7× between the most and least common tile

    six sources on one label map; flowers and seasons topped up with composites

Most and least common tile face in each dataset version, training split
VersionSourcesMost commonLeast commonGap
V1 downloaded38 of characters: 5,097bamboo flower: 8915.7×
V2 deduplicated4red dragon: 4,081bamboo flower: 8864.6×
V3 combined6green dragon: 6,5933 of bamboo: 3,8581.7×
Six sources and 3,198 composites cut the gap between tiles from 5.7× to 1.7×.Training split, tile faces only, recounted from the label files.

That fixed the counts: the flowers and seasons are no longer among the rarest tiles, and it reads them at a real table, like the Winter tile above. They’re still its weakest tiles, though: all four seasons score below its average. I never trained the same model without the pasted examples, so I can’t say how much of the improvement came from them.

Why didn’t I ship the one that scored best?

The model I’d started with was YOLO26. Its licence doesn’t allow it inside a closed-source app without a commercial licence, which cost about $5,000 a year. For a personal project, that was a no.

I didn’t want to pick a replacement by reading about it. I looked at twelve designs, trained four of them all the way, got all four running on an iPhone, and kept one. There are 71 distinct builds of the detector on my disk. Most were thrown away, some the same day I made them.

Then I found out I had timed every one of them wrong. This was a nights-and-weekends project, split across the website, the app, the vision model and the decision model, and somewhere in that juggling I had measured every detector in a development build instead of the optimised one that ships. Rerun properly, the newer transformer designs were 6 to 9 times faster than I had recorded; the one I shipped, Darknet, 1.7 times. So I timed them all again, in the build that ships, and compared them on those numbers.

The strongest candidate on paper was D-FINE. It won the offline scoring, mostly by drawing tighter boxes, which nothing in the app ever reads. On the phone it didn’t hold up. On a test table holding 37 tiles, one version reported 39.

There were 37 tiles on the table.

Darknet · shipped37every tile, and nothing else
D-FINE 51239two of them were not on the table
Two tiles that were never on the table still entered the count.Fixed 37-tile image, iPhone 15 Pro Max, August 2026.

Speed was the other half. Apple’s Neural Engine could run 97% of the shipped model’s operations, but only 43% of that D-FINE’s, so most of its work fell back to slower chips. On the graphics chip, the full-size D-FINE found 5 of the 37 tiles and reported no error. Over a whole hand, that is the difference between a live camera and a slideshow.

What does it do at a real table now?

Twenty frames a second was my floor. On an iPhone 15 Pro Max, the shipped detector can read about 33 frames a second, even with the phone already hot, and it held 27 frames a second for five continuous minutes without the phone heating up. Hold it over a table for a whole hand and it keeps up.

Mahjong Sensei's Tracker view over a wooden table scattered with mahjong tiles. Each tile carries a coloured box and a label such as 7 Dot, Red, West or 9 Bam, and a chip reads 30 tiles in view.
Thirty tiles read in one pass, on a real table in a dim room.Tracker screenshot, shipped model, August 2026.

The irony is that I started playing mahjong to get out and meet people, and now I’m spending those nights trying to ship the app instead.

Try the product

See what it reads at your table.

If you know the tiles better than I do, tell me where it gets them wrong.

Evidence record

Measurements, methods, and limits

These are the July to September 2026 measurements behind the story. The app ships one detector, Darknet Mobile M, and every photo figure above is its output. A score is the model’s confidence that a box holds a particular tile, from 0 to 1. A development build is the unoptimised version of the app I run while working on it; a Release build is the one that ships. Phone figures are from an iPhone 15 Pro Max.

How the photo figures were madeSeptember 23, 2026

The quiz, the walkthrough and both tests show boxes the shipped model drew on this photo, run on a computer instead of the phone, with the app’s own rules: boxes under 0.30 are dropped, overlapping boxes are merged whatever tile they name, and anything at 0.72 or above is accepted without review. The page runs no model.

The photo is one frame, IMG_0157.HEIC, which I took at night under room light, standing, with the whole floor in view. The close-up is a crop of the same pixels, not a second photo. In the full frame the model found 7 boxes, none at or above 0.72, and one of them is on a book. In the crop it found 44, all at or above 0.72, and I checked every name by eye. Before merging, those tiles had collected 17 to 39 overlapping guesses each, 1,308 in all.

The light test multiplies every pixel of the crop by a factor, the same arithmetic as the page’s brightness filter, so what you see is what the model saw. All 44 tiles survive down to 10% and none at 2.5%. The size test pastes the crop, shrunk, onto a canvas of the floor’s average colour. Tile heights are the median box height in the model’s input.

This is one photo, not a benchmark. Darkening in software adds no sensor noise, motion blur or colour shift, so it tests brightness and nothing else. A real dim room could still do more damage than this test shows.

How the Tracker names a tileSeptember 2026

One frame only suggests. The Tracker names a tile after two strong reads of the same face, each scoring at least 0.80, at least half a second apart. If two later strong reads agree on a different face, it clears the name and asks the user rather than switching on its own. A correction by the user holds for that tile until it leaves the table.

The Winter and Autumn captures in the story came from the Lens view, which reads one tile at a time, a minute apart. The first was the first version of the detector, picked from the model menu that development builds have; the second was Darknet Mobile M, the model that ships. An earlier version of this page said both came from the shipped model. They didn’t.

The dataset, pass by passJuly – August 2026
PassSourcesTraining imagesMost common faceLeast common face
V1 · downloaded316,0348 of characters, 5,097bamboo flower, 891
V2 · deduplicated415,159red dragon, 4,081bamboo flower, 886
V3 · combined624,275green dragon, 6,5933 of bamboo, 3,858

Counts are training boxes per tile face, recounted from the label files. The labels also have a class for face-down tiles, and V1 held only six of those. An earlier version of this page compared 5,097 with those six and called the data 850-to-one out of balance. That was true of tile backs, which the app never has to name. Between faces it was 5.7-to-one.

I removed exact duplicates by file hash and near-duplicates by a perceptual hash, which catches the same photo resized or recompressed. One source almost disappeared: its 4,893 training images fell to 16, because nearly all of them were already in another download. I left a seventh download out because its labels were poor. The final set holds 28,369 images and 226,923 labelled boxes across 43 classes: 24,275 for training, 3,467 for validation and 627 for the test. The training part is 18,239 real photos, 3,198 composites and 2,838 repeated copies of rare tiles.

Validation and test changed once, when V3 brought in sources that came with their own splits. From then on they were frozen. The composites and copies went into training only, and I opened the test once, after I had already picked the winner on validation.

The seasons still score lowest on that test: Spring 0.59, Summer 0.61, Autumn 0.58, Winter 0.64, against an average of 0.67 across all 43 classes. I never trained the same model without the composites, so this can’t tell you what they added.

How accuracy was scoredAugust 9, 2026

Detection accuracy is scored by how much a predicted box overlaps the true one. At the looser setting, a box counts if it covers at least half the tile, which in practice means “did you find the right tile”. At the stricter setting it has to fit much more tightly, so it measures how neatly the box was drawn. (Researchers write these as mAP@0.50 and mAP@0.75.) The app only uses which tile and roughly where. Nothing in it reads a box edge, so the stricter score buys the product nothing.

Three detectors compared on two metrics. Their span is 0.0070 on mAP at IoU 0.50, which measures whether the tile was found. Their span is 0.1241 on mAP at IoU 0.75, which rewards a tighter box. Darknet Mobile M finds the tile comparably well and draws the loosest box. 0.800.850.900.95Found the right tileDrew a tight boxDarknetshipped0.9610.803D-FINE0.9650.927YOLO26-m0.9580.923
Detector accuracy at two IoU thresholds, 627-image hidden test split
ModelmAP@0.50mAP@0.75AP@0.30
Darknet0.96100.80270.9657
D-FINE0.96470.92680.9679
YOLO26-m0.95770.92270.9596
Three detectors, two thresholds. The spread is 0.0070 on finding the tile and 0.1241 on box tightness.627-image hidden test, scored through the production pipeline, August 2026.

The chart has three models rather than four because RF-DETR’s tight-box score was never recorded. Between the shipped model and D-FINE alone the gaps are 0.0037 and 0.1241. Every row comes from the same scoring run. Mixing runs would compare a model against itself under a different scorer.

The smaller-input probe. Feeding the shipped model 512-pixel frames instead of 640 sped it up from 18 to 29 frames a second (54.7 to 34.1 ms a frame) and dropped the looser score from 0.9605 to 0.9458, more than the 0.01 I had allowed. 36 of 43 classes got worse, and split by tile size, the smallest tiles lost the most.

Five-minute phone run, both modelsAugust 20, 2026Development build
Frames per second over a five-minute run on one iPhone. The shipped detector starts near 40 and settles around 27, staying above the 20 frames a second the camera needs to feel live. The benchmark winner starts near 19 and settles around 10, below that line for the whole run. 102030401 min2 min3 min4 min5 minfpsDarknet · shippedD-FINE 512below this the camera stops feeling liveDarknet25.1 fps · shippedD-FINE 51210.1 fps
Frames per second by 30-second window, five-minute sustained run, iPhone 15 Pro Max
Model0.5 min1 min1.5 min2 min2.5 min3 min3.5 min4 min4.5 min5 min
Darknet39.726.528.926.828.128.127.327.526.225.1
D-FINE 51218.911.110.810.510.610.810.710.710.510.1
Five minutes of continuous scanning. Only one stays fast enough to feel real-time.iPhone 15 Pro Max, fixed 37-tile image, development build (-Onone).
ModelFramesMean fpsSettles atMedian fpsTilesNeural Engine opsPeak MB
Darknet8,53128.427.1830.037136/140210
D-FINE 5123,44411.510.6611.839385/904224

I ran both models on the same 37-tile image for five minutes each, on the evening of August 20, about ninety minutes apart. The D-FINE row is a mixed-precision 512-pixel version, not the full-size model scored above. The phone never got warm enough for iOS to flag it, and the battery reading didn’t move in either run, so nothing here is a throttling result. The only record of either run is the results panel the bench copies out.

These runs came from a development build (-Onone). That slows every model, and D-FINE’s kind far more than the shipped one (see the table in the search disclosure). So D-FINE’s speeds here are a floor: in the build that ships it can only be faster. Tile counts and where the operations run don’t depend on the build and can be read as they are.

The ship decision rests on a Release sweep instead: the median of 50 timed runs per setting, each started with the phone already hot, so every time is a worst case. Darknet Mobile M ran at about 33 frames a second on the Neural Engine (30.3 ms a frame). D-FINE’s fastest were about 24 (41.6 ms, 512) and 17 (60.5 ms, full size) frames a second, both on the CPU. On the graphics chip, D-FINE returned 5 and 9 tiles of 37, with no crash, no error and nothing in the log. That is why the app won’t run that model on the graphics chip.

CPU time per frame is a rough stand-in for cost to the phone, not energy: iOS has no public way to measure power, or how busy the graphics chip and Neural Engine are. Operation counts are where the model’s plan puts each step before it runs, not where it ran. Peak memory is what iOS charges the app for.

The checks before it shippedAugust 9–13, 2026In the app

Every time a model changes format it can quietly change its answers, so I checked each step against the one before. First the original training output against the ONNX export, then the export against the iPhone’s Core ML model, on 24 fixed test inputs, with and without the Neural Engine. Then I froze the model that won on validation and opened the 627-image test once. It scored 0.967 on the release measure, the looser accuracy score above with a box counting once it covers 30% of the tile, against a bar of 0.77 I set before looking.

The check of the finished app and the later phone run are separate records, not one automated pipeline. The first makes sure the exact App Store build contains the real model rather than a demo fallback. The second shows how it behaves over time, on one phone and one scene.

The five release gates for the shipped detector
GateQuestionRecorded result
Provenance + licenceCan I legally ship every input and weight?Apache-2.0 path documented
Native → ONNX parityDid the first export preserve the model?Outputs matched within tolerance
Core ML parityDoes the phone package agree with ONNX?24 fixed inputs · with and without the Neural Engine
Frozen hidden testDoes the validation winner clear the product metric?0.967 against a bar of 0.77
Built-app qualificationDid the exact Release bundle survive on the device?Bundle preflight + later phone run
Darknet had to clear every check, from licence to the exact Release build on a phone.Release record, August 2026.