T

MaiSay Design

The design canon — every screen, one page.
Private while it is in review.

Voice Studio · 142 screens

MaiSay

Voice Studio — the design canon, every screen, one page.

Read the design ledger →
Laws every screen obeysone pastel, one jobverdicts stand on paper, never a pastelCLOSE is magenta, not amberfour states — the fourth is not scoredink means unscored, never correctblack is the action colourtext floor 12every verdict carries a wordone state per screen — no hidden states

The front door & first run

10

One login for EVERYONE: you sign in, the server returns your role, and the client routes invisibly — an operator lands in the console, everyone else in the app. So the door is warm learner branding with no staff chrome, and there is no separate sign-up screen anywhere.

login One front door for EVERYONE, and it is the FIRST screen — you sign in, the server returns your role, the client routes silently. No sign-up screen exists: continuing creates the account. ONE TAP LEADS — Apple, Google, Facebook — because a cold door must open in one tap; email falls back below the divider. Glyphs are ink monochrome, and nothing here is black: a front door has no committed action, every path on it is a first tap.
login code The six-digit code. One refusal sentence for every kind of wrong, deliberately — a mistyped email and a guessed staff handle get the identical message, so the screen never tells anyone which accounts exist.
onboarding language Which language do we teach you THROUGH? Five live rows — English, Tiếng Việt, Español, Bahasa Indonesia, 한국어 — each in its own script with its sub-line already in that language, and NO coming-soon state anywhere: the mock draws the target set, the build decides which rows render. The choice silently rebranches the product, so the preview is the argument rather than a claim — and it now hangs off the row that caused it: the selected answer and its mint panel are ONE bordered unit, 中国 → Trung Quốc, the head start a Vietnamese learner already owns, with the languages nobody picked continuing underneath.
onboarding mic Earn the permission before the OS asks for it. A real taste of what the microphone unlocks, then the request — and the copy is sentence-true, because what comes next is a five-syllable sentence and not one word. Tap to speak, never hold — and never 'nothing is recorded', because retained audio is what the product is built on.
onboarding take prompt The take, the ask. The target, the example, and a mic waiting — onboarding chrome, not lesson chrome, because there is no run to be 1 of yet. The target is 你好,我来了, founder-picked: it contains the one piece of Chinese nearly everyone owns, the rest is half-guessable off the pinyin, and it says the thing worth saying at minute three — hello, here I come.
onboarding take recording The mic is open and the learner is speaking. The state is said once, on the level meter — the only element that goes still when the mic is dead — and there is NO CLOCK anywhere: duration is a fact about a take you have, not a display while you make one. One button, no cancel.
onboarding first take Their own voice, scored, at minute three — and the MAP leads, not the number. 你好,我来了 comes back as five coloured parts with one named observation — two dips in a row are hard, Mandarin lifts the first: ní hǎo — and the percent demoted to a caption. The accounting line names the grain rather than a count the five boxes never add up to. A number invites comparison; a diagnosis invites trust.
onboarding goal What is driving this, and by when? A reason, not a date picker — and the consequence now hangs off the row that caused it: the selected answer and its peach panel are ONE bordered unit, with the five options the learner did not pick continuing underneath. The no-deadline row stays last and is worded as a real driver rather than a level, because deadline answers lead and “starting from zero” made the why screen read as the band question one screen early.
onboarding band Where are you starting from? Three taps, no placement test — a speaking exam measures pronunciation, which is a weak proxy for the vocabulary band it was setting, and it spent a beginner's first ten minutes on six failures. It sets where you begin, never what you can reach. The consequence is attached to the band that caused it: one bordered unit, the answer on white, the route it opens in peach below it. Jump-ahead now has its counterpart in the same card — placed too high, step back the same way — because a self-reported band is only safe if being wrong is symmetrical.
onboarding plan The closing beat of onboarding, and it asks NOTHING — now drawn so that it LOOKS like it asks nothing. The three answers come back as checklist lines ticking off on the peach under the head Building your plan, because as white option cards they read as one more form. Each tick sits in a 24px pastel coin carrying its domain colour — peach the track, mint the words, rose the speaking — solid and lifted, never outlined, because an outlined circle with a check in it is a radio button; a hairline rule between the lines makes it a finished ledger. Mai presents from the right, mirrored so the palm lands on the headline; beneath sits the twelve-tile wall of everything real the plan comes with, flattened so nothing but the one black door looks tappable. The swiss knife opened, with the learner's blade first — and the blade is walk A's: the trip they said they were taking, and the pack as a bonus stop beside their route rather than a topic dealt inside it.

Home — the path

9

Home is not just navigation, it is the conversion surface: people open the app and read the route to work out what this thing can eventually do for them. So it runs all the way to the last goal, at one node size, with nothing faded and nothing shrunk.

path day one The same home on day one — the zero state, and deliberately a different screen. Nothing is behind her, so every node is untouched and the trail is dotted the whole way; the card's door reads Start, not Continue. Metrics honestly zero without the shame-shaped emptiness of a screen that reads 'you have done nothing'. The eight goals stay at full weight, because that is what a new user is deciding whether to trust. The track strip is gone: the header eyebrow names the track and opens the picker.
path Where am I, and what can this eventually do for me? Day 12: goal 1 is behind her, goal 2 is two steps in. Three node states on one route — cleared, current, ahead — told apart by a joined-up trail, a check, and elevation, with no new hue and no verdict colour. Home is the conversion surface, so the goal card carries the next step as a black action rather than leaving it below the fold, the route still runs all the way to goal 8, and nothing fades or shrinks down the scroll.
path vnext RATIFIED (2026-08-13) — home with the track chips retired: the eyebrow reads RECOMMENDED TRACK and opens the picker. The interest pack arrives as a BONUS stop drawn OFF the trail on a short branch stub — optional, never counted, never blocking, titled as a can-do claim and wearing its own +4 min, because the route minutes count the spine. A node standing in the route's own line that the route's own count ignores was a screen disagreeing with itself.
path review due The same home above 20 words due. The greeting's quiet review line is REPLACED — not joined — by one compact mint offer under the goal card: the standalone Review lesson, a real metered half-speaking lesson built from due words only. The goal card keeps the black bar and stays the primary door; the offer takes the row beneath it at half the height and hands the route back untouched. No verdict colour anywhere near it: due is a fact about the clock, never a grade. At 20 or fewer due the home is exactly path-vnext.
path picker RATIFIED (2026-08-13) — the sheet the eyebrow opens. FOUR live tracks, drawn equal: each with its promise, one real detail and the feature set it alone has (HSK's detail is the syllabus meter — words you can already SAY, never a walked track). Availability is not drawn: the build serves the tracks whose sequences exist, and the mock draws the target state. The foot line is what makes switching safe: your words and your speak-streak come with you.
path afternoon The same home a few hours on — the daytime slot's second study pose. Reading on the floor rather than at the desk: one frame held from breakfast to dinner reads as decoration, two read as a day passing alongside the learner's.
path evening The same home after dark — the cadence law drawn rather than described. One thing changes besides the word evening: Mai winds down, lids heavy behind a small yawn. The swap is time-of-day and never activity-conditional, so this screen renders identically for a learner who spoke twelve times today and one who has not opened the app since breakfast. Never-guilt is the whole test of it.
path night Past about ten — the last slot of the cadence. Only the pose changes; the words still read “Good evening”, because a “good night” at the moment someone is opening the app is a goodbye.
path return Coming back after a long gap — the cadence law's third face, and the screen never-guilt was written for. Mai waves, because a return is an arrival. The broken streak is not printed and not zeroed; no elapsed time is stated anywhere, because '3 weeks away' is the same guilt wearing arithmetic. What takes the slot is the fact that helps: the route is exactly where it was left.

The HSK track — walk B, and the conversion showcase

5

Ready for the New HSK (3.0) — the approved headline, and these are the screens that license it. The first three are the SECOND ONBOARDING WALK: a learner who says 'exam' at ask 3 gets her own consequence, her own ask 4 and her own receipt, and the chain runs end to end. The specific claim leads everywhere: the new exam makes speaking compulsory from New HSK 3, and its level-3 paper opens with the exercise this app scores every day. Drawn to showcase standard on the founder's ruling that a feature can earn its place as a marketing asset — and honest to the boundary: preparation, never simulation, and never a claim to cover both syllabuses.

onboarding goal hsk Ask 3 with the exam selected — the head of walk B. The same six reasons in the same order; what changes is the consequence, because the exam answer swaps the whole sequence rather than reordering it. The peach panel states what the route BECOMES, in the syllabus's own five arcs, with the goal count per arc adding to the seventeen that finish New HSK 1.
onboarding band hsk The same ask 4, shown when the driver was the exam — and the rows now report KNOWLEDGE, not certificates: New to Chinese / basic words and short sentences / 1,000+ words and a simple conversation. Certificate rows over-credited the papers actually in circulation, and a learner read against the wrong word list starts one band too high. The attached consequence places her on the New HSK ladder and names the track as already set, reconciling to the goal with every other HSK screen.
onboarding plan hsk Walk B's receipt: the same assembly, with the HSK track on the first line and TWO checklist lines rather than three — an exam driver gets no interest pack, because the syllabus is already the content and a row invented to keep the rhythm would be the one dishonest line on a screen whose whole claim is that this was assembled for you.
path day one hsk Walk B's reveal — the same route screen with the track already set, named in the header eyebrow. Ruling A1: a learner who declares the exam sees the exam recognize her before the reveal ends, so the goal card carries the one specific claim — the oral paper is compulsory from New HSK 3, and its level-3 paper opens with the exercise scored here every day. Re-cast to the curriculum of record: goal 1 of SEVENTEEN, and goals 2–6 are the sequence's own.
path hsk The HSK track, previewed — and the round's marketing screenshot. The specific claim leads (speaking is compulsory from New HSK 3, whose paper opens with listen-and-repeat), the companion line places us beside her CLASS rather than against it, the word meter states what she can already say out of the syllabus's own 300, and the route is New HSK 1's five arcs carrying all seventeen goals. The 2.0/3.0 distinction is said once, in plain words, and no longer claims to prepare her for both. No HSK progress is asserted: An is browsing, not enrolled.

The first lesson — one run, six exercises

7

Goal 1, step 1: the only lesson a new learner can reach, and the run the first ten minutes ends in. One objective on every screen, one counter from 1 / 6 to 6 / 6, three atoms and no fourth — 我 · 是 · 越南. It ladders into the canonical sentence on purpose: those three atoms are three of its five syllables, and step 4 drills exactly the initial the learner gets wrong on it later. Read these in order — if the run does not walk, nothing else in the gallery matters.

lesson words meet Từ mới — the meet beat, and NOT an exercise. Every genuinely new atom is given before it is asked for: hanzi, pinyin, gloss, audio, the why-line, the Hán-Việt note where one is authored, and one tap onward. No counter and a progress bar that does not move, because a beat takes no slot in the lesson. The fix for “everything in the app is a test”.
lesson match Step 1 of 6. Six tiles, three words, no clock. The densest exercise in the catalog and the only one that cannot be failed, only finished — which is why it opens the first lesson, and why the route's first node points here rather than at a five-syllable sentence. Cleared pairs go ink, never green: nothing here was judged, so nothing claims a verdict.
lesson word check Step 2 of 6. WORD CHECK in the one direction the first lesson runs it: audio → gloss, with no hanzi and no pinyin above the play control. A word check that prints 我 over the options is a reading test wearing an audio button — the job is to hold the SOUND matched thirty seconds ago. Mint, because this is still the library's two minutes.
lesson say wo Step 3 of 6, and the first time the learner speaks inside a lesson. ATOM GRAIN: 我 is one syllable and two parts, drawn in the same fixed frame the sentence lesson uses — same target card, same analyzer, same mic, at the size the first minute of speaking actually is. The progress bar turns rose here and stays there.
lesson listen Step 4 of 6. Can I hear the difference? 是 shì against 四 sì — identical final, identical tone, one initial apart — and it sits at step 4 because step 5 is 是. You cannot produce a distinction you cannot hear, so the learner meets it one screen before they are asked to make it. Listen and choose is one family carrying three names: this, Which tone? and Catch it.
lesson say shi Step 5 of 6, and where step 4 pays off — without a paragraph saying so. The learner picked the curled tongue out of a minimal pair one screen ago; here the sh span comes back green and the delta chip reads 'sh · held'. And it is NOT all clear: tone 4 stayed short, 83%, one part to repair. The celebration is spent once, at step 6.
lesson say yuenan Step 6 of 6 — the climax, and one of the product's two win moments. The first attempt at it stalled 南's rise halfway at 94%; this take comes back at 100% and the part that comes good is the exact part that did not. A repair the learner can hear inside one exercise is what makes the celebration affordable here.

The speaking lesson — the moat

15

One fixed frame, fifteen moments. Three independent verdicts per syllable — initial, final, tone — and the character reports the worst. The failure states matter as much as the happy one: a false red on a pronunciation app is far more damaging than a false green, and an ambiguous error screen is a false red the learner writes themselves.

speaking prompt Before you speak. The same frame with no verdict anywhere — bars on the empty track, no score, because a zeroed score is a lie. The screen a learner sees most often.
speaking arming Between the tap and the tone. The control goes inert and says wait — the screen must never claim recording before the cue, because a learner speaking into a mic that has not opened gets a bad score for a take that never happened.
speaking recording Live, and the LEVEL METER is what says so — it goes still when the mic is dead, which a banner cannot. No clock: duration is a fact about a take you have, not a display while you are making one. One button, tap to stop, no cancel anywhere in this loop.
speaking sending The NETWORK half of the wait, named as itself. 'Your voice is going up' is a problem a person can act on; 'the model is thinking' is not — so the screen says which. No fake progress.
speaking scoring The MODEL half of the wait. The card holds, no spinner theatre, and the listen-back is already available because the audio exists on the device.
speaking What exactly did I get wrong, and how do I fix it? The moat: three independent verdicts per syllable, the character reporting the worst, and one named physical repair.
speaking tone The SECOND repair, and the reason the first screen only names it. 是 has landed; now the rising tone on 南 stalled halfway — and a stalled tone is a picture, not a sentence, so the target-vs-you contour is the whole screen. One change at a time, applied to layout.
speaking retry The win moment. The retry REPLACED the first result in place — no Take 1 / Take 2 list, the old take survives only as a delta line. Every part judged and clear, so this is the one screen where the characters themselves go green.
speaking skip The other branch: the repair pass came back red on the same part, and Skip appears beside it. A red take never walls the lesson (ruled 2026-07-31) — but the black stays on 'Say it again', nothing turns green, and the panel says why skipping is honest: 是 stays marked as not landed and review brings it back within a day or two.
speaking uncertain A part could not be scored. It renders as ABSENCE — ink mark, bar left on the empty track — and the headline withholds its credit: 92%, 12 of 13 parts scored. This screen exists to prove the fourth state is visible. No celebration on a partial.
speaking below floor Below the coverage floor, so there is NO score at all. The card stays pre-speak with an honest note, the copy blames the conditions, and the retry is free.
speaking mic denied The microphone is off at the OS. The record control wears the state and carries Open Settings — a settings problem, not a failure, and never a dark error screen.
speaking no speech Nothing was heard. Blames the room, the distance, the timing — never the learner. This is where trust is won or lost, because the learner's instinct is to assume they did something wrong.
speaking offline Scoring is unreachable. The take is saved on the device and nothing is lost; the practice loop stays warm. No modal, no red, no dead end.
speaking long Twenty-two characters — the stress test. Tokens wrap in reading order, the translation collapses first, pinyin and hanzi never shrink below production scale, and the score never hides. It scrolls, and that is correct.

The other exercise types

23

Speaking is the only hard modality; every other type is a near-free RENDER of the same atom. SEVEN families carrying ten learner-facing names — adding a variant inside a family costs nothing, adding a family is a second product and needs a decision, not a ticket. A lesson is 6–10 of them and always ends on speaking. Nearly every one of these is recognition — and the exception is drawn here too: Say it from memory, where recall happens before the mouth opens, and the outcome is caused by what the learner retrieves.

choice Which reply belongs next? Listen without subtitles, then choose — with the hint behind a deliberate tap, so the ear gets its chance first.
choice wrong The miss. The tapped answer is a fluent sentence answering a DIFFERENT question, and the explanation names that — a comprehension miss, so no analyzer, no bars, no percent.
lesson meaning What does this word mean? WORD CHECK in its hanzi-to-gloss direction — the cue is the characters and nothing else, because hanzi plus pinyin plus audio before the pick hands you three routes to an answer the exercise tests one route to.
lesson recognize Which character is it? WORD CHECK, audio to hanzi — the ear is the whole cue, so the spelling is withheld until the answer. The fix is the stroke difference DRAWN, not described. Recognition is first-class; producing strokes is explicitly out of scope.
lesson word check pinyin How is this one said? WORD CHECK, hanzi to pinyin — the direction scheduled immediately before a Say it, so every distractor is a mistake the learner is about to make out loud. One component, eight directions, three drawings of it.
lesson read character READ THE CHARACTER — the seventh family, the character as an object. Dealt on a word the learner has SPOKEN but never read; the four options are the true readings of four words they know, so a wrong pick is the character wearing another word's voice. The object card below gives structure — radical, strokes, parts — the one fact that doesn't leak the sound. The atom speaks at the marking moment (ruled 2026-08-06), never before.
lesson tone The syllable has played; five contours to choose between, before the mouth ever opens. The character is withheld on purpose — this is perception, not a reading check. All five lines are ink, because a tone is a category and verdict colours are never categories.
lesson catch it Where does the word land? The third name in the listening family and a genuinely third thing: no competing words at all, just one stream and a timeline. A word you know cold in isolation is a word you still cannot find at speed inside eleven other syllables.
lesson catch scene CATCH IT at dialogue grain — two turns of a scene, heard once, then a question about what was actually said. The title is contextual (What did Mei order?), not an eleventh name; no word of the exchange is printed, because the transcript IS the answer.
lesson pattern meet The one beat in a run that ASKS NOTHING — a grammar pattern handed over before the learner needs it: the template as a diagram, and one authored line saying what a speaker DOES with it. A pattern with no line yet simply shows a shorter screen.
lesson fill FILL IT IN — the tenth name, and the HSK gap-fill shape. A sentence with one word missing, read not heard, four words for the slot. The pinyin line is blanked at the same position, because printing chī under a sentence missing 吃 hands over the answer.
lesson fill wrong The wrong word went into the gap. The marked state names the confusion — which verb takes 面条 — with the collocation drawn, not lectured. The first drawn distractor set failed the one-right-answer gate and the review round caught it; the fix is in the head comment as a worked example of why cloze sets are reviewed content.
lesson your turn Mei asks, three replies, one fits. Picking is step 1 and speaking is step 2, and the screen says so rather than pretending the take has happened. The scaffold is not a softening: open production graded against one target marks a right answer wrong.
lesson your turn correct The same screen, marked — recoloured and not one pixel moved (§5.8d). The learner picked a fluent reply to a DIFFERENT question; the teaching arrives BELOW the options, and the fitting reply — never their own wrong pick — goes to the next screen.
lesson your turn speaking Step 2, and it is a screen of its own (founder, 2026-08-02). The pick screen closed with a continue; here the question stays fully readable, the replies are gone, and the fitting reply stands in the speaking card. Nobody scrolls to find a microphone.
lesson picture A photograph replaces the written gloss, so meaning arrives as a scene instead of a translation. Labelled HSK-style and honest about it: the real spoken exam's picture task is open-answer and we hand you the sentence. The one screen that needs a network.
lesson say memory SAY IT · FROM MEMORY — the round's centrepiece, and the one exercise where the learner causes the outcome by retrieving. The cue is the Vietnamese gloss and a picture; pinyin, hanzi and the model audio all wait behind one recorded peek row.
lesson say memory peek The learner opened the pinyin. The answer is on screen, the take is marked Peeked — a state label in the SRS pill's own slot, not an apology — and the screen says nothing in its defence: a peeked take scores pronunciation, and the numbers do the rest.
lesson say memory mismatch A different word came out — and nothing turns red. The take fails part alignment, lands below the coverage floor, and the screen says so honestly: it may have been perfectly good Chinese, it just wasn't this word. NOT SCORED where a percent would sit, a free retry, and the peek row waiting. A false red here is the exact trust-killer the amber doctrine exists to prevent.
conversation A known reply inside a real scene. Scripted, because the scorer needs a known target — your line arrives as a prompt in your own bubble, then gains its verdict in place.
conversation known The same scene met again, and the only screen that draws the gloss fade. A line withdraws its Vietnamese once every atom in it sits at SRS ladder step ≥ 2; a line holding a never-met atom never fades, so stimulus stays comprehensible; and a tap brings any of them back — recorded, costless. The learner’s own reply keeps its pinyin, which is a speaking aid and not a gloss.
conversation recording The take, without leaving the scene. A conversation turn records in the conversation's own chrome — transcript behind, Mei's question still readable, the draft bubble unchanged — because handing the learner to the shared speaking screen mid-scene means step 2 of the goal wearing step 3's counter. Turn badge, no lesson counter, no clock.
conversation feedback How did that turn go — without losing the conversation? A BOTTOM SHEET over the transcript, because a conversation has several scored turns and a lesson exercise has one: the turn stays visible behind it, dismissing lands you exactly where you were, and the next one is a tap rather than two navigations. The signature target-vs-you contour, shown open rather than hidden behind a tap.

One game, every state — Build & say

8

The first family drawn as a STATE MACHINE rather than as a specimen, because the states nobody draws are the states nobody rules — and the one this family was missing was the miss. Walked on the device 2026-08-16: a wrong order got no colour, no diff, and no sign of what the learner had actually built, because the marking swapped their tiles out for the correct sentence in the same seat. Read the six in order and the mechanic is one continuous take — the same four words, the same learner, from the deal to the microphone, with state 3's line being exactly the line state 4 marks. Two rulings come out of it: the built line SURVIVES its own marking (§5.8d — the challenge is never spent, and every other family already kept the learner's answer on screen), and a mis-ordered sentence is CLOSE, not FIX — every word right, one has to move, which is §2.2's definition of the colour word for word — though the WORD stays off the screen, because CLEAR/CLOSE/FIX are analyzer vocabulary and nothing here was heard by the scorer. Which is what lets MAI into the room (founder asked, 2026-08-16): with the analyzer words gone, exclusion 1 governs the speaking room and the analyzer, and step 1 finally has the cue room every other family already had. She is on all five build states or none — a companion who left at the marking would be leaving because the learner got it wrong. And room 2 has her LISTENING — the founder pushed on that the same day and was right: the mic is not live on a prompt, it is live one screen later, and speaking-prompt.html and lesson-say-memory.html already seat her on exactly that screen of a speaking exercise. She stops at speaking-arming.html and does not come back. This is the shape every other family owes a pass in.

lesson build say start State 1 of 5. Four tiles, an empty line, and a control that is OFF with the reason as its label. The strip carries no Example: the assembly IS the question, so the sentence and its recording are withheld until the marking. The tray sits above the line because only the line grows — tap targets that move under the thumb are 5.8d with a different cause, and the Flutter build has these two reversed.
lesson build say first word The first word a learner ever places — and the answer to "nobody finds the tap-gloss". The defect is sharper than not knowing: a WAITING tile's tap places and a SPENT one's asks the meaning, so the learner who experiments correctly concludes *tap = place* and never taps again. The fix is that the mechanic demonstrates itself — the first tile to become spent opens its own gloss, once, ever. No tooltip (explanatory copy on a learner surface, fires once, and it would have to put a state-dependent rule in a sentence). No bounce (a bounce says tap me, and a waiting tap PLACES — it would advertise the half nobody has trouble with). And it writes no GlossTaps row: a demonstration that logs itself as a tap poisons the one instrument that can tell us whether it worked.
lesson build say building Mid-build, at a learner-initiated tap. Two placed, two waiting — and the two rules that only exist here. A spent tile is DRAWN, not deleted, so the tray never changes height. And the meaning switches on when the tile is spent: a waiting tile's one tap places it, a spent one's asks what it means (§5.8f).
lesson build say ready State 3 of 5. The tray reaches empty and one control wakes up in place — Place every word becomes Check the order, same seat, same weight. The answer is a PERMUTATION, so there is no partial submit and no live validation: nothing on this screen grades a line the learner has not committed.
lesson build say marked State 4 of 5 — THE SCREEN THE FAMILY DID NOT HAVE (founder, 2026-08-16). The marking used to swap the learner's tiles out for the correct sentence, so the one artefact a correction needs was destroyed at the moment of correcting. Now the built line STAYS, one tile carries CLOSE — every word right, one has to move — and the sentence that works arrives below it. Close, not FIX: §2.2 defines CLOSE as recognisable-with-a-repair, which is a mis-ordered sentence exactly, and red belongs to the learner who picked a word that does not fit. The tile marked is the MINIMUM MOVE, not the position diff.
lesson build say State 5 of 5, and the family's specimen. One card from the first tile to the take. Four tiles regroup as five syllables in the same frame, every bar still on the empty track — phase 1 landing says nothing about the speech. Turns the catalog's weakest game into a ramp into its spine.
lesson build say speaking Room 2 — and Mai listens in it (2026-08-16). The pose carries a verb and the verb is the one the learner is about to need; she is identical whether the build was accepted or mis-ordered, and she stops dead at the armed mic. What it does NOT carry (founder, 2026-08-15): Your turn keeps its question into the speaking room because a reply answers something; this family's room-1 card is a task recap, and the task is over. The meaning is said once, above the plate.
lesson build say known The SAME board, dealt on day 40 — read it beside lesson-build-say-start. Pinyin is a training wheel and it comes off ONE WORD AT A TIME, decided by the atom's own AtomProgress on from-memory's existing threshold. Not a mode: a learner-chosen difficulty makes AtomProgress mean different things for different people, and that number schedules every review in the app. Not per-exercise either: a board mixes words read forty times with the one being taught, and killing pinyin for the board changes what the exercise MEASURES — order, not character recognition. The mixed board is information, not compromise: the one tile still wearing a label is the one with work left. Show pinyin is the escape hatch, free and unrecorded, because the answer here is the ORDER and pinyin is not it.

How a lesson ends — the completion sequence

11

A lesson does not end on a card, it ends on a SEQUENCE: the score reveal, then your own voice, then where that puts you. Every transient beat carries Skip and advances on a tap anywhere; the final surface carries the doors instead. A beat exists only when its number exists — a run with nothing scoreable simply has no score beat. Read the four beat-1 variants together: the win gesture fires on the all-clear run and NOWHERE else, which is what makes it worth anything on the run that earned it. A goal-finishing run swaps beat 3 for the medallion, and the medallion — one viewport, no scroll — hands on to the recap, which is the final surface and where the evidence lives since the 2026-08-03 split.

complete score Beat 1 — the score reveal, on an ordinary mixed run. The sequence is a GRAMMAR, not a tunnel: two or three beats, every transient one carrying Skip and advancing on a tap anywhere. A beat exists only when its number exists, which is why there is no streak beat: the speak-streak is not computed anywhere yet.
complete score clear Beat 1's all-clear variant, and the one score reveal that spends the win gesture — green reaches the characters and a modest spark burst fires once, because every part of every take was judged and landed. A mixed run gets none of it: not a smaller dose, none. A gesture that also fires with a repair left over is worth nothing on the run that earned it.
complete score partial Beat 1 when the scorer could not judge everything. The same three beats at the same weight — and NO success motion, no green event, no chime: “we scored less of your run than you did” is not a win, and dressing it as one is the false green one level up. The uncertainty is named as the machine's, never as the learner's.
complete no score The run where nothing was scoreable — so beat 1 is OMITTED, not drawn empty. A score beat with no score on it is a hole with a screen built around it. Which leaves the voice beat carrying the whole sequence, and that is exactly what it was ruled grade-blind for.
complete voice Beat 2 — the beat no competitor can play at all, and the one that has to work on a bad day. Grade-blind by ruling: no percentage, no verdict word, no verdict colour. The rose panel is now a listening scene with Mai at 190px, cut by its edges — and this ONE template serves a 41% run and a 100% run alike, because nothing on the beat knows the score.
complete voice first The activation beat — the first-ever scored take, which by ruling outranks lesson completion in celebration weight. It is beat 2 of the completion sequence in its other state: the host swaps it in for the ordinary voice beat on isFirstScoredTake, it fires exactly once, and it is never a step in onboarding.
complete progress Beat 3, and the FINAL SURFACE of an ordinary lesson. It claims the lesson, not the goal: the bar arrives holding the value it had before this lesson and then advances. The skip has landed here, so the slot holds a close instead, and the doors and the share affordance ride this screen.
goal complete The medallion beat — and now ONE VIEWPORT with no scroll (split 2026-08-03: the founder walked it and said it takes too much space). Star, claim, 5/5, and Mai standing on the hero's floor at 168px attending a win the coin still delivers. The evidence left this screen entirely.
goal recap The evidence, and the goal sequence's final surface — the other half of the split. Minutes of the learner's own voice, turns with Mei, the five steps walked, the sentences they can now say with the example playable, then Share and the door to goal 2. Mai is deliberately absent: this screen carries measured numbers.
complete share card What leaves the app — the one surface where a number travels without the screen that explains it. So the coverage line travels WITH the percent, 100% may appear only where the run was full-coverage and every part cleared, and nothing rounds. An affordance off the final surface, never a beat in the sequence.
complete rating The rating MOMENT, not a screen. iOS grants roughly three prompt opportunities a year and the dialog cannot be styled, positioned or detected, so the only decision that is actually ours is which state of which surface is worth spending a request on — and the answer is the goal-completion final surface, gated.

Study sets — the second rail

14

Quizlet for spoken Mandarin, down to the paste-your-own wedge. The shelf is open: anyone may publish, an automatic gate decides whether a set is discoverable, and the badge says only that a human read it. So an author is a thing you can search for, and the owner of a set can rename it, drop a word from it or delete it. The word card sits here rather than with the games: it is the atom rendered for reading, and it asks nothing.

sets The My sets chip — the shelf's default state: due, continue, then the whole collection as Made by you / Learning. The chip row is the screen's one organiser; Mai reads at the learner's own shelf head.
sets popular The Popular chip — the library's front shelf, Netflix-shaped in house materials: a spotlight rail of the most-loved certified sets, then the grid. No rank numerals; stars are display-only, and a printed ordinal would be a leaderboard.
sets empty What does nothing look like? The house empty state — browsing Mai, one plain fact, one way forward — on the one screen that can genuinely produce a zero result. Never drawn on an error.
sets search Who made this? The shelf opened to everyone, so the author became a thing you can look for — the person leads the results and their sets follow, matched on the handle as well as the display name, because the handle is the half that is unique.
set detail Is this list worth my time? The set page's BROWSE state — the catalog item and nothing else: category, claim, the average with its count, and the phrases. No progress, no focus, no rating control, because none of them exist until study does.
set detail studying The same page once study begins, and the split is the design: every user-study fact — how far in, what is due, where it picks up, what the focus is set to — in ONE Your-study panel that is simply not drawn before there is a study. Per-phrase status, the trust story behind the badge, and a continue action that names the exact phrase it resumes at.
set manage The owner's own door onto a set they made: rename it, re-describe it, take a word out, take it off the shelf, delete it. A title edit re-enters the gate and a word leaving does not — and a word already learned stays learned.
lesson words The atom study card — NOT an exercise. Everything one word carries, rendered for reading and opened from a study set: hanzi, pinyin, gloss, the Han-Viet reading and the decomposition. The single biggest unfair advantage a Vietnamese learner has.
set paste Can I just paste my own list? The import wedge — and its load-bearing middle step is sense disambiguation: xing vs hang are different words and the system never guesses. A pasted phrase stays one practice target and is never silently exploded.
set add words The weekly ritual, and the screen the class learner actually lives in — this week's handout pasted into a set they already have. Same accounting as the import, plus the one fact it cannot know: what was already here. The duplicate is not an error and says so — “Your progress on it stays.”
due today Eight chapter-sets, one button. The cross-set queue for a learner whose daily time is six-sevenths drilling and one-seventh new words. Due is a fact about the clock, never a grade, so no verdict colour appears anywhere on it.
nothing due The app says you are done. No filler lesson, no medallion — a number to plan around and three real doors, one of which is free and unmetered. Padding an empty queue would be the worst thing we could do with the meter.
set paste large The import moment for a real three-month class list — 488 lines, 476 phrases, 10 already listed and 2 it could not use. You cannot draw 476 rows, so the exceptions get the space and the accepted rows get a count.
set focus 476 phrases is a wall if it opens at word one. Recent 30 / Recent 100 / Everything over the order the class taught them — and focus scopes what is NEW, never what is due, because forgetting does not respect a filter.

Study sets — the games, one screen per name

11

The whole deck a set lesson deals from, one canonical screen per reserved name so the ten can be reviewed at one look. These screens also live with their home groups above — this section is the review contact sheet, not a second copy of the canon. The recognition names deal on NEW words; Catch it, Fill it in, Your turn, Build & say and the picture cue pin onto RETURNING words only (one stored-content exercise per lesson), which is why a brand-new set shows the first half of this row and the second half arrives with the reviews. Say it appears twice deliberately: from memory is the name's study-set centrepiece, and the picture cue is the same name wearing the atom's picture.

lesson match Step 1 of 6. Six tiles, three words, no clock. The densest exercise in the catalog and the only one that cannot be failed, only finished — which is why it opens the first lesson, and why the route's first node points here rather than at a five-syllable sentence. Cleared pairs go ink, never green: nothing here was judged, so nothing claims a verdict.
lesson word check Step 2 of 6. WORD CHECK in the one direction the first lesson runs it: audio → gloss, with no hanzi and no pinyin above the play control. A word check that prints 我 over the options is a reading test wearing an audio button — the job is to hold the SOUND matched thirty seconds ago. Mint, because this is still the library's two minutes.
lesson tone The syllable has played; five contours to choose between, before the mouth ever opens. The character is withheld on purpose — this is perception, not a reading check. All five lines are ink, because a tone is a category and verdict colours are never categories.
lesson listen Step 4 of 6. Can I hear the difference? 是 shì against 四 sì — identical final, identical tone, one initial apart — and it sits at step 4 because step 5 is 是. You cannot produce a distinction you cannot hear, so the learner meets it one screen before they are asked to make it. Listen and choose is one family carrying three names: this, Which tone? and Catch it.
lesson read character READ THE CHARACTER — the seventh family, the character as an object. Dealt on a word the learner has SPOKEN but never read; the four options are the true readings of four words they know, so a wrong pick is the character wearing another word's voice. The object card below gives structure — radical, strokes, parts — the one fact that doesn't leak the sound. The atom speaks at the marking moment (ruled 2026-08-06), never before.
lesson catch it Where does the word land? The third name in the listening family and a genuinely third thing: no competing words at all, just one stream and a timeline. A word you know cold in isolation is a word you still cannot find at speed inside eleven other syllables.
lesson fill FILL IT IN — the tenth name, and the HSK gap-fill shape. A sentence with one word missing, read not heard, four words for the slot. The pinyin line is blanked at the same position, because printing chī under a sentence missing 吃 hands over the answer.
lesson say memory SAY IT · FROM MEMORY — the round's centrepiece, and the one exercise where the learner causes the outcome by retrieving. The cue is the Vietnamese gloss and a picture; pinyin, hanzi and the model audio all wait behind one recorded peek row.
lesson picture A photograph replaces the written gloss, so meaning arrives as a scene instead of a translation. Labelled HSK-style and honest about it: the real spoken exam's picture task is open-answer and we hand you the sentence. The one screen that needs a network.
lesson build say State 5 of 5, and the family's specimen. One card from the first tile to the take. Four tiles regroup as five syllables in the same frame, every bar still on the empty track — phase 1 landing says nothing about the speech. Turns the catalog's weakest game into a ramp into its spine.
lesson your turn Mei asks, three replies, one fits. Picking is step 1 and speaking is step 2, and the screen says so rather than pretending the take has happened. The scaffold is not a softening: open production graded against one target marks a right answer wrong.

Quiet review — the free rail

2

For the moments a learner physically cannot speak: commute, shared room. Review-only, streak-neutral, unmetered — it is NOT a lesson and must not dress like one, so there is no counter, no celebration and no charge. The meter sells speaking, quiet review contains none, and its end state hands straight back to the words waiting for a voice.

Progress & the wall

4

Evidence rather than points, and a paywall that meters the one scarce unit — two lessons a day, spent at the first scored take, unlimited speaking inside each — while never gating a single piece of content.

Teacher tools — the recording sitting

20

The internal shell for the recording study, in Vietnamese. Protect the corpus, forgive the room: what gets stored and how it is labelled fails closed; what goes wrong with a student standing there gets a retry and a sentence. The capture screens are the learner's speaking UI with everything that scores removed — the receipt is the only status they carry.

tool home Where do I go, and who still needs recording? Two doors over a name-first tri-state roster. The state that matters: a student can be 23 of 23 and still not finished, because takes are queued on the phone — and that row must never say saved.
tool setup Set one student up. The app draws the code, the teacher types the name. Start is gated on three answers and the hint names ONE missing thing at a time; the four cohort answers are covariates — recorded, never acted on.
tool setup duplicate Is this the same student? A diacritic-folded near-match, two answers and no default, and nothing is ever merged. A duplicate here splits one person's audio across two codes, permanently and undetectably — so it fails closed.
tool session starting The two or three seconds while the sentence list is fetched and the sitting is registered. Names the two calls instead of spinning, and no elapsed count — a number ticking upward turns an ordinary wait into a thing to watch.
tool unavailable The server would not open the sitting, so nothing may be recorded against it. One retry, one sentence to say to the student, and Chau's name — no offline path, because a take with no sitting to belong to has nowhere to land. Takes already on the phone are safe and keep retrying on their own.
tool preflight Three seconds of the room before anyone speaks. Microphone, route, storage, battery — all measured. Battery and 'couldn't check' are warnings, never blockers.
tool preflight not armed The one blocker nobody in the room can clear: the server refused a test write, so every take of the morning would be lost. No fix list and no dead Start — it says stop, message Chau, and your old recordings are still safe.
tool preflight mic denied A permission to grant. Three taps in Settings, written as three rows, and everything else on the phone is shown passing so nobody goes hunting.
tool preflight mic dead Three seconds of digital silence — the check's whole reason to exist. Another app is holding the mic or the mute switch is on; 3 s here saves 28 dead takes.
tool preflight headset The blocker that sounds fine in the room. Bluetooth drops the take onto an 8000 Hz call line and the detail is gone for good. The only overridable one, and the override is stamped on the sitting so audio is never mislabelled.
tool record prompt The learner's speaking screen with everything that scores removed. The sentence renders completely uncoloured — capture is not scoring, and errors are data, never marked.
tool record arming Wait for the tone. The teacher must not tell a student to speak while the take has not begun.
tool record recording Live capture. One button, no cancel, no per-sentence skip — an aborted take is corpus data, not something a teacher may discard.
tool record saved The receipt, which is what replaces the score here. Green means the bytes landed and says nothing about the speech. Listen back plays what actually went up; then next, or re-record.
tool record retake The take-quality guard re-armed the mic for the same sentence. Soft guidance, never a grade — the 'forgive the room' side of the line: retry, and tell the teacher.
tool record pending The wifi died. The take is safe on the phone and the flow ADVANCES, but the ledger says pending and the sitting cannot close until it lands. Saying 'saved' over audio on a device is the exact permanent, undetectable mislabelling the corpus rule exists to stop.
tool record noise The noise block — the same five sentences read in a deliberately noisy room. Skipping is never silent: it requires a reason, because a skip that leaves no trace is a hole in the data nobody would ever find.
tool record fault The one screen whose answer is STOP. The server refused the take itself — the two things a retry genuinely cannot fix. The control goes down and stays down, dimmed not hidden, so it reads as 'stopped on purpose'; message Chau, and every take so far is safe.
tool session complete Can I let the student go? The question the teacher actually has, as the headline. Completeness is derived from the takes that exist, never from a timestamp.
tool signed out The sign-in ran out, mid-anything. One screen for the recording and rating surfaces both. Not an alert — nothing failed. The load-bearing line: takes spooled on this phone upload only under the account that recorded them, so 'sign in again' means 'as cô Hường'. No self-service account flow; message Chau.

Rating — the ground truth

5

Three external teachers judge recorded audio one syllable at a time, and this is what the scorer is validated against. Independent, blind to what the model thought, and free to abstain — an honest 'cannot judge' is worth far more than a guess. No student identity appears anywhere.

Admin — deliberately small

6

Reads the corpus; writes only people-and-assignment rows. It may not edit or delete a capture, take, participant, session or catalog — destructive corpus operations stay deliberate SQL, because an admin UI that can mutate corpus rows is a corpus-integrity hazard with a comfortable interface on it.

admin records Every teacher's sittings on one list — the read the operator surfaces deliberately refuse. Each row stamps WHICH teacher recorded it; the teacher pills filter, never rank. Newest first, receipts as the only colour, and every row opens the sitting.
admin sitting Did the first student actually work? One sitting's takes in order, with the numbers only a real recording produces, and a play control on every take. Reads the corpus; writes nothing to it.
admin sitting edit Correcting a learner's registration after the fact — the group above all, because the group is what a stratum count is made of. The facts that can never move are shown as rows rather than omitted, and every changed field is audited, said before the button.
admin cohort What is still MISSING. The corpus against the protocol's 24 — A8·B8·C6·Z2 — with the recruiting instruction under each group and C named as the one nothing substitutes for. Counts learners, not sittings, and only on a finished ledger. What blocks a queue build today is filed separately: recruiting more people fixes none of it.
admin grading How the grading round is going: the all-ears count that gates analysis, then one row per slot — progress, cannot-judge tally, last-graded — with the stalled ear named for a phone call. Aggregates only, never a rater's answers. Binding stays on Teachers.
admin teachers Who holds which capacity, and which rating slot. The capacities are an independent SET, not a hierarchy. An unresolved slot is shown explicitly, because that is exactly what makes a queue build refuse to run.

The system

3

The tokens and primitives every screen is built from, and Mai. Rulings live in DESIGN.md.