The design canon — every screen, one page.
Private while it is in review.
Voice Studio — the design canon, every screen, one page.
Read the design ledger →One login for EVERYONE: you sign in, the server returns your role, and the client routes invisibly — an operator lands in the console, everyone else in the app. So the door is warm learner branding with no staff chrome, and there is no separate sign-up screen anywhere.
Home is not just navigation, it is the conversion surface: people open the app and read the route to work out what this thing can eventually do for them. So it runs all the way to the last goal, at one node size, with nothing faded and nothing shrunk.
Ready for the New HSK (3.0) — the approved headline, and these are the screens that license it. The first three are the SECOND ONBOARDING WALK: a learner who says 'exam' at ask 3 gets her own consequence, her own ask 4 and her own receipt, and the chain runs end to end. The specific claim leads everywhere: the new exam makes speaking compulsory from New HSK 3, and its level-3 paper opens with the exercise this app scores every day. Drawn to showcase standard on the founder's ruling that a feature can earn its place as a marketing asset — and honest to the boundary: preparation, never simulation, and never a claim to cover both syllabuses.
Goal 1, step 1: the only lesson a new learner can reach, and the run the first ten minutes ends in. One objective on every screen, one counter from 1 / 6 to 6 / 6, three atoms and no fourth — 我 · 是 · 越南. It ladders into the canonical sentence on purpose: those three atoms are three of its five syllables, and step 4 drills exactly the initial the learner gets wrong on it later. Read these in order — if the run does not walk, nothing else in the gallery matters.
One fixed frame, fifteen moments. Three independent verdicts per syllable — initial, final, tone — and the character reports the worst. The failure states matter as much as the happy one: a false red on a pronunciation app is far more damaging than a false green, and an ambiguous error screen is a false red the learner writes themselves.
Speaking is the only hard modality; every other type is a near-free RENDER of the same atom. SEVEN families carrying ten learner-facing names — adding a variant inside a family costs nothing, adding a family is a second product and needs a decision, not a ticket. A lesson is 6–10 of them and always ends on speaking. Nearly every one of these is recognition — and the exception is drawn here too: Say it from memory, where recall happens before the mouth opens, and the outcome is caused by what the learner retrieves.
The first family drawn as a STATE MACHINE rather than as a specimen, because the states nobody draws are the states nobody rules — and the one this family was missing was the miss. Walked on the device 2026-08-16: a wrong order got no colour, no diff, and no sign of what the learner had actually built, because the marking swapped their tiles out for the correct sentence in the same seat. Read the six in order and the mechanic is one continuous take — the same four words, the same learner, from the deal to the microphone, with state 3's line being exactly the line state 4 marks. Two rulings come out of it: the built line SURVIVES its own marking (§5.8d — the challenge is never spent, and every other family already kept the learner's answer on screen), and a mis-ordered sentence is CLOSE, not FIX — every word right, one has to move, which is §2.2's definition of the colour word for word — though the WORD stays off the screen, because CLEAR/CLOSE/FIX are analyzer vocabulary and nothing here was heard by the scorer. Which is what lets MAI into the room (founder asked, 2026-08-16): with the analyzer words gone, exclusion 1 governs the speaking room and the analyzer, and step 1 finally has the cue room every other family already had. She is on all five build states or none — a companion who left at the marking would be leaving because the learner got it wrong. And room 2 has her LISTENING — the founder pushed on that the same day and was right: the mic is not live on a prompt, it is live one screen later, and speaking-prompt.html and lesson-say-memory.html already seat her on exactly that screen of a speaking exercise. She stops at speaking-arming.html and does not come back. This is the shape every other family owes a pass in.
A lesson does not end on a card, it ends on a SEQUENCE: the score reveal, then your own voice, then where that puts you. Every transient beat carries Skip and advances on a tap anywhere; the final surface carries the doors instead. A beat exists only when its number exists — a run with nothing scoreable simply has no score beat. Read the four beat-1 variants together: the win gesture fires on the all-clear run and NOWHERE else, which is what makes it worth anything on the run that earned it. A goal-finishing run swaps beat 3 for the medallion, and the medallion — one viewport, no scroll — hands on to the recap, which is the final surface and where the evidence lives since the 2026-08-03 split.
Quizlet for spoken Mandarin, down to the paste-your-own wedge. The shelf is open: anyone may publish, an automatic gate decides whether a set is discoverable, and the badge says only that a human read it. So an author is a thing you can search for, and the owner of a set can rename it, drop a word from it or delete it. The word card sits here rather than with the games: it is the atom rendered for reading, and it asks nothing.
The whole deck a set lesson deals from, one canonical screen per reserved name so the ten can be reviewed at one look. These screens also live with their home groups above — this section is the review contact sheet, not a second copy of the canon. The recognition names deal on NEW words; Catch it, Fill it in, Your turn, Build & say and the picture cue pin onto RETURNING words only (one stored-content exercise per lesson), which is why a brand-new set shows the first half of this row and the second half arrives with the reviews. Say it appears twice deliberately: from memory is the name's study-set centrepiece, and the picture cue is the same name wearing the atom's picture.
For the moments a learner physically cannot speak: commute, shared room. Review-only, streak-neutral, unmetered — it is NOT a lesson and must not dress like one, so there is no counter, no celebration and no charge. The meter sells speaking, quiet review contains none, and its end state hands straight back to the words waiting for a voice.
Evidence rather than points, and a paywall that meters the one scarce unit — two lessons a day, spent at the first scored take, unlimited speaking inside each — while never gating a single piece of content.
The internal shell for the recording study, in Vietnamese. Protect the corpus, forgive the room: what gets stored and how it is labelled fails closed; what goes wrong with a student standing there gets a retry and a sentence. The capture screens are the learner's speaking UI with everything that scores removed — the receipt is the only status they carry.
Three external teachers judge recorded audio one syllable at a time, and this is what the scorer is validated against. Independent, blind to what the model thought, and free to abstain — an honest 'cannot judge' is worth far more than a guess. No student identity appears anywhere.
Reads the corpus; writes only people-and-assignment rows. It may not edit or delete a capture, take, participant, session or catalog — destructive corpus operations stay deliberate SQL, because an admin UI that can mutate corpus rows is a corpus-integrity hazard with a comfortable interface on it.