The design canon — every screen, one page.
Private while it is in review.
Voice Studio — the design canon, every screen, one page.
Read the design ledger →One login for EVERYONE: you sign in, the server returns your role, and the client routes invisibly — an operator lands in the console, everyone else in the app. So the door is warm learner branding with no staff chrome, and there is no separate sign-up screen anywhere.
Home is not just navigation, it is the conversion surface: people open the app and read the route to work out what this thing can eventually do for them. So it runs all the way to the last goal, at one node size, with nothing faded and nothing shrunk.
Ready for the new HSK (3.0) — the approved headline, and these are the screens that license it. The specific claim leads everywhere: the new exam makes speaking compulsory from HSK 3, and its level-3 paper opens with the exercise this app scores every day. Drawn to showcase standard on the founder's ruling that a feature can earn its place as a marketing asset — and honest to the boundary: preparation, never simulation.
Goal 1, step 1: the only lesson a new learner can reach, and the run the first ten minutes ends in. One objective on every screen, one counter from 1 / 6 to 6 / 6, three atoms and no fourth — 我 · 是 · 越南. It ladders into the canonical sentence on purpose: those three atoms are three of its five syllables, and step 4 drills exactly the initial the learner gets wrong on it later. Read these in order — if the run does not walk, nothing else in the gallery matters.
One fixed frame, fifteen moments. Three independent verdicts per syllable — initial, final, tone — and the character reports the worst. The failure states matter as much as the happy one: a false red on a pronunciation app is far more damaging than a false green, and an ambiguous error screen is a false red the learner writes themselves.
Speaking is the only hard modality; every other type is a near-free RENDER of the same atom. SEVEN families carrying ten learner-facing names — adding a variant inside a family costs nothing, adding a family is a second product and needs a decision, not a ticket. A lesson is 6–10 of them and always ends on speaking. Nearly every one of these is recognition — and the exception is drawn here too: Say it from memory, where recall happens before the mouth opens, and the outcome is caused by what the learner retrieves.
A lesson does not end on a card, it ends on a SEQUENCE: the score reveal, then your own voice, then where that puts you. Every transient beat carries Skip and advances on a tap anywhere; the final surface carries the doors instead. A beat exists only when its number exists — a run with nothing scoreable simply has no score beat. Read the four beat-1 variants together: the win gesture fires on the all-clear run and NOWHERE else, which is what makes it worth anything on the run that earned it. A goal-finishing run swaps beat 3 for the medallion, and the medallion — one viewport, no scroll — hands on to the recap, which is the final surface and where the evidence lives since the 2026-08-03 split.
Quizlet for spoken Mandarin, down to the paste-your-own wedge. Discovery shows official and verified sets only; user sets are shareable by link and never discoverable. The word card sits here rather than with the games: it is the atom rendered for reading, and it asks nothing.
For the moments a learner physically cannot speak: commute, shared room. Review-only, streak-neutral, unmetered — it is NOT a lesson and must not dress like one, so there is no counter, no celebration and no charge. The meter sells speaking, quiet review contains none, and its end state hands straight back to the words waiting for a voice.
Evidence rather than points, and a paywall that meters the one scarce unit — two lessons a day, spent at the first scored take, unlimited speaking inside each — while never gating a single piece of content.
The internal shell for the recording study, in Vietnamese. Protect the corpus, forgive the room: what gets stored and how it is labelled fails closed; what goes wrong with a student standing there gets a retry and a sentence. The capture screens are the learner's speaking UI with everything that scores removed — the receipt is the only status they carry.
Three external teachers judge recorded audio one syllable at a time, and this is what the scorer is validated against. Independent, blind to what the model thought, and free to abstain — an honest 'cannot judge' is worth far more than a guess. No student identity appears anywhere.
Reads the corpus; writes only people-and-assignment rows. It may not edit or delete a capture, take, participant, session or catalog — destructive corpus operations stay deliberate SQL, because an admin UI that can mutate corpus rows is a corpus-integrity hazard with a comfortable interface on it.
The tokens and primitives every screen is built from, and Mai. Rulings live in DESIGN.md.