T

MaiSay Design

The design canon — every screen, one page.
Private while it is in review.

Voice Studio · 115 screens

MaiSay

Voice Studio — the design canon, every screen, one page.

Read the design ledger →
Laws every screen obeysone pastel, one jobverdicts stand on paper, never a pastelCLOSE is magenta, not amberfour states — the fourth is not scoredink means unscored, never correctblack is the action colourtext floor 12every verdict carries a wordone state per screen — no hidden states

The front door & first run

11

One login for EVERYONE: you sign in, the server returns your role, and the client routes invisibly — an operator lands in the console, everyone else in the app. So the door is warm learner branding with no staff chrome, and there is no separate sign-up screen anywhere.

login One front door for EVERYONE, and it is the FIRST screen — you sign in, the server returns your role, the client routes silently. No sign-up screen exists: continuing creates the account. ONE TAP LEADS — Apple, Google, Facebook — because a cold door must open in one tap; email falls back below the divider. Glyphs are ink monochrome, and nothing here is black: a front door has no committed action, every path on it is a first tap.
login code The six-digit code. One refusal sentence for every kind of wrong, deliberately — a mistyped email and a guessed staff handle get the identical message, so the screen never tells anyone which accounts exist.
onboarding language Which language do we teach you THROUGH? The one decision that silently rebranches the product — Vietnamese unlocks the Han-Viet scaffold, and the preview is the argument: ~60% of Vietnamese vocabulary is Sino-Vietnamese, so the learner already half-knows thousands of words they have never seen.
onboarding mic Earn the permission before the OS asks for it. A real taste of what the microphone unlocks, then the request. Tap to speak, never hold — and never 'nothing is recorded', because retained audio is what the product is built on.
onboarding take prompt Game 2, the ask. The target, the model, and a mic waiting — onboarding chrome, not lesson chrome, because there is no run to be 1 of yet. 越南 is the target on purpose: the learner's own country, and the word the first lesson comes back to at 100% six exercises later.
onboarding take recording The mic is open and the learner is speaking. The state is said once, on the level meter — the only element that goes still when the mic is dead — and there is NO CLOCK anywhere: duration is a fact about a take you have, not a display while you make one. One button, no cancel.
onboarding first take Their own voice, scored, at minute three. 越南 — the learner's own country — two syllables, five parts, honest at 94% with 南's rise stalled halfway. A rendering of the analyzer is what every competitor shows; this is it running on YOU.
onboarding goal What is driving this, and by when? A reason, not a date picker — and choosing one shows the consequence: a three-week plan front-loads survival.
onboarding hanviet ask Game 1, the moment the learner acts. A play button and four Vietnamese words — and NO hanzi and no pinyin anywhere on screen, because with the writing showing the pick is a reading exercise and the phonetic bridge is an assertion instead of evidence. Non-blocking, and still possible to get wrong: four real options, the answer at C.
onboarding hanviet Game 1, before there is an account: WORD CHECK in its audio → gloss direction. 韩国 is heard, not read — a learner two minutes in cannot read it — and the bridge then says why they knew it anyway: 韩 Hàn, 国 Quốc, one character one syllable. Nothing is scored and nothing can be got wrong: the first thing the app does is GIVE something, not ask.
onboarding band Where are you starting from? Three taps, no placement test — a speaking exam measures pronunciation, which is a weak proxy for the vocabulary band it was setting, and it spent a beginner's first ten minutes on six failures. It sets where you begin, never what you can reach.

Home — the path

4

Home is not just navigation, it is the conversion surface: people open the app and read the route to work out what this thing can eventually do for them. So it runs all the way to the last goal, at one node size, with nothing faded and nothing shrunk.

path day one The same home on day one — the zero state, and deliberately a different screen. Nothing is behind her, so every node is untouched and the trail is dotted the whole way; the card's door reads Start, not Continue. Metrics honestly zero without the shame-shaped emptiness of a screen that reads 'you have done nothing'. The eight goals stay at full weight, because that is what a new user is deciding whether to trust.
path Where am I, and what can this eventually do for me? Day 12: goal 1 is behind her, goal 2 is two steps in. Three node states on one route — cleared, current, ahead — told apart by a joined-up trail, a check, and elevation, with no new hue and no verdict colour. Home is the conversion surface, so the goal card carries the next step as a black action rather than leaving it below the fold, the route still runs all the way to goal 8, and nothing fades or shrinks down the scroll.
path evening The same home after dark — the cadence law drawn rather than described. One thing changes besides the word evening: Mai winds down, lids heavy behind a small yawn. The swap is time-of-day and never activity-conditional, so this screen renders identically for a learner who spoke twelve times today and one who has not opened the app since breakfast. Never-guilt is the whole test of it.
path return Coming back after a long gap — the cadence law's third face, and the screen never-guilt was written for. Mai waves, because a return is an arrival. The broken streak is not printed and not zeroed; no elapsed time is stated anywhere, because '3 weeks away' is the same guilt wearing arithmetic. What takes the slot is the fact that helps: the route is exactly where it was left.

The HSK track — the conversion showcase

2

Ready for the new HSK (3.0) — the approved headline, and these are the screens that license it. The specific claim leads everywhere: the new exam makes speaking compulsory from HSK 3, and its level-3 paper opens with the exercise this app scores every day. Drawn to showcase standard on the founder's ruling that a feature can earn its place as a marketing asset — and honest to the boundary: preparation, never simulation.

The first lesson — one run, six exercises

7

Goal 1, step 1: the only lesson a new learner can reach, and the run the first ten minutes ends in. One objective on every screen, one counter from 1 / 6 to 6 / 6, three atoms and no fourth — 我 · 是 · 越南. It ladders into the canonical sentence on purpose: those three atoms are three of its five syllables, and step 4 drills exactly the initial the learner gets wrong on it later. Read these in order — if the run does not walk, nothing else in the gallery matters.

lesson words meet Từ mới — the meet beat, and NOT an exercise. Every genuinely new atom is given before it is asked for: hanzi, pinyin, gloss, audio, the why-line, the Hán-Việt note where one is authored, and one tap onward. No counter and a progress bar that does not move, because a beat takes no slot in the lesson. The fix for “everything in the app is a test”.
lesson match Step 1 of 6. Six tiles, three words, no clock. The densest exercise in the catalog and the only one that cannot be failed, only finished — which is why it opens the first lesson, and why the route's first node points here rather than at a five-syllable sentence. Cleared pairs go ink, never green: nothing here was judged, so nothing claims a verdict.
lesson word check Step 2 of 6. WORD CHECK in the one direction the first lesson runs it: audio → gloss, with no hanzi and no pinyin above the play control. A word check that prints 我 over the options is a reading test wearing an audio button — the job is to hold the SOUND matched thirty seconds ago. Mint, because this is still the library's two minutes.
lesson say wo Step 3 of 6, and the first time the learner speaks inside a lesson. ATOM GRAIN: 我 is one syllable and two parts, drawn in the same fixed frame the sentence lesson uses — same target card, same analyzer, same mic, at the size the first minute of speaking actually is. The progress bar turns rose here and stays there.
lesson listen Step 4 of 6. Can I hear the difference? 是 shì against 四 sì — identical final, identical tone, one initial apart — and it sits at step 4 because step 5 is 是. You cannot produce a distinction you cannot hear, so the learner meets it one screen before they are asked to make it. Listen and choose is one family carrying three names: this, Which tone? and Catch it.
lesson say shi Step 5 of 6, and where step 4 pays off — without a paragraph saying so. The learner picked the curled tongue out of a minimal pair one screen ago; here the sh span comes back green and the delta chip reads 'sh · held'. And it is NOT all clear: tone 4 stayed short, 83%, one part to repair. The celebration is spent once, at step 6.
lesson say yuenan Step 6 of 6 — the climax, and one of the product's two win moments. The learner said 越南 once before, in onboarding, at 94% with 南's rise stalling halfway; six exercises later the same two syllables come back at 100% and the part that comes good is the exact part that did not. A callback is what makes the celebration affordable here.

The speaking lesson — the moat

15

One fixed frame, fifteen moments. Three independent verdicts per syllable — initial, final, tone — and the character reports the worst. The failure states matter as much as the happy one: a false red on a pronunciation app is far more damaging than a false green, and an ambiguous error screen is a false red the learner writes themselves.

speaking prompt Before you speak. The same frame with no verdict anywhere — bars on the empty track, no score, because a zeroed score is a lie. The screen a learner sees most often.
speaking arming Between the tap and the tone. The control goes inert and says wait — the screen must never claim recording before the cue, because a learner speaking into a mic that has not opened gets a bad score for a take that never happened.
speaking recording Live, and the LEVEL METER is what says so — it goes still when the mic is dead, which a banner cannot. No clock: duration is a fact about a take you have, not a display while you are making one. One button, tap to stop, no cancel anywhere in this loop.
speaking sending The NETWORK half of the wait, named as itself. 'Your voice is going up' is a problem a person can act on; 'the model is thinking' is not — so the screen says which. No fake progress.
speaking scoring The MODEL half of the wait. The card holds, no spinner theatre, and the listen-back is already available because the audio exists on the device.
speaking What exactly did I get wrong, and how do I fix it? The moat: three independent verdicts per syllable, the character reporting the worst, and one named physical repair.
speaking tone The SECOND repair, and the reason the first screen only names it. 是 has landed; now the rising tone on 南 stalled halfway — and a stalled tone is a picture, not a sentence, so the target-vs-you contour is the whole screen. One change at a time, applied to layout.
speaking retry The win moment. The retry REPLACED the first result in place — no Take 1 / Take 2 list, the old take survives only as a delta line. Every part judged and clear, so this is the one screen where the characters themselves go green.
speaking skip The other branch: the repair pass came back red on the same part, and Skip appears beside it. A red take never walls the lesson (ruled 2026-07-31) — but the black stays on 'Say it again', nothing turns green, and the panel says why skipping is honest: 是 stays marked as not landed and review brings it back within a day or two.
speaking uncertain A part could not be scored. It renders as ABSENCE — ink mark, bar left on the empty track — and the headline withholds its credit: 92%, 12 of 13 parts scored. This screen exists to prove the fourth state is visible. No celebration on a partial.
speaking below floor Below the coverage floor, so there is NO score at all. The card stays pre-speak with an honest note, the copy blames the conditions, and the retry is free.
speaking mic denied The microphone is off at the OS. The record control wears the state and carries Open Settings — a settings problem, not a failure, and never a dark error screen.
speaking no speech Nothing was heard. Blames the room, the distance, the timing — never the learner. This is where trust is won or lost, because the learner's instinct is to assume they did something wrong.
speaking offline Scoring is unreachable. The take is saved on the device and nothing is lost; the practice loop stays warm. No modal, no red, no dead end.
speaking long Twenty-two characters — the stress test. Tokens wrap in reading order, the translation collapses first, pinyin and hanzi never shrink below production scale, and the score never hides. It scrolls, and that is correct.

The other exercise types

22

Speaking is the only hard modality; every other type is a near-free RENDER of the same atom. SEVEN families carrying ten learner-facing names — adding a variant inside a family costs nothing, adding a family is a second product and needs a decision, not a ticket. A lesson is 6–10 of them and always ends on speaking. Nearly every one of these is recognition — and the exception is drawn here too: Say it from memory, where recall happens before the mouth opens, and the outcome is caused by what the learner retrieves.

choice Which reply belongs next? Listen without subtitles, then choose — with the hint behind a deliberate tap, so the ear gets its chance first.
choice wrong The miss. The tapped answer is a fluent sentence answering a DIFFERENT question, and the explanation names that — a comprehension miss, so no analyzer, no bars, no percent.
lesson meaning What does this word mean? WORD CHECK in its hanzi-to-gloss direction — the cue is the characters and nothing else, because hanzi plus pinyin plus audio before the pick hands you three routes to an answer the exercise tests one route to.
lesson recognize Which character is it? WORD CHECK, audio to hanzi — the ear is the whole cue, so the spelling is withheld until the answer. The fix is the stroke difference DRAWN, not described. Recognition is first-class; producing strokes is explicitly out of scope.
lesson word check pinyin How is this one said? WORD CHECK, hanzi to pinyin — the direction scheduled immediately before a Say it, so every distractor is a mistake the learner is about to make out loud. One component, eight directions, three drawings of it.
lesson tone The syllable has played; five contours to choose between, before the mouth ever opens. The character is withheld on purpose — this is perception, not a reading check. All five lines are ink, because a tone is a category and verdict colours are never categories.
lesson catch it Where does the word land? The third name in the listening family and a genuinely third thing: no competing words at all, just one stream and a timeline. A word you know cold in isolation is a word you still cannot find at speed inside eleven other syllables.
lesson catch scene CATCH IT at dialogue grain — two turns of a scene, heard once, then a question about what was actually said. The title is contextual (What did Mei order?), not an eleventh name; no word of the exchange is printed, because the transcript IS the answer.
lesson fill FILL IT IN — the tenth name, and the HSK gap-fill shape. A sentence with one word missing, read not heard, four words for the slot. The pinyin line is blanked at the same position, because printing chī under a sentence missing 吃 hands over the answer.
lesson fill wrong The wrong word went into the gap. The marked state names the confusion — which verb takes 面条 — with the collocation drawn, not lectured. The first drawn distractor set failed the one-right-answer gate and the review round caught it; the fix is in the head comment as a worked example of why cloze sets are reviewed content.
lesson build say One card from the first tile to the take. Four tiles regroup as five syllables in the same frame, every bar still on the empty track — phase 1 landing says nothing about the speech. Turns the catalog's weakest game into a ramp into its spine.
lesson your turn Mei asks, three replies, one fits. Picking is step 1 and speaking is step 2, and the screen says so rather than pretending the take has happened. The scaffold is not a softening: open production graded against one target marks a right answer wrong.
lesson your turn correct The same screen, marked — recoloured and not one pixel moved (§5.8d). The learner picked a fluent reply to a DIFFERENT question; the teaching arrives BELOW the options, and the fitting reply — never their own wrong pick — goes to the next screen.
lesson your turn speaking Step 2, and it is a screen of its own (founder, 2026-08-02). The pick screen closed with a continue; here the question stays fully readable, the replies are gone, and the fitting reply stands in the speaking card. Nobody scrolls to find a microphone.
lesson picture A photograph replaces the written gloss, so meaning arrives as a scene instead of a translation. Labelled HSK-style and honest about it: the real spoken exam's picture task is open-answer and we hand you the sentence. The one screen that needs a network.
lesson say memory SAY IT · FROM MEMORY — the round's centrepiece, and the one exercise where the learner causes the outcome by retrieving. The cue is the Vietnamese gloss and a picture; pinyin, hanzi and the model audio all wait behind one recorded peek row.
lesson say memory peek The learner opened the pinyin. The answer is on screen, the take is marked Peeked — a state label in the SRS pill's own slot, not an apology — and the screen says nothing in its defence: a peeked take scores pronunciation, and the numbers do the rest.
lesson say memory mismatch A different word came out — and nothing turns red. The take fails part alignment, lands below the coverage floor, and the screen says so honestly: it may have been perfectly good Chinese, it just wasn't this word. NOT SCORED where a percent would sit, a free retry, and the peek row waiting. A false red here is the exact trust-killer the amber doctrine exists to prevent.
conversation A known reply inside a real scene. Scripted, because the scorer needs a known target — your line arrives as a prompt in your own bubble, then gains its verdict in place.
conversation known The same scene met again, and the only screen that draws the gloss fade. A line withdraws its Vietnamese once every atom in it sits at SRS ladder step ≥ 2; a line holding a never-met atom never fades, so stimulus stays comprehensible; and a tap brings any of them back — recorded, costless. The learner’s own reply keeps its pinyin, which is a speaking aid and not a gloss.
conversation recording The take, without leaving the scene. A conversation turn records in the conversation's own chrome — transcript behind, Mei's question still readable, the draft bubble unchanged — because handing the learner to the shared speaking screen mid-scene means step 2 of the goal wearing step 3's counter. Turn badge, no lesson counter, no clock.
conversation feedback How did that turn go — without losing the conversation? A BOTTOM SHEET over the transcript, because a conversation has several scored turns and a lesson exercise has one: the turn stays visible behind it, dismissing lands you exactly where you were, and the next one is a tap rather than two navigations. The signature target-vs-you contour, shown open rather than hidden behind a tap.

How a lesson ends — the completion sequence

11

A lesson does not end on a card, it ends on a SEQUENCE: the score reveal, then your own voice, then where that puts you. Every transient beat carries Skip and advances on a tap anywhere; the final surface carries the doors instead. A beat exists only when its number exists — a run with nothing scoreable simply has no score beat. Read the four beat-1 variants together: the win gesture fires on the all-clear run and NOWHERE else, which is what makes it worth anything on the run that earned it. A goal-finishing run swaps beat 3 for the medallion, and the medallion — one viewport, no scroll — hands on to the recap, which is the final surface and where the evidence lives since the 2026-08-03 split.

complete score Beat 1 — the score reveal, on an ordinary mixed run. The sequence is a GRAMMAR, not a tunnel: two or three beats, every transient one carrying Skip and advancing on a tap anywhere. A beat exists only when its number exists, which is why there is no streak beat: the speak-streak is not computed anywhere yet.
complete score clear Beat 1's all-clear variant, and the one score reveal that spends the win gesture — green reaches the characters and a modest spark burst fires once, because every part of every take was judged and landed. A mixed run gets none of it: not a smaller dose, none. A gesture that also fires with a repair left over is worth nothing on the run that earned it.
complete score partial Beat 1 when the scorer could not judge everything. The same three beats at the same weight — and NO success motion, no green event, no chime: “we scored less of your run than you did” is not a win, and dressing it as one is the false green one level up. The uncertainty is named as the machine's, never as the learner's.
complete no score The run where nothing was scoreable — so beat 1 is OMITTED, not drawn empty. A score beat with no score on it is a hole with a screen built around it. Which leaves the voice beat carrying the whole sequence, and that is exactly what it was ruled grade-blind for.
complete voice Beat 2 — the beat no competitor can play at all, and the one that has to work on a bad day. Grade-blind by ruling: no percentage, no verdict word, no verdict colour. The rose panel is now a listening scene with Mai at 190px, cut by its edges — and this ONE template serves a 41% run and a 100% run alike, because nothing on the beat knows the score.
complete voice first The activation beat — the first-ever scored take, which by ruling outranks lesson completion in celebration weight. It replaces the ordinary voice beat and fires exactly once, wherever that first take happened.
complete progress Beat 3, and the FINAL SURFACE of an ordinary lesson. It claims the lesson, not the goal: the bar arrives holding the value it had before this lesson and then advances. The skip has landed here, so the slot holds a close instead, and the doors and the share affordance ride this screen.
goal complete The medallion beat — and now ONE VIEWPORT with no scroll (split 2026-08-03: the founder walked it and said it takes too much space). Star, claim, 5/5, and Mai standing on the hero's floor at 168px attending a win the coin still delivers. The evidence left this screen entirely.
goal recap The evidence, and the goal sequence's final surface — the other half of the split. Minutes of the learner's own voice, turns with Mei, the five steps walked, the sentences they can now say with the example playable, then Share and the door to goal 2. Mai is deliberately absent: this screen carries measured numbers.
complete share card What leaves the app — the one surface where a number travels without the screen that explains it. So the coverage line travels WITH the percent, 100% may appear only where the run was full-coverage and every part cleared, and nothing rounds. An affordance off the final surface, never a beat in the sequence.
complete rating The rating MOMENT, not a screen. iOS grants roughly three prompt opportunities a year and the dialog cannot be styled, positioned or detected, so the only decision that is actually ours is which state of which surface is worth spending a request on — and the answer is the goal-completion final surface, gated.

Study sets — the second rail

5

Quizlet for spoken Mandarin, down to the paste-your-own wedge. Discovery shows official and verified sets only; user sets are shareable by link and never discoverable. The word card sits here rather than with the games: it is the atom rendered for reading, and it asks nothing.

Quiet review — the free rail

2

For the moments a learner physically cannot speak: commute, shared room. Review-only, streak-neutral, unmetered — it is NOT a lesson and must not dress like one, so there is no counter, no celebration and no charge. The meter sells speaking, quiet review contains none, and its end state hands straight back to the words waiting for a voice.

Progress & the wall

4

Evidence rather than points, and a paywall that meters the one scarce unit — two lessons a day, spent at the first scored take, unlimited speaking inside each — while never gating a single piece of content.

Teacher tools — the recording sitting

20

The internal shell for the recording study, in Vietnamese. Protect the corpus, forgive the room: what gets stored and how it is labelled fails closed; what goes wrong with a student standing there gets a retry and a sentence. The capture screens are the learner's speaking UI with everything that scores removed — the receipt is the only status they carry.

tool home Where do I go, and who still needs recording? Two doors over a name-first tri-state roster. The state that matters: a student can be 23 of 23 and still not finished, because takes are queued on the phone — and that row must never say saved.
tool setup Set one student up. The app draws the code, the teacher types the name. Start is gated on three answers and the hint names ONE missing thing at a time; the four cohort answers are covariates — recorded, never acted on.
tool setup duplicate Is this the same student? A diacritic-folded near-match, two answers and no default, and nothing is ever merged. A duplicate here splits one person's audio across two codes, permanently and undetectably — so it fails closed.
tool session starting The two or three seconds while the sentence list is fetched and the sitting is registered. Names the two calls instead of spinning, and no elapsed count — a number ticking upward turns an ordinary wait into a thing to watch.
tool unavailable The server would not open the sitting, so nothing may be recorded against it. One retry, one sentence to say to the student, and Chau's name — no offline path, because a take with no sitting to belong to has nowhere to land. Takes already on the phone are safe and keep retrying on their own.
tool preflight Three seconds of the room before anyone speaks. Microphone, route, storage, battery — all measured. Battery and 'couldn't check' are warnings, never blockers.
tool preflight not armed The one blocker nobody in the room can clear: the server refused a test write, so every take of the morning would be lost. No fix list and no dead Start — it says stop, message Chau, and your old recordings are still safe.
tool preflight mic denied A permission to grant. Three taps in Settings, written as three rows, and everything else on the phone is shown passing so nobody goes hunting.
tool preflight mic dead Three seconds of digital silence — the check's whole reason to exist. Another app is holding the mic or the mute switch is on; 3 s here saves 28 dead takes.
tool preflight headset The blocker that sounds fine in the room. Bluetooth drops the take onto an 8000 Hz call line and the detail is gone for good. The only overridable one, and the override is stamped on the sitting so audio is never mislabelled.
tool record prompt The learner's speaking screen with everything that scores removed. The sentence renders completely uncoloured — capture is not scoring, and errors are data, never marked.
tool record arming Wait for the tone. The teacher must not tell a student to speak while the take has not begun.
tool record recording Live capture. One button, no cancel, no per-sentence skip — an aborted take is corpus data, not something a teacher may discard.
tool record saved The receipt, which is what replaces the score here. Green means the bytes landed and says nothing about the speech. Listen back plays what actually went up; then next, or re-record.
tool record retake The take-quality guard re-armed the mic for the same sentence. Soft guidance, never a grade — the 'forgive the room' side of the line: retry, and tell the teacher.
tool record pending The wifi died. The take is safe on the phone and the flow ADVANCES, but the ledger says pending and the sitting cannot close until it lands. Saying 'saved' over audio on a device is the exact permanent, undetectable mislabelling the corpus rule exists to stop.
tool record noise The noise block — the same five sentences read in a deliberately noisy room. Skipping is never silent: it requires a reason, because a skip that leaves no trace is a hole in the data nobody would ever find.
tool record fault The one screen whose answer is STOP. The server refused the take itself — the two things a retry genuinely cannot fix. The control goes down and stays down, dimmed not hidden, so it reads as 'stopped on purpose'; message Chau, and every take so far is safe.
tool session complete Can I let the student go? The question the teacher actually has, as the headline. Completeness is derived from the takes that exist, never from a timestamp.
tool signed out The sign-in ran out, mid-anything. One screen for the recording and rating surfaces both. Not an alert — nothing failed. The load-bearing line: takes spooled on this phone upload only under the account that recorded them, so 'sign in again' means 'as cô Hường'. No self-service account flow; message Chau.

Rating — the ground truth

5

Three external teachers judge recorded audio one syllable at a time, and this is what the scorer is validated against. Independent, blind to what the model thought, and free to abstain — an honest 'cannot judge' is worth far more than a guess. No student identity appears anywhere.

Admin — deliberately small

2

Reads the corpus; writes only people-and-assignment rows. It may not edit or delete a capture, take, participant, session or catalog — destructive corpus operations stay deliberate SQL, because an admin UI that can mutate corpus rows is a corpus-integrity hazard with a comfortable interface on it.

The system

3

The tokens and primitives every screen is built from, and Mai. Rulings live in DESIGN.md.

More

2

Not yet grouped — add it to APP_GROUPS/WEB_GROUPS in scripts/build-gallery.py.