T

Toni Speak Design

The design canon — every screen, one page.
Private while it is in review.

Voice Studio · 86 screens

Toni Speak

Voice Studio — the design canon, every screen, one page.

Read the design ledger →
Laws every screen obeysone pastel, one jobverdicts stand on paper, never a pastelCLOSE is magenta, not amberfour states — the fourth is not scoredink means unscored, never correctblack is the action colourtext floor 12every verdict carries a wordone state per screen — no hidden states

The front door & first run

11

One login for EVERYONE: you sign in, the server returns your role, and the client routes invisibly — an operator lands in the console, everyone else in the app. So the door is warm learner branding with no staff chrome, and there is no separate sign-up screen anywhere.

onboarding language Which language do we teach you THROUGH? The one decision that silently rebranches the product — Vietnamese unlocks the Han-Viet scaffold, and the preview is the argument: ~60% of Vietnamese vocabulary is Sino-Vietnamese, so the learner already half-knows thousands of words they have never seen.
onboarding goal What is driving this, and by when? A reason, not a date picker — and choosing one shows the consequence: a three-week plan front-loads survival.
onboarding hanviet ask Game 1, the moment the learner acts. A play button and four Vietnamese words — and NO hanzi and no pinyin anywhere on screen, because with the writing showing the pick is a reading exercise and the phonetic bridge is an assertion instead of evidence. Non-blocking, and still possible to get wrong: four real options, the answer at C.
onboarding hanviet Game 1, before there is an account: WORD CHECK in its audio → gloss direction. 韩国 is heard, not read — a learner two minutes in cannot read it — and the bridge then says why they knew it anyway: 韩 Hàn, 国 Quốc, one character one syllable. Nothing is scored and nothing can be got wrong: the first thing the app does is GIVE something, not ask.
login One front door for EVERYONE — you sign in, the server returns your role, the client routes silently. No sign-up screen exists: continuing creates the account. ONE TAP LEADS — Apple, Google, Facebook — because that is what let login move behind the first game; email falls back below the divider. Glyphs are ink monochrome, and nothing here is black: a front door has no committed action, every path on it is a first tap.
login code The six-digit code. One refusal sentence for every kind of wrong, deliberately — a mistyped email and a guessed staff handle get the identical message, so the screen never tells anyone which accounts exist.
onboarding mic Earn the permission before the OS asks for it. A real taste of what the microphone unlocks, then the request. Tap to speak, never hold — and never 'nothing is recorded', because retained audio is what the product is built on.
onboarding take prompt Game 2, the ask. The target, the model, and a mic waiting — onboarding chrome, not lesson chrome, because there is no run to be 1 of yet. 越南 is the target on purpose: the learner's own country, and the word the first lesson comes back to at 100% six exercises later.
onboarding take recording The mic is open and the learner is speaking. The state is said once, on the level meter — the only element that goes still when the mic is dead — and there is NO CLOCK anywhere: duration is a fact about a take you have, not a display while you make one. One button, no cancel.
onboarding first take Their own voice, scored, at minute three. 越南 — the learner's own country — two syllables, five parts, honest at 94% with 南's rise stalled halfway. A rendering of the analyzer is what every competitor shows; this is it running on YOU.
onboarding band Where are you starting from? Three taps, no placement test — a speaking exam measures pronunciation, which is a weak proxy for the vocabulary band it was setting, and it spent a beginner's first ten minutes on six failures. It sets where you begin, never what you can reach.

Home — the path

2

Home is not just navigation, it is the conversion surface: people open the app and read the route to work out what this thing can eventually do for them. So it runs all the way to the last goal, at one node size, with nothing faded and nothing shrunk.

The first lesson — one run, six exercises

7

Goal 1, step 1: the only lesson a new learner can reach, and the run the first ten minutes ends in. One objective on every screen, one counter from 1 / 6 to 6 / 6, three atoms and no fourth — 我 · 是 · 越南. It ladders into the canonical sentence on purpose: those three atoms are three of its five syllables, and step 4 drills exactly the initial the learner gets wrong on it later. Read these in order — if the run does not walk, nothing else in the gallery matters.

lesson match Step 1 of 6. Six tiles, three words, no clock. The densest exercise in the catalog and the only one that cannot be failed, only finished — which is why it opens the first lesson, and why the route's first node points here rather than at a five-syllable sentence. Cleared pairs go ink, never green: nothing here was judged, so nothing claims a verdict.
lesson word check Step 2 of 6. WORD CHECK in the one direction the first lesson runs it: audio → gloss, with no hanzi and no pinyin above the play control. A word check that prints 我 over the options is a reading test wearing an audio button — the job is to hold the SOUND matched thirty seconds ago. Mint, because this is still the library's two minutes.
lesson say wo Step 3 of 6, and the first time the learner speaks inside a lesson. ATOM GRAIN: 我 is one syllable and two parts, drawn in the same fixed frame the sentence lesson uses — same target card, same analyzer, same mic, at the size the first minute of speaking actually is. The progress bar turns rose here and stays there.
lesson listen Step 4 of 6. Can I hear the difference? 是 shì against 四 sì — identical final, identical tone, one initial apart — and it sits at step 4 because step 5 is 是. You cannot produce a distinction you cannot hear, so the learner meets it one screen before they are asked to make it. Listen and choose is one family carrying three names: this, Which tone? and Catch it.
lesson say shi Step 5 of 6, and where step 4 pays off — without a paragraph saying so. The learner picked the curled tongue out of a minimal pair one screen ago; here the sh span comes back green and the delta chip reads 'sh · held'. And it is NOT all clear: tone 4 stayed short, 83%, one part to repair. The celebration is spent once, at step 6.
lesson say yuenan Step 6 of 6 — the climax, and one of the product's two win moments. The learner said 越南 once before, in onboarding, at 94% with 南's rise stalling halfway; six exercises later the same two syllables come back at 100% and the part that comes good is the exact part that did not. A callback is what makes the celebration affordable here.
lesson complete Did I actually get better? Not a points tally — proof. The sentence that was not clean and now is, the parts cleared, the minutes of your own voice.

The speaking lesson — the moat

14

One fixed frame, thirteen moments. Three independent verdicts per syllable — initial, final, tone — and the character reports the worst. The failure states matter as much as the happy one: a false red on a pronunciation app is far more damaging than a false green, and an ambiguous error screen is a false red the learner writes themselves.

speaking prompt Before you speak. The same frame with no verdict anywhere — bars on the empty track, no score, because a zeroed score is a lie. The screen a learner sees most often.
speaking arming Between the tap and the tone. The control goes inert and says wait — the screen must never claim recording before the cue, because a learner speaking into a mic that has not opened gets a bad score for a take that never happened.
speaking recording Live, and the LEVEL METER is what says so — it goes still when the mic is dead, which a banner cannot. No clock: duration is a fact about a take you have, not a display while you are making one. One button, tap to stop, no cancel anywhere in this loop.
speaking sending The NETWORK half of the wait, named as itself. 'Your voice is going up' is a problem a person can act on; 'the model is thinking' is not — so the screen says which. No fake progress.
speaking scoring The MODEL half of the wait. The card holds, no spinner theatre, and the listen-back is already available because the audio exists on the device.
speaking What exactly did I get wrong, and how do I fix it? The moat: three independent verdicts per syllable, the character reporting the worst, and one named physical repair.
speaking tone The SECOND repair, and the reason the first screen only names it. 是 has landed; now the rising tone on 南 stalled halfway — and a stalled tone is a picture, not a sentence, so the target-vs-you contour is the whole screen. One change at a time, applied to layout.
speaking retry The win moment. The retry REPLACED the first result in place — no Take 1 / Take 2 list, the old take survives only as a delta line. Every part judged and clear, so this is the one screen where the characters themselves go green.
speaking uncertain A part could not be scored. It renders as ABSENCE — ink mark, bar left on the empty track — and the headline withholds its credit: 92%, 12 of 13 parts scored. This screen exists to prove the fourth state is visible. No celebration on a partial.
speaking below floor Below the coverage floor, so there is NO score at all. The card stays pre-speak with an honest note, the copy blames the conditions, and the retry is free.
speaking mic denied The microphone is off at the OS. The record control wears the state and carries Open Settings — a settings problem, not a failure, and never a dark error screen.
speaking no speech Nothing was heard. Blames the room, the distance, the timing — never the learner. This is where trust is won or lost, because the learner's instinct is to assume they did something wrong.
speaking offline Scoring is unreachable. The take is saved on the device and nothing is lost; the practice loop stays warm. No modal, no red, no dead end.
speaking long Twenty-two characters — the stress test. Tokens wrap in reading order, the translation collapses first, pinyin and hanzi never shrink below production scale, and the score never hides. It scrolls, and that is correct.

The other exercise types

14

Speaking is the only hard modality; every other type is a near-free RENDER of the same atom. SEVEN families carrying nine learner-facing names — adding a variant inside a family costs nothing, adding a family is a second product and needs a decision, not a ticket. A lesson is 6–10 of them and always ends on speaking. And every one of these is recognition: the only true RECALL exercise in the product is the one where you open your mouth.

choice Which reply belongs next? Listen without subtitles, then choose — with the hint behind a deliberate tap, so the ear gets its chance first.
choice wrong The miss. The tapped answer is a fluent sentence answering a DIFFERENT question, and the explanation names that — a comprehension miss, so no analyzer, no bars, no percent.
lesson meaning What does this word mean? WORD CHECK in its hanzi-to-gloss direction — the cue is the characters and nothing else, because hanzi plus pinyin plus audio before the pick hands you three routes to an answer the exercise tests one route to.
lesson recognize Which character is it? WORD CHECK, audio to hanzi — the ear is the whole cue, so the spelling is withheld until the answer. The fix is the stroke difference DRAWN, not described. Recognition is first-class; producing strokes is explicitly out of scope.
lesson word check pinyin How is this one said? WORD CHECK, hanzi to pinyin — the direction scheduled immediately before a Say it, so every distractor is a mistake the learner is about to make out loud. One component, eight directions, three drawings of it.
lesson tone The syllable has played; five contours to choose between, before the mouth ever opens. The character is withheld on purpose — this is perception, not a reading check. All five lines are ink, because a tone is a category and verdict colours are never categories.
lesson catch it Where does the word land? The third name in the listening family and a genuinely third thing: no competing words at all, just one stream and a timeline. A word you know cold in isolation is a word you still cannot find at speed inside eleven other syllables.
lesson build say One card from the first tile to the take. Four tiles regroup as five syllables in the same frame, every bar still on the empty track — phase 1 landing says nothing about the speech. Turns the catalog's weakest game into a ramp into its spine.
lesson your turn Mei asks, three replies, one fits. Picking is step 1 and speaking is step 2, and the screen says so rather than pretending the take has happened. The scaffold is not a softening: open production graded against one target marks a right answer wrong.
lesson your turn correct The ruling, drawn. The learner picked a fluent reply to a DIFFERENT question — and the microphone is handed the canonical fitting reply, never their own wrong pick, because rehearsing an inappropriate answer aloud is the failure this screen prevents.
lesson picture A photograph replaces the written gloss, so meaning arrives as a scene instead of a translation. Labelled HSK-style and honest about it: the real spoken exam's picture task is open-answer and we hand you the sentence. The one screen that needs a network.
conversation A known reply inside a real scene. Scripted, because the scorer needs a known target — your line arrives as a prompt in your own bubble, then gains its verdict in place.
conversation recording The take, without leaving the scene. A conversation turn records in the conversation's own chrome — transcript behind, Mei's question still readable, the draft bubble unchanged — because handing the learner to the shared speaking screen mid-scene means step 2 of the goal wearing step 3's counter. Turn badge, no lesson counter, no clock.
conversation feedback How did that turn go — without losing the conversation? A BOTTOM SHEET over the transcript, because a conversation has several scored turns and a lesson exercise has one: the turn stays visible behind it, dismissing lands you exactly where you were, and the next one is a tap rather than two navigations. The signature target-vs-you contour, shown open rather than hidden behind a tap.

Goal complete — the second win moment

1

A practice room that never celebrates reads flat. The celebration is motion and scale plus green reaching the characters, spent at the all-clear take (step 6 of the run above) and at goal completion, and NOWHERE else. One real moment beats eight decorated ones.

Study sets — the second rail

4

Quizlet for spoken Mandarin, down to the paste-your-own wedge. Discovery shows official and verified sets only; user sets are shareable by link and never discoverable. The word card sits here rather than with the games: it is the atom rendered for reading, and it asks nothing.

Progress & the wall

4

Evidence rather than points, and a paywall that meters the one scarce unit — two lessons a day, spent at the first scored take, unlimited speaking inside each — while never gating a single piece of content.

Teacher tools — the recording sitting

20

The internal shell for the recording study, in Vietnamese. Protect the corpus, forgive the room: what gets stored and how it is labelled fails closed; what goes wrong with a student standing there gets a retry and a sentence. The capture screens are the learner's speaking UI with everything that scores removed — the receipt is the only status they carry.

tool home Where do I go, and who still needs recording? Two doors over a name-first tri-state roster. The state that matters: a student can be 23 of 23 and still not finished, because takes are queued on the phone — and that row must never say saved.
tool setup Set one student up. The app draws the code, the teacher types the name. Start is gated on three answers and the hint names ONE missing thing at a time; the four cohort answers are covariates — recorded, never acted on.
tool setup duplicate Is this the same student? A diacritic-folded near-match, two answers and no default, and nothing is ever merged. A duplicate here splits one person's audio across two codes, permanently and undetectably — so it fails closed.
tool session starting The two or three seconds while the sentence list is fetched and the sitting is registered. Names the two calls instead of spinning, and no elapsed count — a number ticking upward turns an ordinary wait into a thing to watch.
tool unavailable The server would not open the sitting, so nothing may be recorded against it. One retry, one sentence to say to the student, and Chau's name — no offline path, because a take with no sitting to belong to has nowhere to land. Takes already on the phone are safe and keep retrying on their own.
tool preflight Three seconds of the room before anyone speaks. Microphone, route, storage, battery — all measured. Battery and 'couldn't check' are warnings, never blockers.
tool preflight not armed The one blocker nobody in the room can clear: the server refused a test write, so every take of the morning would be lost. No fix list and no dead Start — it says stop, message Chau, and your old recordings are still safe.
tool preflight mic denied A permission to grant. Three taps in Settings, written as three rows, and everything else on the phone is shown passing so nobody goes hunting.
tool preflight mic dead Three seconds of digital silence — the check's whole reason to exist. Another app is holding the mic or the mute switch is on; 3 s here saves 28 dead takes.
tool preflight headset The blocker that sounds fine in the room. Bluetooth drops the take onto an 8000 Hz call line and the detail is gone for good. The only overridable one, and the override is stamped on the sitting so audio is never mislabelled.
tool record prompt The learner's speaking screen with everything that scores removed. The sentence renders completely uncoloured — capture is not scoring, and errors are data, never marked.
tool record arming Wait for the tone. The teacher must not tell a student to speak while the take has not begun.
tool record recording Live capture. One button, no cancel, no per-sentence skip — an aborted take is corpus data, not something a teacher may discard.
tool record saved The receipt, which is what replaces the score here. Green means the bytes landed and says nothing about the speech. Listen back plays what actually went up; then next, or re-record.
tool record retake The take-quality guard re-armed the mic for the same sentence. Soft guidance, never a grade — the 'forgive the room' side of the line: retry, and tell the teacher.
tool record pending The wifi died. The take is safe on the phone and the flow ADVANCES, but the ledger says pending and the sitting cannot close until it lands. Saying 'saved' over audio on a device is the exact permanent, undetectable mislabelling the corpus rule exists to stop.
tool record noise The noise block — the same five sentences read in a deliberately noisy room. Skipping is never silent: it requires a reason, because a skip that leaves no trace is a hole in the data nobody would ever find.
tool record fault The one screen whose answer is STOP. The server refused the take itself — the two things a retry genuinely cannot fix. The control goes down and stays down, dimmed not hidden, so it reads as 'stopped on purpose'; message Chau, and every take so far is safe.
tool session complete Can I let the student go? The question the teacher actually has, as the headline. Completeness is derived from the takes that exist, never from a timestamp.
tool signed out The sign-in ran out, mid-anything. One screen for the recording and rating surfaces both. Not an alert — nothing failed. The load-bearing line: takes spooled on this phone upload only under the account that recorded them, so 'sign in again' means 'as cô Hường'. No self-service account flow; message Chau.

Rating — the ground truth

5

Three external teachers judge recorded audio one syllable at a time, and this is what the scorer is validated against. Independent, blind to what the model thought, and free to abstain — an honest 'cannot judge' is worth far more than a guess. No student identity appears anywhere.

Admin — deliberately small

2

Reads the corpus; writes only people-and-assignment rows. It may not edit or delete a capture, take, participant, session or catalog — destructive corpus operations stay deliberate SQL, because an admin UI that can mutate corpus rows is a corpus-integrity hazard with a comfortable interface on it.

The system

2

The tokens and primitives every screen is built from, and Toni. Rulings live in DESIGN.md.