Listen mode — building the one thing that actually lost the interview
The owner failed an interview for a specific reason: he couldn't parse fast native English in real time and the conversation broke down. The whole platform was speaking-heavy, and its listening surfaces showed the partner's line as text — so the ear never trained. Built a no-subtitles fast-listening drill that reproduces the interview condition.
After the rejection, the diagnosis was clean and painful: it wasn't knowledge or speaking, it was that he couldn't follow fast native English in real time, so the conversation stalled. That reframed everything — the platform had become a speaking gym (talk, sparring, AI-audit, phrases, voice gates), but the skill that actually lost the interview was listening, and there was no surface for it. Worse, the existing listening output (the partner's reply in /talk, /sparring) is shown as text, so the natural move is to read it — the ear never gets trained.
So /listen hides the text. Each turn plays a natural fast-native line as audio only — no captions — at 1 to 2x with a selectable accent. He has to understand it and answer by mic. Only after he answers does the line get revealed, with a comprehension score, the Korean gist of what it meant, the exact words and linking he likely missed (the "gonna", the swallowed function words), and a one-line listening tip — then a slightly harder line. Replays are capped at two so he can't lean on re-listening; the point is first-pass comprehension, the way an interview works. It reuses the speed-controlled TTS, the VAD mic loop, and transcription already built for the speaking modes.
Verified live: it opens with something like "So can you tell me about a project where you had to balance cost and quality?", and a vague answer comes back scored 2/5 with the gist and a harder follow-up.
The honest framing that went with it: fast-listening is a separate, trainable skill, not a sign of being slow — it's the latest-blooming ESL skill and it's about ear-exposure hours, not intelligence. The tool reproduces the pressure; the real cure is volume (fast audio, no captions, every day) plus the recovery phrases so a missed word never breaks the conversation again.
The reusable lesson: a listening drill must hide the transcript until after the attempt, or the user reads instead of listens — and the failure that loses the interview deserves its own surface, not a side effect of a speaking tool.