MendSpeech (Meaning-Preserving Dictation)
A real-time dictation system built around one narrow question: can it make a transcript easier to read without changing what the speaker said, and what would that cost in latency? Built so far: a seeded acoustic-damage suite, a frozen 30-utterance benchmark across five speakers, Wav2Vec2 recognition with greedy CTC decoding, and the alignment algorithm implemented from first principles. The planned system adds cache-aware streaming ASR, confidence calibrated against correctness, a conservative editor that may format a transcript but never rewrite it, bounded post-training of that editor, and a correlated per-stage latency budget.
PyTorch · Torchaudio · Streaming ASR · CTC decoding · Confidence calibration · SpeechDamageBench