Field diary

The transcript was finished. Five minutes of speech were missing.

Listening to the recording again showed that the blank stretches contained speech the software had failed to transcribe.

The transcript file had been created without an error, but the text didn't look right. I listened to the recording again. People were speaking during stretches where the transcript was blank. The longest gap was more than five minutes.

On September 7, we checked the same recording again. Another recognition model could transcribe those sections. The gaps couldn't simply be explained by silence or poor recording quality.

The fault was in the GPU quantization path used by Paraformer. It failed on some audio segments, and the surrounding library returned empty text rather than surfacing the error. Processing continued, the progress bar reached the end, and an incomplete transcript was saved.

We moved that Chinese recognition path to a safer CPU mode and reran the recording. The output grew from roughly 7,600 Chinese characters to 19,800. The largest gap fell to about nine seconds. A coverage check was added so that widespread empty speech segments would raise an error rather than produce another apparently finished file.

Earlier I had been wondering why the GPU wasn't being used more. After seeing the missing content, utilization felt less urgent. I needed the spoken words to make it onto the page. Those measurements came from this particular recording; they weren't a promise about every recording I would try.

Keep reading