I had used online speech recognition for years. Upload a recording, wait a little, get the text back, then correct it myself. That process had saved me a lot of time, and I was used to it.
When I started looking more closely at how I handled work material, though, there was another question to answer first: could this recording leave my computer? Some material could not go to an outside service. However convenient a tool was, I had to settle that question before using it.
What I wanted had barely changed. Turn a recording into text. Replay the parts that were unclear. Make corrections and export a Word document. This time, I wanted the whole process to stay on my own computer.
So I asked Codex to help me build a transcription application. I had no development experience. I did not know which model to choose or how to make a desktop window. I started by explaining how I would actually use the software.
For the first version, I asked for four things
Let me drag in a recording. Separate the speakers. Let me change their names. Export the result as TXT or Word.
Meeting minutes could wait. What I needed first was a transcript I could check; another tool could help with the later editing. Those four requests sounded ordinary, but I knew that if any one of them was awkward, the application would struggle to become part of my daily work.
The first approach was a Windows desktop application. We tried Paraformer for Chinese and Whisper for recordings mixing Chinese and English. Speaker separation used pyannote, the window used PySide6, and Word export used an existing component too.
At the time, those were mostly names Codex was explaining to me. I cared about whether I could see progress after putting a file in, and whether there would be a usable transcript at the end. AI connected the components. I tried the application with recordings I actually needed to process.
After a few attempts, more requests began to appear.
I needed a stop button. Sometimes I might have chosen the wrong model; sometimes I just wanted to see the part already transcribed. I needed a way to stop instead of waiting for the entire recording. The number of speakers should be detected automatically, with an option to choose it myself when I knew how many people were there. Once transcription finished, the area for choosing a file and a model should shrink and give the text more room.
I could not have written a complete design for all of that at the beginning. Using the application showed me which panel took up too much space and which action was missing a button.
Changing a name turned out to be two different things
Once the voices were separated, the transcript showed labels such as “Speaker 1” and “Speaker 2.” The software could judge that several passages probably belonged to the same voice. It did not automatically know the person's name. I still needed to check the recording and assign the names.
Then I found that changing a name had to cover two separate situations.
If I had confirmed that “Speaker 1” meant the same person throughout the transcript, I wanted to replace that label with their name everywhere. But if the model had assigned just one passage to the wrong speaker, changing that row should affect only that row.
The original behavior mixed those actions together. A correction to one passage could change labels across the whole document. Both operations sounded like “renaming,” yet they changed different things. They needed to be separate and clear.
The player needed similar attention. When I moved the playback position, the text below should follow. Clicking the time beside a passage should start the recording there. I did not want to listen while hunting through a long transcript for the paragraph I was hearing.
A long contribution from one person needed paragraph breaks. An incorrectly recognized word needed an option to replace it throughout the text. When I exported a checked transcript to Word, the paragraphs and corrected names had to come with it. The exported document had to reflect what I had just edited on screen.
I could not write those functions myself. I could tell very clearly which behavior would get in the way of my work.
The speakers looked right, but a large part of the content was missing
The experience that changed this project was a long recording.
It ran for more than an hour, with several people taking turns. The application finished, produced a document, and separated the speakers. At first glance, the workflow seemed complete. Reading further, I found large empty stretches, both earlier and later in the transcript. Some of the meaning was clearly wrong too.
My feedback to Codex was direct: the speaker count and separation were broadly right, but the content was far off and much of it was missing. I wanted the model problem fixed first. I did not want a long explanation of all the possible reasons it might have gone wrong.
A check against the original recording confirmed that there was speech in the gaps. Another recognition route could also transcribe sample passages from those stretches. The recording contained material that this run had failed to deliver.
The development notes kept the comparison from that particular case. For a recording of about 71 minutes, the original GPU route produced roughly 7,647 characters. The longest gap in its transcript timeline was about 312 seconds.
More than five minutes. Set against the recording, that stopped being an abstract technical number. There was speech within that stretch, and the transcript had failed to capture it.
I had also made an understandable assumption: with a graphics card installed, running on the GPU ought to be better. This time, neither GPU activity nor a process continuing to the end guaranteed a complete transcript.
We first went back to the CPU to recover the missing text
Codex followed those blank stretches through the processing and found the problem in the quantized GPU route used by Paraformer at the time. It could fail when handling speech blocks of different lengths, while a third-party library treated some of those errors as silence or noise. In other words, recognition had failed, but the result reaching the application looked as though no one had spoken in that passage.
The interface did not clearly tell me that a passage had failed. It moved on to the next one and eventually delivered a document that appeared finished.
The same sample passages produced text consistently on the CPU route. For that repair, we first used CPU INT8 rather than continuing to push GPU utilization higher.
The recorded full rerun took about 125.8 seconds. Of 336 detected speech intervals, 334 produced text, giving roughly 19,800 characters in total. That was over ten thousand characters more than the earlier result.
Those figures describe one recording and the two processing routes tested at that time. They do not mean every recording of mine runs at that speed, or that every character was correct. Names, numbers, and important statements still needed checking against the audio. Large gaps in the transcript needed to be visible as a problem, even if the process had reached its final step.
We then added a completeness check. Of the intervals where speech had been detected, how many had actually produced text? If the proportion was far too low, the application should report an error instead of quietly saving an incomplete result. One threshold set then was 70 percent. That measured coverage of detected speech intervals; it was not a recognition accuracy score.
The messages I wanted were practical too: which part had produced no text, why processing had stopped, and whether the completed work had been saved. When something went wrong, I needed to understand what I still had and where to go next.
The application also had to fit on my computer
Once the models had been downloaded, disk space became a problem as well. My C drive was already short of room. I asked Codex to move the project's workspace and models to the D drive. Recording projects, results, and logs needed a place of their own too.
That might not look like a central transcription feature, but it directly affected whether I could keep using the application. Processing audio locally would not help me if it filled the system drive.
GPU memory also needed a sensible sequence. Transcribe the text first, then separate the speakers, releasing the earlier model's resources where necessary. If speaker separation failed, the text already recognized still had to survive. Stopping a job should preserve what had been completed as well.
The original recording needed to stay intact. Volume adjustment or noise reduction should use a temporary copy. If I later wanted to check a sentence or try a different model, I still needed the same original recording available.
Automatic project saving, reopening earlier results, and continuing to edit names and text all became part of the application. Those first four requests had led to much more work than I had imagined.
I still need to check the recording
The software later took the name “知微见著” (Zhiwei Jianzhuo). Transcription became one part of it, with translation and text editing tools added alongside it.
But when I return to a recording, the transcript is still what matters most to me. I need to open it, keep editing the text, correct speaker labels, and check important content against the original audio.
Running locally does not make the result automatically correct. Speaker numbers can change when a recording is processed again. Recognition can mishear words. Names, numbers, and what a sentence actually means still need a person to look at them.
Starting to build software with AI did not suddenly teach me model inference or turn me into a programmer. I took the inconveniences I had put up with in other tools and began asking for them to be fixed, one at a time.
When a stop button was missing, I asked for one. When correcting one speaker changed the whole transcript, I asked for two separate actions. When the application said it had finished but left a five-minute gap, we checked the source recording to find out what had happened.
By that point, I had a working transcript I could organize, check, save, and export on my own computer. The application could develop further. What it delivered to me already needed to make its behavior understandable and preserve the corrections I had made.

