On September 18, I asked again why local translation felt so slow. I'd press the button and wait more than ten seconds. The screen showed that something was happening, but I couldn't tell which part was taking the time.
We timed the same text in separate stages. The 27B run took about 16.2 seconds: 12.1 to load the model and 3.7 to generate the text. The 12B run took about 8.5 seconds, with more than half of that spent loading too.
The program started the model afresh for each translation. During much of what I had thought of as translation time, it hadn't started translating yet.
Keeping the model loaded would save that wait on the next request, but I also used the GPU for transcription and Qwen. I didn't want translation to keep that memory indefinitely. Now I could see what a quicker response would cost in shared resources.
I also asked for automatic detection as the first source-language option. I didn't need to choose Chinese or English every time I opened the page. The target language should still be my choice. Small changes like that become noticeable when you use a tool several times a day.

