The project that got me working on my home computer again was Strata. It was started by Niko, who uses the name Niko1221 on GitHub. The model comes from the Qwen team. Niko and the contributors provide the engine that runs it on a personal computer, with the code released under the MIT License.
The first release appeared on September 24, 2026. Its ready-made Windows engine mainly targeted newer NVIDIA cards. From the beginning, the aim was to run the 125B Qwen3.8-Flash-Next model on ordinary PCs, with a script to get the installation started.
Support grew to cover more equipment, including AMD cards and multiple GPUs. People kept working on problems encountered during installation and use. The October 8 update was still improving short requests, multi-GPU operation and behavior when memory was tight. For older cards, community developers had also published V100 porting and repair notes.
Those records were what led me to try my own V100.
NVIDIA introduced the V100 generation in 2017. It started life as data-centre equipment. Used cards have since found their way into the hands of people like me. In Chinese second-hand hardware circles, they often get called “foreign junk.” Mine sits in an external enclosure connected to a second-hand laptop. Seeing this old equipment run a new large model was enough to make me curious.
I had already been using a 27B local model for a while: organizing transcripts, translating and finding information in long documents. When I found Strata, I wanted to see how much more the equipment I owned could do.
Getting it onto my own computer
Knowing someone in the community had made it work was a useful starting point. Installation on my computer still took some persistence. A driver check failed, compilation stopped, then the linker produced a long series of errors.
I used DeepSeek Harness (DSH) for this project. I have never learned to program, so I gave it the complete errors and asked it to investigate. Then I tried its changes. I also had transcription software and other tools already working on the computer, and wanted to keep them usable while I experimented.
This took time. When the engine finally started, replies were slow. I wanted to know whether the setup was usable and whether it was worth spending more time on.
A log entry showing only about 1GB of free VRAM eventually led us to an old background program.
Earlier, I had added a watchdog to restore the model after it was unloaded. After the switch to Strata, that program kept following its old rule and loading the 27B model again. As the new engine tried to use the card, the old program took the memory back.
I had been worrying that the V100 was simply too old. Finding this cause was a relief. We changed the watchdog and let the new engine manage its own loading. Replies became much smoother.
Looking back, it felt familiar from everyday software use. A feature had been useful, and I had become accustomed to it being there. Changing what ran underneath meant coming back to check that feature as well.
That is when I feel most involved in the development. DeepSeek Harness handles the code, while I have to remember why we added something and try what it now does on my computer. An error message doesn't tell that whole history for me.
What stayed useful after the first run
Once it was running, I brought back tasks I had actually used before. The model could take a long piece of text, extract the requested information and return a format the next program could read. Some requests that had previously produced no result now returned something.
I opened the results to check them. Had the material really reached the model? Could I find an extracted number in the source? Had the resulting file been saved? Seeing a reply was encouraging. Getting something I could use made me want to open the tool again the next day.
Some tests also made me revise my first judgment. A batch of strange answers turned out to have received almost empty input. The preprocessing had dropped the text and left separators. I had suspected the model, while the failure was earlier in the process.
In a test asking for performance notes on individual podcast lines, the old 27B answer was better tailored to each line. The newer model repeated similar templates. I still needed to see how it behaved on the particular work in front of it.
I tried a quantization setting that used more resources, too. After downloading it, preparing it and running the old tasks again, I switched back. That run used more RAM and was slower. I hadn't found enough practical benefit to keep it.
By then I felt less urgency about changing models. A configuration that worked comfortably deserved some time in daily use. When a task exposed a real limit, I could decide what to try next. Chasing every release could leave many files on the computer without giving me another useful result.
I would rather spend time on the everyday details: getting the complete material in, finding the output afterwards and having a clue to follow when something fails. Those details make it easier for the model to become a tool I actually use.
Keeping the old computer
When software felt awkward, I used to put up with the extra clicks or look for another tool. Now I can describe what I want to AI and try the changed version myself. That has given me more things I want to make.
The laptop, enclosure and V100 arrived in stages. I wrote about that in the earlier story. When I bought them, I didn't expect to end up with a model centre, a job queue and eventually a 125B model.
Open source lets me continue from work other people have already done. Someone writes the engine, someone gets an older card working, someone publishes the error they ran into. I take that work back to my own computer and gradually fit it into my life.
It has made me a little more patient with second-hand equipment. There are limits and installation can take work. If it does what I need, I can keep using it. With a limited budget, I am happy to try that route first.
Local inference removes the separately billed cloud requests for this part of the work, and the material can be processed on my computer. Disk space, RAM, electricity and time still cost something. I used DeepSeek Harness for this development work. I judge the expense by what I can actually use.
I still couldn't write this software independently. But I have tried the old card doing useful work at home. I will keep using it and fixing the problems I encounter. When a task finally gives me a reason to buy something else, I can consider it then.
