bk99.de entertain the web since 1997

What I learned from 2.9 billion tokens

Summary

In September I made AI speak locally, got a 35-billion-parameter model running on a single graphics card and, according to OpenAI, used 2.9 billion tokens in the cloud. A present for a friend still did not get finished. This post shows where the waiting time of a voice AI really comes from, why an HTTP 200 can lie and what a token counter reveals about AI agents.

Empty-handed to Munich

I actually had a clear plan: for my meeting with Thomas in Munich, I wanted to bring him a new DeskHop, a small switch that moves keyboard and mouse between two computers. I put many evenings and a great many tokens into it, including with GPT-6 Astra. I still arrived in Munich empty-handed.

So this month was not a pure success. But it showed me where AI really helps today, where it costs time and where it reaches its limits. That is exactly what I want to share here.

A voice from a few seconds

The most impressive experiment was a small read-aloud bot. It read me news items from the heise newsticker in the voices of Lucy and David. A few seconds of voice recording were enough as a sample.

Behind it is OpenVoice. A first model turns the text into speech, a second one then transfers the timbre from the short recording. Because the two steps are separate, the recording does not even have to be in the same language as the text being read.

What you can take away: A few seconds of your voice are enough for a convincing clone today. Every voice message and every video is therefore also raw material. A familiar voice on the phone no longer proves who is really speaking; a family code word agreed in advance does a better job.

Why GLaDOS needed 22 seconds

The second voice is the German GLaDOS, a freely available Piper model. I used it to build a voice chat that runs entirely at home: Whisper turns my question into text, Gemma 3 writes the answer, Piper speaks it. No audio leaves the house.

In the first test the answer took almost 22 seconds. The model was not too slow. It was waiting: my document processing used the same graphics card with a different context size, and the requests were queuing. With two parallel slots in Ollama and a uniform context size, the first text arrived after 1.6 seconds. A single value sped up speech output itself: with eight instead of one compute thread, it dropped from 3.1 to 0.7 seconds for a short sentence.

What you can take away: For voice assistants, measure the time to the first word. If it is bad, look for shared resources first before you buy a faster model.

A big model on a small card

For longer texts I got Qwen3.6-35B running on an RTX 5070 Ti with 16 GB of graphics memory. That sounds impossible, but it works because it is a mixture-of-experts model: for each token only a small part of the network computes, while the other experts wait in normal RAM. In my test the system processed a good 18,000 tokens of input in just under eight seconds.

The most important lesson came from an error. The server answered requests that were too long with status 200, meaning “all good”. The actual error message only appeared in the running data stream. Any monitoring that only looks at the status would have missed the error.

What you can take away: A green light is not proof. For AI services, check the result itself, not just the status code.

The agent that went silent

In the background runs Hermes, an AI agent that is supposed to summarise Hacker News and heise for me every day. Through the router Omniroute it automatically picks a suitable model from various providers.

A check revealed that none of its six scheduled jobs was running anymore. The agent itself was reachable and looked healthy. But every job failed right at the start, because the scheduler tried to create a systemd scope that did not exist in its environment. The summaries simply stopped coming.

What you can take away: For automations, monitor the result, for example “Did a summary arrive today?”, and not just whether the service is running.

What 2.9 billion tokens really mean

The biggest number of the month comes from the OpenAI web interface: 2.9 billion tokens in September. For 1.31 billion of them I was able to analyse the logs of Codex on my admin machine. The result surprised me.

Only 0.35 percent of them were answers from the AI. Almost everything else was input, and 94.5 percent of that input came from the cache. An agent re-reads its entire context at every step: files, logs, the conversation so far. So billions of tokens do not mean billions of written words, but context re-read billions of times.

What you can take away: For agents, context size and caching determine cost and speed. Small, clearly scoped tasks are therefore cheaper than an agent that constantly reads everything.

Why the DeskHop still did not get finished

Back to the present for Thomas. My variant hDeskHop was meant to use two directly mounted RP2040 chips instead of ready-made controller boards and to work even with just one connected computer. Each side has its own ground and powers the other through a galvanically isolated branch.

Digitally, the design is finished: 221 components on 62 by 72 millimetres, zero errors in the KiCad checks, firmware that builds and passes its tests. But my actual goal was a device about as cheap as the original DIY build. The components alone cost around 53 US dollars, some parts were not available in sufficient quantity, and the power margins are only calculated, not measured. A first prototype was never made.

What you can take away: AI can produce schematics, layouts and firmware at an astonishing pace. It cannot order parts, solder or measure. So set the cost target and the first prototype early, before the next billion tokens flow into fine-tuning.

Numbers at a glance

  • OpenVoice has been under the MIT licence in versions 1 and 2 since April 2024.
  • The German GLaDOS voice is a freely available Piper model of about 114 MB.
  • Qwen3.6-35B-A3B in NVFP4 format comprises about 23.5 GB of model files and processed 18,249 input tokens in 7.7 seconds in my test.
  • Of the 1.31 billion Codex tokens analysed for September, 4.5 million were output.
  • The hDeskHop design needs 57 different manufacturer parts.

References

Remarks

  • All times are single measurements on my hardware and not a general performance comparison.

View OpenVoice on GitHub