Six Hundred Forty-Six and One
In the last dispatch we wrote honestly about what the new speech layer does not have. Sixty languages is a lot for a product and not much for what we are after. That admission was not a surrender: the rare-language line is alive. It just became clear where the border runs between what you can buy and what you have to train yourself.
Six hundred and forty-six language codes — that is the dictionary of the engine we started from, and it is our working horizon. It is not a claim that we speak all of them. We measured the quality ourselves and published the result: far fewer of them are usable. A map of the territory, not a flag on the summit.
Right now we are training our own model on continental Portuguese. The choice is not accidental. For the big services Portuguese almost always means Brazilian, and the European variety comes along as an afterthought: swallowed vowels, a different rhythm, different words. A person in Porto hears it in the first sentence. We live here, so for us this is not an academic question.
The model will run on our own infrastructure and inside a client’s private setup — the audio does not travel to anyone else’s cloud. And we are building it in full duplex: not “the agent has finished, now it is your turn,” but a normal conversation where you can interrupt and still be heard.
Six hundred and forty-six is a target, not a status line. European languages come first and are being trained now; African and Arabic ones are next in the plan. The engineering is not what holds us back — the team has two decades behind it in telephony and knows what a phone line does to a model.
What we are short of is compute. Training a language properly takes GPU time, and the gap between the languages we want to cover and the hardware we own is the real bottleneck right now. We and our partners currently have requests in for additional GPU capacity; until they are confirmed, we go language by language rather than in parallel.
There have been no live calls on it yet, and we are not announcing dates. When the numbers arrive — latency, accuracy, cost per minute — we will publish them here, including the ones we do not like. That is how this journal works.