Kalahari Labs Open-Sources a Swahili Language Model Small Enough to Run on a Phone
The 1.4-billion-parameter model runs offline on mid-range Android handsets and was trained on a corpus Kalahari says it licensed rather than scraped.
Kalahari Labs released Mto-1.4B on Monday under a permissive license, along with the tokenizer, the training recipe, and — unusually — a full manifest of its training corpus. The model is designed to run on-device, and in Afrikons' own testing it generated coherent Swahili responses on a three-year-old handset with 4GB of RAM and airplane mode switched on.
The corpus is the headline. Rather than scraping the open web, Kalahari says it licensed 2.1 billion tokens from Tanzanian and Kenyan publishers, radio transcript archives, and three university linguistics departments, paying a reported $1.8 million in aggregate. Contributors receive a share of any commercial licensing revenue the lab collects on derivative work.
"There is a version of this where we scraped everything and shipped six months earlier," said Kalahari research lead Josephine Mbatia. "We would also have had no answer when a publisher asked us what we had taken. We wanted to be able to answer that question."
For developers, the practical appeal is latency and cost: no round trip, no per-token bill, no dependency on a connection. For the broader field, Mto is a test of whether a licensed-data model can stay competitive with scraped-data rivals — and whether anyone is willing to pay the premium for a clean provenance story.
Why the offline-first bet is reshaping African AI · The Build Loop
When you purchase through links in our articles, we may earn a small commission. This doesn't affect our editorial independence.
Wanjiru Kamau
AI and Enterprise Reporter
Wanjiru Kamau reports on applied AI and enterprise software for Afrikons from Nairobi. She spent five years as a systems analyst before turning to journalism, and still reads changelogs for fun.
View BioLoading the next article


