BurkimbIA

non-profit
Activity Feed

AI & ML interests

LLM, Text2Text, Text2Speech, Moore Language

Recent Activity

sawadogosalif  updated a Space 3 days ago
burkimbia/README
sawadogosalif  published a model 3 days ago
burkimbia/tengsoaba-4b
sawadogosalif  updated a model 4 days ago
burkimbia/tengsoaba-1.7b
View all activity

Organization Card

Ni yibeogo 🫡 🇧🇫

Who we are A community of engineers, researchers, students and professionals building open Burkinabè AI, aligned with our values, identity and culture.

What we build Open datasets and models for Mooré (mos), a language of Burkina Faso with almost no digital presence. Other languages of the country will follow, and the pipeline is built so they can.

Why it matters Mooré is low-resource. The bottleneck is never the architecture, it is the data. Everything we publish is meant to move that constraint for everyone working on the language, not only for us.


What we have published

Translation, French to Mooré and back NLLB fine-tuned at several sizes, and Mistral-7B.

Language models for Mooré Instruction-tuned models trained on a set of structured tasks: translation both ways, spelling correction, quality judgment, terminology and standardisation.

On top of them we build a conversational assistant, which is a different thing from a translator: a translator turns an input into a parallel output, an assistant answers a question whose answer is written nowhere. A general-purpose backbone is adapted to Mooré through continued pretraining and instruction alignment, with a replay mix so the model gains the language without losing the reasoning it already has. The mix is composed by capability rather than by language, and code is in it as a carrier of structured reasoning, not as a programming language.

Speech recognition Whisper fine-tuned for Mooré, several generations, served in production.

Speech synthesis Spark-TTS with a BiCodec codec, plus XTTSv2, VITS and ParlerTTS.

Data An aligned French-Mooré text corpus, and a speech corpus taken from ingestion through diarisation to consolidation.


Benchmarks, and why we publish them

A score only means something if another team can reproduce it. We publish the test sets, not just the numbers.

Collections


Responsible use Resources are released for research and educational use. Contributors must respect consent, provenance and licensing requirements. Commercial use requires permission from the BurkimbIA community.

Mooré output should be reviewed by a speaker before downstream use. Our models can produce word forms that look correct and are not attested anywhere in the corpus, which is hard to detect without one.

Get involved

Become a Burkimbila. Collect data, improve models, write documentation, deploy, or run evaluations. Every contribution counts, and the one that helps most is a Mooré speaker willing to review output.

BurkimbIA, building AI for Burkina Faso, by Burkina Faso 🇧🇫