Language data
Text, speech, terminology and digital archives, with language variety and provenance recorded.
Data & archivesFollow our research priorities, find Buuni on GitHub and talk to us about the language resources your work needs.
Useful language resources need documentation: where the data comes from, what it covers, how it may be used and what its limits are.
We will list individual datasets, models and benchmarks here as they become available with that supporting information.
Our public organisation is buunilabs. Visit it for any repositories we make available, and check each repository’s own documentation and licence before using its contents.
Visit buunilabs on GitHubThese are areas for research and collaboration. They are not a catalog of released products or performance guarantees.
Text, speech, terminology and digital archives, with language variety and provenance recorded.
Data & archivesTranslation, speech and knowledge tools adapted to Somali and to the needs of a specific audience.
Models & knowledgeTesting with Somali speakers, including language coverage, factual accuracy and practical limits.
Evaluation & reviewTell us what you need, how you plan to use it and which language variety matters for your work.