Research
Foundational research on multilingual NLP and long-tail language modeling. Selected publications:
Models
Machine translation and language models for the long tail of languages and domains. From our current work:
OkaMT-1.3B
Machine Translation Model
We introduced machine translation for four languages not currently supported by commercial systems. Together with English, OkaMT-1.3B translates between all five, 20 directional pairs in total. Oshikwanyama ⟷ English is our strongest.
Try it on OkaLex →OkaLM
Kwanyama Language Models
OkaLM is the first family of publicly available large language models for Kwanyama. Available in three sizes (1B, 3B, 8B parameters) to suit different use cases, from lightweight applications to more capable generation.
🤗 View on Hugging Face →Data
Datasets we collect and openly release for the long tail of languages and domains. From our current work:
Datasets
Open Data
We collect and openly release datasets for languages with little or no existing data. Available on GitHub.
Browse on GitHub →Oshikwanyama Lexicon
Bilingual Dictionary
A carefully curated Oshikwanyama–English bilingual lexicon, now searchable online on OkaLex, a first for the language.
Search on OkaLex →Applications
Application areas where we are building first-of-their-kind systems for the long tail of languages and domains.
Machine Translation
Question Answering
Structured Data Representation
Who we are
Founded in 2021, Okalai held its first AI school in 2022 and grew from those schools into a research program building first-of-their-kind models for the long tail of languages and domains. We have trained the first LLMs and machine-translation systems for several languages that previously had none.
Okalai AI was founded by Ndapa Nakashole, who serves as Chief Scientist. Ndapa is an Associate Professor of Computer Science at the University of California, San Diego (UCSD). Her research focuses on Natural Language Processing (NLP), and Artificial Intelligence (AI) more broadly.
Get in touch: hello@okalai.org