State of the art, Indigenous-built technology
AI Anima is an initiative dedicated to building a modern language and knowledge model that learns, processes, and communicates primarily in the Cross Lake dialect of Cree. Our technology is built right at home in Cross Lake, Manitoba and supports the continuity of our cultural worldview into the age of artificial intelligence. Our research contributes to a range of advanced AI applications designed for native speakers, new speakers, and everyone in between.
Equipped with our technology, our community is empowered to use our language in digital spaces. As our models mature, they will support community programming, create employment pathways for Indigenous youth, and contribute to broader conversations about Indigenous data sovereignty and ethical AI.
Speech Database
All of our technologies rest on our huge database of spoken Cree. Our novel approach to data collection involves per-minute payment for any fluent speakers who submit audio through our bespoke recording app. Any fluent community member can contribute in three ways:
- Recording directly in the app
- Uploading audio or video files
- Making phone calls in Cree directly in the app
Users can even split earnings when multiple speakers are in a recording. Our secure database is built and managed right at home, and also acts as a repository of stories where the contributor agrees to share a recording with the community.
End-to-end, our custom software handles audio processing and storage, app user profiles, payments, transcription file management, and orderly data handoffs to our model training systems.
Speech Recognition
Every minute of spoken Cree in our database is reviewed by our team of transcribers who annotate the audio submissions. These annotated audio files are used to train our speech recognition model. The results include:
- Speech to text technology, or automatic transcription
- Synthetic generative Cree voices that work with our large language model
Translation
Every minute of spoken Cree in our database is reviewed by our team of translators who annotate the audio submissions with morpheme-level translations, word-level translations, sentence-level translations, and 'expanded' translations. This massive number of Cree-English pairs enables our language model to learn how to build more nuanced bridges between Cree and English meaning.
Language Model
Our language model is trained on multiple tiers of transcribed Cree speech. It learns Cree morphemes, words, and sentences in Syllabics. It learns morpheme-level translations, word-level translations, sentence-level translations, and 'expanded' translations to English. The morpheme-level training is important since Cree is a polysynthetic language and our morphological analyser allows the model to generate high quality Cree through context-aware morpheme sequencing.