Skip to main content

How this works

The whole project, explained

Teaching AI to speak Igala, with the people who speak it. This page lays out why we started, who is doing the work, how the method fits together, and what the pilot has shown so far. The numbers on this page update themselves from the live annotation platform.

Why we started

Frontier AI fails the world's small languages

Ask today's best AI models to work in Igala and you get fluent-looking nonsense: wrong spelling, wrong words, the language quietly swapped for a neighbour like Yoruba.

Igala has around two million speakers in Kogi State, Nigeria. It is tonal, so a shift in pitch can change what a word means. The major models were trained with almost none of it, so they guess, and they guess confidently.

We put real model output in front of native Igala speakers and asked a simple question: is either answer good enough? The figure beside this text is their verdict. It is not a rounding error. It is the size of the gap.

Nobody could point to that gap before, because nobody was measuring it. That is what this project changes: it turns “the models are bad at our language” into something you can count, and hold every model accountable for.

of AI answers rejected by native speakers

Share of blind comparisons where speakers judged both AI attempts inadequate. Live from the pilot.

Who is doing this

Led by the community that speaks the language

This is built with a community, not on one. The people who speak Igala decide what good looks like, score the models, and own what gets built. In the Wikitongues model, communities come to the archive - the opposite of extraction.

Convenes and builds

Wikitongues

A nonprofit archive working to document every language in the world. Wikitongues runs the initiative and builds the annotation platform behind it.

Lead the language work

The Igala Wikimedians

A community of Igala speakers and Wikimedia contributors based in Abuja, Nigeria. They write the reference answers, correct the models, and set the standard for what counts as real Igala.

Guide the method

Academic and industry advisors

Researchers from New York University and Google Research advise on evaluation design and methods for low-resource languages, alongside an independent linguistics lead.

How it works

One episode, one flywheel, one ladder

The method has three moving parts: how a single judgement is made, how those judgements compound into a better model, and how we prove the model is actually improving.

1. Inside one annotation episode

Every judgement follows the same four steps. Writing an answer first, before any AI is shown, is the key move: it captures how a speaker would really say it, uncontaminated by the model's phrasing.

  1. Write your own answer

    The speaker answers the prompt in their own Igala first, before any AI output appears.

  2. Compare two AI attempts, blind

    Two model answers appear side by side with no labels, so no brand or reputation can sway the choice.

  3. Explain the choice in English

    The speaker says, in plain English, why one answer is better, or why both fall short.

  4. Score the winner

    A short rubric captures spelling, grammar, word choice, tone marks, meaning, and whether it is really Igala at all.

2. The data flywheel

Each episode leaves behind gold: correct answers and clear judgements. That gold trains a better model, which is judged again by the community. Every turn of the wheel raises the floor.

Theflywheel12345each turn raises the floor
  1. Community gold. Speakers write correct Igala and correct the models. This is the raw material.

  2. Fine-tuning. That gold teaches an open model to prefer real Igala over its confident guesses.

  3. Blind arena. The new model goes back in front of speakers, unlabelled, against the others.

  4. A better model. The judgements show where it improved and where it still fails.

  5. Back to the community. The remaining gaps set the next round of prompts. The wheel turns again.

3. The method ladder

We climb one rung at a time, and never grade ourselves on our own homework. A frozen set of questions is locked away as the exam, so a good score cannot be faked.

1Benchmark2SFT3DPOone rung at a time
  1. Benchmark. First, measure. Score today's models on Igala so there is an honest baseline to beat.

  2. SFT. Supervised fine-tuning: show an open model thousands of correct, community-written answers.

  3. DPO. Preference tuning: teach it from the community's own better-versus-worse judgements. This is the finisher, not the teacher.

The frozen exam

A held-out set of community-authored questions, sealed off from training. It is never used to teach the model, only to test it, so improvement is real and not memorised.

What we have learned

Early findings from the pilot

The pilot is small and honest. Here is what the data already shows. The counts below come straight from the live platform, so they move as the work continues.

99%
of AI answers rejected by native speakers
300
prompts in the evaluation bank
43
questions locked in the frozen exam
548
community gold answers collected
481
blind comparisons judged
8
native speakers contributing

Off-target output, before and after

41%3.1%non-Igala content in model output, after one plain instruction and before any fine-tuning

Loading the latest numbers from the annotation platform...

The models almost always fail

In blind comparisons, native speakers judged both AI answers inadequate the overwhelming majority of the time, and they were confident about it. There is no defensible winner among today's models yet. That absence is itself the finding.

Models reach for the wrong language

When a model does not know an Igala word, it does not stay silent. It borrows from Yoruba, Igbo, or Nigerian Pidgin and presents the result as Igala. A speaker caught it live: the word a model gave for “morning” was not Igala at all. A simple instruction that names and forbids this cut off-target output by roughly an order of magnitude, before any fine-tuning.

Speakers agree on words, not always on spelling

A quiet but useful surprise. Independent speakers picked the same correct word almost every time, while writing it with visibly different spelling and tone marks. So the first thing to align is not vocabulary or meaning. It is spelling conventions.

What is next

Where this goes

The pilot runs toward a public launch this autumn.

  1. A community-tuned model in blind testing

    The first model fine-tuned on community gold goes back into the blind arena, judged against the frontier models on the frozen exam.

  2. A public launch in Ghana

    The initiative and the first Igala model leaderboard launch publicly at the Wikimedia Foundation conference in October 2026.

  3. A method others can reuse

    A written account of what worked, so the next community can run the same playbook for their language without starting from scratch.

The evidence

Further reading

The thinking behind the benchmark and the methods it draws on, annotated for a general reader.

Evaluation & the floating-motifs problem

How do you measure whether a model truly speaks a language, rather than producing fluent-looking text? This is the question the benchmark answers.

  • The eight evaluation dimensions

    Wikitongues AI

    A working framework for where models fail at Igala - from orthography to authenticity - turned into a rubric the community can score against.

    • framework
    • benchmark
  • Open question: does community-written Igala beat translated text?

    To come

    Wikitongues AI

    A live research question for the pilot: whether text written by speakers outperforms text back-translated from English when teaching and testing a model.

    • research question

Methods for low-resource & oral languages

Techniques for building language technology where there is little written data and a strong oral tradition.

  • Google Research (with the Gates Foundation)

    An openly licensed (CC-BY-4.0) speech dataset spanning many African languages - the kind of community-rooted resource that makes underserved-language AI possible.

    • dataset
    • CC-BY-4.0
    • speech
  • Adapting models to low-resource languages

    To come

    To be annotated

    A placeholder for the methods literature on fine-tuning, retrieval, and corpus-building for languages with limited written data.

Community-rooted language AI

Peers and analogs building language technology with, not on, communities.

  • Community-led language technology: peers and analogs

    To come

    To be annotated

    A placeholder for projects that put communities in charge of what gets built and how their language is represented.

Igala & the corpus

Sources on the Igala language and the corpus feeding the pilot.

  • The Igala Wikipedia corpus

    To come

    Ikala Wikimedians

    Community-written Igala from the Igala Wikipedia, fed into the models as a first corpus for the pilot.

    • corpus