Skip to main content

Wikitongues AI

Teaching AI to speak the world’s underserved languages, one community at a time. Starting with Igala.

Two tracks, one principle. A tutor that helps people learn and use Igala authentically, and the first public benchmark of how well today’s AI models actually speak it. Both led by the community that speaks the language.

What this is

A scoreboard for languages the models keep getting wrong

The problem, made concrete

Ask a model to work in Igala and you get fluent nonsense - wrong spelling, wrong meaning, the language confused with neighbours like Idoma. A native speaker who tried it put it plainly: everything was wrong. The failure is real, but until now it has been invisible.

Track one

The tutor

A tool that helps people learn and use Igala authentically, with the community correcting it as it goes.

Track two

The benchmark

The lever. A public scoreboard that turns "the models are bad at our language" into something measurable, holds every model accountable, and is reusable for the next language.

The eight dimensions we measure

Where models fail at Igala, explained for a lay reader.

  • Orthography & spelling

    Are words written correctly, with the right characters and tone marks?

  • Grammar, morphology & tone

    Does it handle Igala structure and tone, where a shift in pitch changes meaning?

  • Lexicon

    Are the words actually Igala, not borrowed from neighbours like Idoma?

  • Dialectal fidelity

    Does it keep the community’s variety instead of collapsing to a prestige standard?

  • Register & honorifics

    Does it use the right level of respect and formality for the moment?

  • Idioms, metaphor & motifs

    Does it understand figures of speech, or read them flatly and literally?

  • Cultural knowledge & values

    Does it know the customs, beliefs, and values behind the words?

  • Authenticity vs translationese

    Does it sound community-written, or like text back-translated from English?

Community-led by design

Communities decide what good looks like, score the models, and own what is built. In the Wikitongues model, communities come to the archive - the opposite of extraction.

The benchmark

The leaderboard

How the major models score across the eight dimensions. Higher is better; every score is paired with a number, not just colour.

Illustrative sample - not real results. The first Igala leaderboard launches in October 2026.

Wikitongues AI Igala benchmark. Illustrative sample - not real results. The first Igala leaderboard launches in October 2026. Columns 1 to 8 are the evaluation dimensions listed above.
ModelOrthography & spellingGrammar, morphology & toneLexiconDialectal fidelityRegister & honorificsIdioms, metaphor & motifsCultural knowledge & valuesAuthenticity vs translationeseOverall
Claude625560485752645658
Gemini544750414844554549
ChatGPT504447384542514346
Gemma362932243027352831

Where we are

The timeline

Working back from the public launch in Ghana this October.

  1. Completed:

    June 2026

    Done

    Kickoff with the Igala community

    Project launches with Agnes, community lead of the Ikala Wikimedians in Abuja: a live prototype, standing weekly sessions, and an evaluation framework shaped with our linguistics lead.

  2. Completed:

    By June 22, 2026

    Done

    Prototype and question bank to the advisory council

    A refined prototype and tightened question bank go to the research advisory council, annotators gain an inline edit tool to correct outputs directly, and the Igala Wikipedia corpus is fed into the models.

  3. Completed:

    July 2026

    Done

    Lock the rubric, run the first benchmark

    We lock the scoring rubric and prompt buckets with the community, then run a baseline benchmark of the major models - ChatGPT, Gemini, and Claude - on Igala.

  4. Completed:

    August 2026

    Done

    Working session and dataset expansion

    A heads-down working session in San Francisco expands the benchmark dataset, designs the community feedback pass, and drafts the menu of data-ownership options.

  5. Completed:

    September 2026

    Done

    This site goes live

    The initiative goes public on the web with a clear menu of data-ownership options for the community and the first funder update.

  6. Upcoming:

    October 4-5, 2026

    Upcoming

    Public launch in Ghana

    The initiative and the first Igala model leaderboard launch publicly at the Wikimedia Foundation conference in Ghana.

The evidence

Research

The thinking behind the benchmark and the methods it draws on - annotated for a general reader.

Evaluation & the floating-motifs problem

How do you measure whether a model truly speaks a language, rather than producing fluent-looking text? This is the question the benchmark answers.

  • The eight evaluation dimensions

    Wikitongues AI

    A working framework for where models fail at Igala - from orthography to authenticity - turned into a rubric the community can score against.

    • framework
    • benchmark
  • Open question: does community-written Igala beat translated text?

    To come

    Wikitongues AI

    A live research question for the pilot: whether text written by speakers outperforms text back-translated from English when teaching and testing a model.

    • research question

Methods for low-resource & oral languages

Techniques for building language technology where there is little written data and a strong oral tradition.

  • Google Research (with the Gates Foundation)

    An openly licensed (CC-BY-4.0) speech dataset spanning many African languages - the kind of community-rooted resource that makes underserved-language AI possible.

    • dataset
    • CC-BY-4.0
    • speech
  • Adapting models to low-resource languages

    To come

    To be annotated

    A placeholder for the methods literature on fine-tuning, retrieval, and corpus-building for languages with limited written data.

Community-rooted language AI

Peers and analogs building language technology with, not on, communities.

  • Community-led language technology: peers and analogs

    To come

    To be annotated

    A placeholder for projects that put communities in charge of what gets built and how their language is represented.

Igala & the corpus

Sources on the Igala language and the corpus feeding the pilot.

  • The Igala Wikipedia corpus

    To come

    Ikala Wikimedians

    Community-written Igala from the Igala Wikipedia, fed into the models as a first corpus for the pilot.

    • corpus

Questions

FAQ, including language rights

The questions communities, funders, and the curious ask us most.

What about language rights? Who owns the data the community creates?Draft - pending sign-off

The community does. Igala speakers decide what good Igala looks like, score the models, and own what we build together. We are working with the community on a clear menu of data-ownership options, so the value, the consent, and the rights stay with the people who speak the language rather than being extracted from them. This wording is still being finalized with the community and our team.

How is this different from a company scraping our language?Draft - pending sign-off

Scraping takes language without asking. This is the opposite: the community comes first, decides what is shared, sets the standard for what counts as good, and stays in control of the result. Nothing is built about a community without that community leading it.

Is the data open source or proprietary?Draft - pending sign-off

That decision is still open and we are making it with the community, not for them. We are weighing open and proprietary options and will publish a clear menu of choices. The principle is fixed even though the mechanism is not: the community decides.

Who decides what good Igala looks like?

Native speakers. The community defines the standard, reviews the models, and corrects them. We build the tools; the community is the authority on the language.

Can my community get involved or be next?

Yes. Igala is the pilot, and the benchmark is built to be reused for the next language and the one after that. If your community wants to take part, reach out to us at hello@wikitongues.org.

Who is behind this and who funds it?

Wikitongues, the open archive working to document every language in the world, in partnership with the Igala community. The work is community-led by design and supported by early funders backing the first phase of the Igala pilot.

What models does it use, and does it send our data to them?Draft - pending sign-off

The benchmark measures the major models - including ChatGPT, Gemini, and Claude - and the tutor currently runs on a swappable backend. How community data is handled is part of the data-ownership conversation underway with the community, and we will state the specifics plainly once they are settled.

Get involved

Support the initiative

Your gift funds the Igala pilot and the benchmark that follows. Donations run through Wikitongues - add a note to direct your gift to the AI initiative.

Want your community to be next? Email hello@wikitongues.org.