Models guide

Add your language to Kuma.

Kuma transcribes with Whisper models in the GGML format used by whisper.cpp. To work in a language, you need a model that supports it. There are four ways to get one, from easiest to most involved.

Route 1

Built-in model

Download a standard multilingual model. Covers about ten African languages.

Route 2

Import a model

Someone already made a GGML model for your language. Load it in one step.

Route 3

Convert a fine-tune

A Hugging Face model exists, but not in GGML. Convert it once.

Route 4

Train your own

No model exists yet. Fine-tune Whisper, then convert and share.

1. Use a built-in multilingual model

Open the Models panel in Kuma and download a multilingual model such as small, medium or large-v3. Bigger models are more accurate but slower. One multilingual model serves every language Whisper supports, so you only download it once.

Stock Whisper gives usable results for a handful of African languages, including Swahili, Yoruba, Hausa, Amharic, Afrikaans, Somali, Shona, Lingala and Malagasy, though quality varies. For anything else, use one of the routes below.

2. Import a community model

If someone has already trained and converted a model for your language, you can load it directly. In Kuma, open Models, then Add a community model:

The imported model then appears as a selectable model. The African languages catalog inside Kuma lists known community models, with honest labels for which ones are ready to import and which still need converting.

Tip: a GGML whisper model file begins with the bytes ggml and usually has ggml- in its name. If a download is actually an HTML page or the wrong format, Kuma rejects it instead of loading a broken model.

3. Convert a fine-tuned Whisper model to GGML

Most community fine-tunes on Hugging Face are published in the Transformers (PyTorch) format, not GGML, so Kuma can't load them directly. Convert one with whisper.cpp's conversion script. You need Python with torch and transformers installed.

  1. Get whisper.cpp
    git clone https://github.com/ggml-org/whisper.cpp
    cd whisper.cpp
  2. Download the Transformers model (example: a Swahili fine-tune)
    git lfs install
    git clone https://huggingface.co/USER/whisper-swahili models/whisper-swahili
  3. Convert it to GGML
    python3 models/convert-h5-to-ggml.py \
        models/whisper-swahili ./ ./models

    This writes models/ggml-model.bin.

  4. Optional: quantize it to make the file smaller and faster, with little loss in quality
    cmake -B build && cmake --build build --config Release
    ./build/bin/quantize models/ggml-model.bin \
        models/ggml-model-q5_0.bin q5_0
  5. Import into Kuma

    In Models → Add a community model, import ggml-model.bin (or the quantized file). Give it a clear name so you recognise the language later.

Exact file names and build steps change over time. The authoritative reference is whisper.cpp's models/README, which documents convert-h5-to-ggml.py and quantization.

4. Fine-tune Whisper on your language

If no model exists for your language yet, this is the step that actually moves things forward. In outline:

  1. Gather labelled audio — short clips paired with accurate transcripts. Good sources are Mozilla Common Voice, Google FLEURS, and your own recordings. A few hours already helps; more is better.
  2. Fine-tune a base model — start from openai/whisper-small or whisper-medium and fine-tune with Hugging Face Transformers. The community fine-tuning guide is a solid starting point.
  3. Convert to GGML — follow Route 3 above, then import into Kuma.
  4. Share it — publish the model on Hugging Face so others can use it, and tell us so it can be listed in the in-app catalog.

5. Give your corrections back

Every transcript you fix is training data. In Kuma, use Contribute dataset to export your corrected clips as a ready-to-share speech dataset, then donate it to an open project like Common Voice or Masakhane. The data gap behind African speech recognition closes one contribution at a time.


Stuck, or want your language added to the catalog? Open an issue on GitHub and we'll help.