Home

Attributions

Where the open-data decks in the library come from.

The Kotoba Dōjō library holds two kinds of deck: decks written for the app, and decks built from openly licensed data. This page records the second kind — what we took from each source, where it came from, and the licence it carries.

Wiktionary

For each concept in a starter deck we look up the target-language word in the English Wiktionary translation tables, read through kaikki.org’s machine-readable extraction of them. The word on the front of an open-data card comes from there.

Source: English Wiktionary via kaikki.org · Licence: CC BY-SA 4.0

FrequencyWords

When a translation table offers several words for one concept, we choose between them by how common each word actually is. The rankings come from hermitdave’s FrequencyWords lists, derived from the OpenSubtitles corpus. Nothing from the lists is printed on a card — they only decide which candidate wins.

Source: hermitdave/FrequencyWords · Licence: CC BY-SA 4.0

Tatoeba

The example sentence on an open-data card is taken from Tatoeba, a collection of sentences contributed and translated by volunteers. Sentences are used as published, unaltered.

Source: Tatoeba · Licence: CC BY 2.0 FR

Share-alike

Decks derived from the CC BY-SA sources above are themselves offered under CC BY-SA 4.0. The library they sit in is not paywalled — every account, free or Pro, can open them.

Decks written for the dojo

The rest of the library was written for Kotoba Dōjō and carries no third-party licence. Each deck records its own origin, so you can always tell which kind you are reading.

Last updated: 29 July 2026