User

Olafbot

Revision as of 17:15, 27 February 2021 by Olaf (talk | contribs)

Wiktionary Bots.png

The bot, created by Olaf, updates various lists of missing audio recordings every night. Much more active in Polish Wiktionary.

Lists named "Lemmas-without-audio-sorted-by-number-of-wiktionaries" are created in the following way:

  • For a given language, the bot traverses categories on all wiktionaries and a few open dictionaries and collects statistics - for each lemma it counts dictionaries that describe this word in this language. This is something the bot has been doing for 11 years, generating different lists for Polish Wiktionary.
  • Titles written in wrong alphabets are removed.
  • Lemmas with audio recording in Commons are also removed from this set. Not only files created with LiLi are removed, but also other recordings found in the "pronunciation" category for a given language or in its subcategories.
  • For a few languages, minor corrections are done, in order to extract the set of dictionary lemmas, if possible without inflected forms.
  • The resulting list is sorted descending by the number of dictionaries and limited to 5000 entries.