How these word lists are made
Most vocabulary pages do not say where their words came from or who checked them. This one does, because that is the only claim worth making: not that we are experts in German, but that you can see exactly what is behind every line.
Which words end up on a page
Words are ranked by how often their forms occur in a corpus of everyday speech, and the list is cut at a rank, not chosen by taste. A word that does not appear in real speech does not make it in. Senses marked in the dictionary as archaic, outdated, historical, rare, elevated, vulgar or derogatory are removed before anything else happens.
Published word lists from language institutes are copyrighted, so none are used here. The selection is our own.
What comes from a dictionary, unchanged
Article and gender, plural, the three verb forms, the auxiliary, whether a prefix separates, case government, comparative and superlative, and the IPA transcription. None of this is generated. It is read from German and English Wiktionary, and where the two disagree the entry is dropped rather than guessed.
What is generated, and what checks it
Example sentences, collocations and the short notes are generated, then put through automatic gates before publication: every word form is checked against an index of over a million real German forms, sentence length and vocabulary ceiling are checked against the level, and officialese is rejected. Examples that fail are not published.
What the gates do not check is whether a sentence sounds natural to a native speaker. That is the honest limit of this method, and it is why every page has a way to report a mistake.
Who builds this
Denis Kucherenko, who makes the Deka vocabulary apps, and who is learning German himself. Not a native speaker and not a teacher. The grammar on these pages comes from dictionaries; the tooling, the selection and the checks are mine.
Licences
Wiktionary data is used under CC BY-SA and credited on every page that uses it. Audio is generated with one natural German text-to-speech voice (Google WaveNet), the same for every word and example. The frequency list (hermitdave/FrequencyWords) is CC BY-SA.