🇨🇿 Czech 11 million speakers

The 5,000 most common Czech words

Ranked by how often they actually turn up in spoken dialogue, not by how useful somebody guessed they would be.
Every number below is counted from a corpus, and the corpus is named at the bottom.

What each step buys you

Learn the first 100 words of Czech and you follow 43.9% of everyday speech. The first 1000 take you to 70.6%. After that every thousand words buys less than the thousand before it, which is exactly why the order matters.

Words knownSpeech followed
1018.4%
5035.3%
10043.9%
25054.8%
50063.1%
1,00070.6%
2,00077.3%
3,00081.0%
5,00085.2%
8,00088.9%
10,00090.5%
15,00093.3%
20,00095.2%

Comfortable reading starts around 98%. Everything before that is still worth doing, but the first hundred words buy you a foothold, not comprehension.

The first 100 words of Czech

In order. The percentage is the share of ordinary speech you cover once you know that word and every word above it.

  1. 1to3.62%
  2. 2se6.07%
  3. 3je8.44%
  4. 4a10.46%
  5. 5že12.07%
  6. 6jsem13.6%
  7. 7na15.13%
  8. 8co16.43%
  9. 9v17.45%
  10. 10si18.44%
  11. 11tak19.28%
  12. 12ale20.06%
  13. 13ne20.82%
  14. 14s21.54%
  15. 15jsi22.16%
  16. 1622.77%
  17. 17mi23.37%
  18. 1823.95%
  19. 19o24.51%
  20. 20do25.02%
  21. 21jak25.52%
  22. 22ty25.98%
  23. 23z26.43%
  24. 24jo26.85%
  25. 25jste27.25%
  26. 2627.65%
  27. 27když28.04%
  28. 28jen28.42%
  29. 29jsme28.8%
  30. 30jako29.18%
  31. 31tu29.55%
  32. 32za29.92%
  33. 33ho30.28%
  34. 34ti30.63%
  35. 35by30.97%
  36. 36tady31.31%
  37. 37ano31.63%
  38. 38pro31.95%
  39. 39dobře32.27%
  40. 40není32.56%
  41. 41byl32.85%
  42. 4233.14%
  43. 43něco33.42%
  44. 44proč33.7%
  45. 45teď33.97%
  46. 46no34.23%
  47. 47k34.49%
  48. 48tam34.75%
  49. 49tohle35.01%
  50. 50mám35.27%
  51. 51jsou35.52%
  52. 52ten35.77%
  53. 53být36.01%
  54. 54toho36.24%
  55. 55bude36.47%
  56. 56bych36.68%
  57. 57vás36.9%
  58. 58nic37.11%
  59. 59takže37.31%
  60. 6037.51%
  61. 61ve37.71%
  62. 62vám37.91%
  63. 63nebo38.11%
  64. 64kdo38.3%
  65. 65aby38.49%
  66. 66tom38.68%
  67. 67byla38.87%
  68. 68ji39.05%
  69. 69bylo39.23%
  70. 70kde39.41%
  71. 71ještě39.58%
  72. 72i39.75%
  73. 73po39.93%
  74. 74pane40.09%
  75. 75protože40.26%
  76. 76než40.43%
  77. 77můj40.59%
  78. 78jeho40.75%
  79. 79od40.9%
  80. 80myslím41.06%
  81. 81tím41.22%
  82. 82jestli41.37%
  83. 83máš41.52%
  84. 84měl41.68%
  85. 85vím41.83%
  86. 86nás41.98%
  87. 87všechno42.13%
  88. 88možná42.28%
  89. 89víš42.42%
  90. 90nikdy42.57%
  91. 91prosím42.71%
  92. 92moc42.86%
  93. 93mu43.0%
  94. 94vy43.14%
  95. 95ani43.28%
  96. 96chci43.42%
  97. 97taky43.55%
  98. 98nevím43.67%
  99. 9943.79%
  100. 100pokud43.91%

Open all 5,000 words in the codex →

In the codex you tap the ones you already know and the number moves. It stays in your browser, no account.

Where Czech comes from

Proto-Indo-European Balto-Slavic Slavic West Slavic Czech

Spread very fast and very recently, which is why Slavic languages are still unusually close to each other.

See the whole family tree →

Words Czech has that English never built

  • lítost LEE-tostThe torment of suddenly seeing your own misery

All the untranslatable words →

Where the numbers come from

Frequencies counted over the OpenSubtitles 2018 dialogue corpus, 228,940,096 tokens of Czech in total. Lists compiled by Hermit Dave, released under CC BY-SA 3.0. Subtitles skew towards conversation, which is the point: this is the language people speak, not the language people publish.