🇱🇻 Latvian 2 million speakers

The 5,000 most common Latvian words

Ranked by how often they actually turn up in spoken dialogue, not by how useful somebody guessed they would be.
Every number below is counted from a corpus, and the corpus is named at the bottom.

What each step buys you

Learn the first 100 words of Latvian and you follow 44.1% of everyday speech. The first 1000 take you to 66.9%. After that every thousand words buys less than the thousand before it, which is exactly why the order matters.

Words knownSpeech followed
1015.7%
5036.1%
10044.1%
25053.8%
50060.5%
1,00066.9%
2,00073.3%
3,00077.0%
5,00081.7%
8,00086.0%
10,00088.0%
15,00091.4%
20,00093.7%

Comfortable reading starts around 98%. Everything before that is still worth doing, but the first hundred words buy you a foothold, not comprehension.

The first 100 words of Latvian

In order. The percentage is the share of ordinary speech you cover once you know that word and every word above it.

  1. 1ir2.83%
  2. 2es5.49%
  3. 3tu7.34%
  4. 4un8.97%
  5. 5tas10.2%
  6. 6ka11.4%
  7. 7man12.53%
  8. 8to13.65%
  9. 9vai14.67%
  10. 10ko15.68%
  11. 11ar16.64%
  12. 12par17.54%
  13. 13kas18.42%
  14. 1419.28%
  15. 1520.11%
  16. 16viņš20.92%
  17. 17tev21.66%
  18. 1822.38%
  19. 19uz23.07%
  20. 20no23.75%
  21. 21mēs24.43%
  22. 22nav25.08%
  23. 23labi25.72%
  24. 2426.36%
  25. 25bet26.93%
  26. 26jūs27.49%
  27. 27lai28.03%
  28. 28ja28.53%
  29. 29viņa29.02%
  30. 30mani29.5%
  31. 31bija29.98%
  32. 32viņu30.43%
  33. 33esmu30.87%
  34. 34esi31.3%
  35. 35tevi31.72%
  36. 36mums32.14%
  37. 37ne32.47%
  38. 38tikai32.8%
  39. 39tad33.12%
  40. 40viņi33.43%
  41. 41kad33.74%
  42. 42jums34.04%
  43. 43arī34.33%
  44. 44viss34.62%
  45. 45pie34.9%
  46. 46kur35.17%
  47. 47nu35.42%
  48. 48vēl35.66%
  49. 49tik35.9%
  50. 50jau36.14%
  51. 51tur36.37%
  52. 52te36.61%
  53. 53kaut36.84%
  54. 54būs37.07%
  55. 55paldies37.29%
  56. 56visu37.51%
  57. 57viņam37.73%
  58. 58tagad37.95%
  59. 59pēc38.17%
  60. 60šeit38.38%
  61. 61kāpēc38.58%
  62. 62būtu38.77%
  63. 63ļoti38.96%
  64. 64savu39.14%
  65. 65mūsu39.32%
  66. 66lūdzu39.49%
  67. 67taču39.66%
  68. 68šo39.82%
  69. 69gan39.98%
  70. 70mans40.14%
  71. 71ak40.29%
  72. 72kāds40.45%
  73. 73jūsu40.6%
  74. 74zinu40.75%
  75. 75daudz40.9%
  76. 76jo41.05%
  77. 77varbūt41.19%
  78. 78esam41.34%
  79. 79kungs41.48%
  80. 80cik41.63%
  81. 81zini41.77%
  82. 82tās41.91%
  83. 83mana42.04%
  84. 84tiešām42.17%
  85. 85esat42.3%
  86. 86domāju42.43%
  87. 87pa42.55%
  88. 88var42.68%
  89. 89varētu42.81%
  90. 90būt42.93%
  91. 91tam43.05%
  92. 92tāpēc43.18%
  93. 93visi43.3%
  94. 94nezinu43.42%
  95. 95šī43.54%
  96. 96nekad43.65%
  97. 97viens43.77%
  98. 98līdz43.89%
  99. 99pat44.0%
  100. 100aiziet44.12%

Open all 5,000 words in the codex →

In the codex you tap the ones you already know and the number moves. It stays in your browser, no account.

Where Latvian comes from

Proto-Indo-European Balto-Slavic Baltic Latvian

Two branches that kept more of the original grammar than almost anyone else, cases included.

See the whole family tree →

Where the numbers come from

Frequencies counted over the OpenSubtitles 2018 dialogue corpus, 2,271,911 tokens of Latvian in total. Lists compiled by Hermit Dave, released under CC BY-SA 3.0. Subtitles skew towards conversation, which is the point: this is the language people speak, not the language people publish.