๐Ÿ‡ฌ๐Ÿ‡ช Georgian 4 million speakers

The 5,000 most common Georgian words

Ranked by how often they actually turn up in spoken dialogue, not by how useful somebody guessed they would be.
Every number below is counted from a corpus, and the corpus is named at the bottom.

What each step buys you

Learn the first 100 words of Georgian and you follow 46.5% of everyday speech. The first 1000 take you to 68.7%. After that every thousand words buys less than the thousand before it, which is exactly why the order matters.

Words knownSpeech followed
1021.2%
5039.3%
10046.5%
25055.8%
50062.3%
1,00068.7%
2,00075.2%
3,00078.9%
5,00083.4%
8,00087.4%
10,00089.2%
15,00092.3%
20,00094.4%

Comfortable reading starts around 98%. Everything before that is still worth doing, but the first hundred words buy you a foothold, not comprehension.

The first 100 words of Georgian

In order. The percentage is the share of ordinary speech you cover once you know that word and every word above it.

  1. 1แƒ”แƒ4.3%
  2. 2แƒ•7.56%
  3. 3แƒœแƒ•10.39%
  4. 4แƒŸแƒ•12.68%
  5. 5แƒ—14.53%
  6. 6แƒœแƒ16.36%
  7. 7แƒŸแƒ—17.69%
  8. 8แƒฑแƒ19.0%
  9. 9แƒฆแƒ•20.09%
  10. 10แƒ แƒ—21.17%
  11. 11แƒ“แƒฒ22.24%
  12. 12แƒšแƒ—23.27%
  13. 13แƒ24.19%
  14. 14แƒ’25.09%
  15. 15แƒ›แƒ—25.98%
  16. 16แƒคแƒ•26.81%
  17. 17แƒ แƒฒแƒ’แƒ27.64%
  18. 18แƒฒแƒ 28.39%
  19. 19แƒŸ29.06%
  20. 20แƒ™แƒแƒ™แƒ’แƒฒ29.69%
  21. 21แƒ แƒ•30.22%
  22. 22แƒœแƒฒ30.73%
  23. 23แƒ›แƒ•31.19%
  24. 24แƒŸแƒงแƒ›31.62%
  25. 25แƒแƒฑ31.99%
  26. 26แƒแƒ™แƒฒ32.35%
  27. 27แƒฅแƒ แƒฒ32.69%
  28. 28แƒ”แƒฒแƒ‘แƒžแƒ•33.04%
  29. 29แƒ’แƒฒ33.38%
  30. 30แƒ แƒžแƒฌแƒ‘แƒ’แƒ33.72%
  31. 31แƒŸแƒ34.04%
  32. 32แƒœแƒฌแƒ›แƒ34.36%
  33. 33แƒ›แƒฒแƒ–แƒ•34.67%
  34. 34แƒ›แƒก34.97%
  35. 35แƒœแƒ—35.27%
  36. 36แƒŸแƒแƒ›แƒฒ35.57%
  37. 37แƒ’แƒ—35.86%
  38. 38แƒแƒฒ36.15%
  39. 39แƒฌ36.44%
  40. 40แƒ‘แƒ•แƒฅแƒ•36.72%
  41. 41แƒ“แƒ—37.01%
  42. 42แƒ แƒแƒ™แƒ37.3%
  43. 43แƒ แƒฒแƒ37.58%
  44. 44แƒ—แƒ›แƒ37.86%
  45. 45แƒ แƒกแƒ™38.13%
  46. 46แƒฑแƒœแƒแƒ›38.37%
  47. 47แƒ›แƒœแƒฒแƒ“แƒฒ38.6%
  48. 48แƒฒแƒ”38.83%
  49. 49แƒ”แƒฒ39.06%
  50. 50แƒ”แƒ•แƒ™แƒ39.29%
  51. 51แƒ™แƒแƒ แƒฒ39.51%
  52. 52แƒŸแƒฒ39.73%
  53. 53แƒŸแƒ•แƒ“แƒ39.95%
  54. 54แƒ™แƒแƒ™40.16%
  55. 55แƒ แƒฒแƒ˜40.37%
  56. 56แƒ‘แƒ—40.58%
  57. 57แƒœแƒ•แƒฆแƒฒ40.78%
  58. 58แƒฑแƒแƒฆแƒฒ40.96%
  59. 59แƒก41.14%
  60. 60แƒ›แƒ•แƒœ41.31%
  61. 61แƒ™แƒแƒ™แƒฒ41.48%
  62. 62แƒ แƒ•แƒ‘41.65%
  63. 63แƒ’แƒŸแƒ—แƒคแƒ™แƒฒ41.81%
  64. 64แƒŸแƒ แƒ•41.98%
  65. 65แƒŸแƒ›แƒ•42.14%
  66. 66แƒ แƒฒ42.3%
  67. 67แƒฑแƒœแƒแƒ•แƒฅ42.45%
  68. 68แƒ—แƒšแƒ—42.6%
  69. 69แƒ—แƒŸแƒ™แƒแƒ›42.75%
  70. 70แƒ”แƒฒแƒ‘แƒžแƒฒ42.89%
  71. 71แƒฒแƒฆแƒ•43.03%
  72. 72แƒ™แƒงแƒ”แƒ•43.17%
  73. 73แƒแƒŸ43.31%
  74. 74แƒ›แƒฒแƒ“แƒ43.44%
  75. 75แƒ™แƒฒแƒ“แƒแƒ แƒฒ43.57%
  76. 76แƒ‘แƒšแƒแƒ“แƒฒแƒ”แƒแƒžแƒฌ43.7%
  77. 77แƒ—แƒ›แƒแƒ›43.83%
  78. 78แƒœแƒแƒšแƒ—43.96%
  79. 79แƒ’แƒžแƒ•แƒ›แƒ•44.09%
  80. 80แƒ™แƒฒ44.22%
  81. 81แƒœแƒ•แƒ“แƒฒ44.35%
  82. 82แƒฒ44.48%
  83. 83แƒ’แƒŸแƒ—แƒคแƒ™แƒ—44.6%
  84. 84แƒแƒžแƒฒแƒŸแƒ แƒฒ44.73%
  85. 85แƒŸแƒกแƒ›44.85%
  86. 86แƒ แƒฒแƒฑแƒ—44.97%
  87. 87แƒ แƒแƒ›45.09%
  88. 88แƒฒแƒ’แƒ45.21%
  89. 89แƒŸแƒแƒ›45.33%
  90. 90แƒแƒ แƒ45.44%
  91. 91แƒ™แƒฒแƒ•แƒ แƒฒ45.56%
  92. 92แƒ›แƒฒ45.67%
  93. 93แƒแƒžแƒ—45.78%
  94. 94แƒ แƒฌ45.9%
  95. 95แƒ แƒแƒฑแƒ—46.0%
  96. 96แƒ•แƒ”แƒ—แƒœ46.11%
  97. 97แƒ›แƒฒแƒ–แƒ•แƒฅ46.22%
  98. 98แƒ™แƒฒแƒ˜46.33%
  99. 99แƒœแƒ—แƒฆแƒฒ46.43%
  100. 100แƒฑแƒแƒฆแƒฒแƒ แƒฒ46.53%

Open all 5,000 words in the codex →

In the codex you tap the ones you already know and the number moves. It stays in your browser, no account.

Where Georgian comes from

Kartvelian Georgian

Georgian and three relatives, in the Caucasus, related to nothing outside it.

See the whole family tree →

Words Georgian has that English never built

  • แƒจแƒ”แƒ›แƒแƒ›แƒ”แƒญแƒแƒ›แƒ she-mo-me-JA-moI accidentally ate the whole thing

All the untranslatable words →

Where the numbers come from

Frequencies counted over the OpenSubtitles 2018 dialogue corpus, 1,090,991 tokens of Georgian in total. Lists compiled by Hermit Dave, released under CC BY-SA 3.0. Subtitles skew towards conversation, which is the point: this is the language people speak, not the language people publish.