The 5,000 most common Bulgarian words
Ranked by how often they actually turn up in spoken dialogue, not by how useful somebody guessed they would be.
Every number below is counted from a corpus, and the corpus is named at the bottom.
What each step buys you
Learn the first 100 words of Bulgarian and you follow 52.2% of everyday speech. The first 1000 take you to 74.0%. After that every thousand words buys less than the thousand before it, which is exactly why the order matters.
| Words known | Speech followed | |
|---|---|---|
| 10 | 22.7% | |
| 50 | 44.0% | |
| 100 | 52.2% | |
| 250 | 61.6% | |
| 500 | 67.9% | |
| 1,000 | 74.0% | |
| 2,000 | 79.8% | |
| 3,000 | 83.1% | |
| 5,000 | 87.0% | |
| 8,000 | 90.5% | |
| 10,000 | 92.0% | |
| 15,000 | 94.5% | |
| 20,000 | 96.1% |
Comfortable reading starts around 98%. Everything before that is still worth doing, but the first hundred words buy you a foothold, not comprehension.
The first 100 words of Bulgarian
In order. The percentage is the share of ordinary speech you cover once you know that word and every word above it.
- 1да4.66%
- 2е7.73%
- 3не10.62%
- 4се12.94%
- 5и14.87%
- 6на16.77%
- 7си18.36%
- 8за19.87%
- 9ще21.35%
- 10че22.67%
- 11това23.97%
- 12ли25.24%
- 13в26.51%
- 14ти27.6%
- 15от28.64%
- 16ми29.59%
- 17с30.52%
- 18какво31.4%
- 19го32.22%
- 20съм32.87%
- 21но33.49%
- 22аз34.07%
- 23те34.61%
- 24трябва35.1%
- 25добре35.58%
- 26ме36.05%
- 27са36.51%
- 28няма36.95%
- 29може37.36%
- 30ако37.77%
- 31тук38.17%
- 32така38.53%
- 33а38.89%
- 34ви39.25%
- 35като39.6%
- 36много39.95%
- 37той40.28%
- 38по40.61%
- 39има40.93%
- 40беше41.25%
- 41нещо41.55%
- 42му41.86%
- 43как42.15%
- 44защо42.43%
- 45ни42.71%
- 46я42.99%
- 47знам43.25%
- 48само43.5%
- 49всичко43.75%
- 50мен44.0%
- 51сега44.25%
- 52теб44.48%
- 53тя44.7%
- 54което44.91%
- 55искам45.12%
- 56мога45.33%
- 57още45.54%
- 58когато45.75%
- 59до45.95%
- 60ги46.15%
- 61този46.34%
- 62просто46.54%
- 63сте46.73%
- 64или46.92%
- 65къде47.1%
- 66би47.28%
- 67тази47.46%
- 68всички47.64%
- 69благодаря47.81%
- 70нали47.99%
- 71сме48.16%
- 72нищо48.33%
- 73толкова48.49%
- 74знаеш48.65%
- 75един48.81%
- 76защото48.97%
- 77след49.12%
- 78там49.28%
- 79време49.43%
- 80който49.58%
- 81искаш49.72%
- 82мисля49.87%
- 83кой50.01%
- 84малко50.16%
- 85имам50.3%
- 86преди50.43%
- 87хайде50.57%
- 88вече50.7%
- 89тези50.83%
- 90моля50.96%
- 91о51.09%
- 92някой51.22%
- 93каза51.35%
- 94него51.48%
- 95наистина51.6%
- 96ние51.72%
- 97вие51.84%
- 98които51.96%
- 99й52.07%
- 100повече52.18%
Open all 5,000 words in the codex →
In the codex you tap the ones you already know and the number moves. It stays in your browser, no account.
Where Bulgarian comes from
Proto-Indo-European › Balto-Slavic › Slavic › South Slavic › Bulgarian
Cut off from the rest by Hungarian and Romanian arriving in between.
Where the numbers come from
Frequencies counted over the OpenSubtitles 2018 dialogue corpus, 236,068,393 tokens of Bulgarian in total. Lists compiled by Hermit Dave, released under CC BY-SA 3.0. Subtitles skew towards conversation, which is the point: this is the language people speak, not the language people publish.