๐Ÿ‡ฏ๐Ÿ‡ต Japanese 125 million speakers

The 5,000 most common Japanese words

Ranked by how often they actually turn up in spoken dialogue, not by how useful somebody guessed they would be.
Every number below is counted from a corpus, and the corpus is named at the bottom.

What each step buys you

Learn the first 100 words of Japanese and you follow 67.0% of everyday speech. The first 1000 take you to 84.8%. After that every thousand words buys less than the thousand before it, which is exactly why the order matters.

Words knownSpeech followed
1031.8%
5059.2%
10067.0%
25074.8%
50080.0%
1,00084.8%
2,00089.1%
3,00091.4%
5,00094.0%
8,00096.1%
10,00096.9%
15,00098.1%
20,00098.7%

Comfortable reading starts around 98%. Everything before that is still worth doing, but the first hundred words buy you a foothold, not comprehension.

The first 100 words of Japanese

In order. The percentage is the share of ordinary speech you cover once you know that word and every word above it.

  1. 1ใ„4.25%
  2. 2ใฎ8.25%
  3. 3ใฏ11.68%
  4. 4ใฆ14.97%
  5. 5ใŸ18.09%
  6. 6ใซ21.12%
  7. 7ใช24.05%
  8. 8ใ‚’26.73%
  9. 9ใ‚‹29.29%
  10. 10ใ 31.79%
  11. 11ใŒ34.09%
  12. 12ใง36.14%
  13. 13ใ—38.05%
  14. 14ใฃ39.94%
  15. 15ใ‹41.28%
  16. 16ใจ42.55%
  17. 17ใ‚‚43.64%
  18. 18ใ‚“44.66%
  19. 19ใ™45.66%
  20. 20ใ‚ˆ46.65%
  21. 21ใ†47.4%
  22. 22ใพ48.1%
  23. 23็ง48.73%
  24. 24ใ‚‰49.35%
  25. 25ใ49.96%
  26. 26ไฝ•50.55%
  27. 27ใ‚51.07%
  28. 28ใ‹ใ‚‰51.58%
  29. 29ใ‚52.04%
  30. 30ใ‚Œ52.49%
  31. 31ใญ52.92%
  32. 32ใ™ใ‚‹53.33%
  33. 33ใใ‚Œ53.74%
  34. 34ๅฝผ54.14%
  35. 35ใ•54.54%
  36. 36ใ‚54.92%
  37. 37ใ“ใจ55.31%
  38. 38ใใ†55.65%
  39. 39ใ‚ใชใŸ56.0%
  40. 40ใŠ56.32%
  41. 41ใฉใ†56.64%
  42. 42ๅ›56.95%
  43. 43ใ›57.26%
  44. 44่จ€57.56%
  45. 45ไฟบ57.85%
  46. 46ใ‚„58.14%
  47. 47ใ˜ใ‚ƒ58.41%
  48. 48ใ‚Š58.67%
  49. 49ใ‹ใฃ58.93%
  50. 50่กŒ59.18%
  51. 51ใ‚ˆใ†59.42%
  52. 52ใใ‚Œ59.65%
  53. 53ใ“ใ‚Œ59.88%
  54. 54ใ“ใฎ60.1%
  55. 55ไบ‹60.31%
  56. 56ใ“ใ“60.52%
  57. 57ไบบ60.73%
  58. 58็Ÿฅ60.93%
  59. 59ใ61.13%
  60. 60ใฐ61.33%
  61. 61ใใฎ61.53%
  62. 62่ชฐ61.73%
  63. 63ๅฝผๅฅณ61.93%
  64. 64ๆ€62.12%
  65. 65่ฆ‹62.31%
  66. 66ใฃใฆ62.5%
  67. 67ใŠๅ‰62.69%
  68. 68่ฉฑ62.88%
  69. 69ใ ใฃ63.05%
  70. 70ใค63.23%
  71. 71ใž63.41%
  72. 72ใ•ใ‚“63.58%
  73. 73ใ‚‚ใ†63.74%
  74. 74ๆฅ63.89%
  75. 75ๅƒ•64.03%
  76. 76ใ—ใ‚‡64.18%
  77. 77ใงใ64.32%
  78. 78ใ ใ‘64.47%
  79. 79ใชใ‚‰64.61%
  80. 80่€…64.75%
  81. 81ใ‚ใ‹64.89%
  82. 82ไปŠ65.03%
  83. 83ๅˆ†ใ‹65.17%
  84. 84ใŸใก65.29%
  85. 85่ž65.41%
  86. 86ๅ‡บ65.53%
  87. 87ใพใง65.65%
  88. 88ใ‘ใฉ65.76%
  89. 89ใฉใ“65.88%
  90. 90ใ‚‰ใ‚Œ65.99%
  91. 91ๆญป66.1%
  92. 92ใธ66.22%
  93. 93ๆ™‚66.33%
  94. 94ๆฎบ66.43%
  95. 95้”66.54%
  96. 96ใ‚ใ‚66.65%
  97. 97ๅ‰66.75%
  98. 98ไธญ66.86%
  99. 99ๅฟ…่ฆ66.95%
  100. 100ใฟ67.05%

Japanese does not separate words with spaces, so what is counted here are the segments the corpus splits on, not whole words. The order is still right and the coverage is still real, but read the list as building blocks rather than dictionary entries.

Open all 5,000 words in the codex →

In the codex you tap the ones you already know and the number moves. It stays in your browser, no account.

Where Japanese comes from

Japonic Japanese

One family, essentially one language.
Its relationship to anything else is still unsettled.

See the whole family tree →

Words Japanese has that English never built

  • ๆœจๆผใ‚Œๆ—ฅ ko-mo-REH-beeSunlight filtering through the leaves of trees
  • ็ฉ่ชญ tsoon-DOH-kooBuying books and letting them pile up unread
  • ็‰ฉใฎๅ“€ใ‚Œ MO-no no a-WA-rehThe gentle sadness of things, because they end
  • ๆฃฎๆž—ๆตด shin-rin-YO-kooBathing in the forest, as a treatment
  • ็”Ÿใ็”ฒๆ– ee-kee-GUYThe reason you get up, small enough to be true
  • ไพ˜ๅฏ‚ WAH-bee SAH-beeThe beauty of what is worn, uneven and passing
  • ้‡‘็ถ™ใŽ kin-TSOO-gheeRepairing with gold, so the break shows

All the untranslatable words →

Where the numbers come from

Frequencies counted over the OpenSubtitles 2018 dialogue corpus, 13,228,940 tokens of Japanese in total. Lists compiled by Hermit Dave, released under CC BY-SA 3.0. Subtitles skew towards conversation, which is the point: this is the language people speak, not the language people publish.