The 5,000 most common Chinese (simplified) words
Ranked by how often they actually turn up in spoken dialogue, not by how useful somebody guessed they would be.
Every number below is counted from a corpus, and the corpus is named at the bottom.
What each step buys you
Learn the first 100 words of Chinese (simplified) and you follow 46.8% of everyday speech. The first 1000 take you to 72.3%. After that every thousand words buys less than the thousand before it, which is exactly why the order matters.
| Words known | Speech followed | |
|---|---|---|
| 10 | 23.1% | |
| 50 | 38.8% | |
| 100 | 46.8% | |
| 250 | 57.6% | |
| 500 | 65.4% | |
| 1,000 | 72.3% | |
| 2,000 | 78.8% | |
| 3,000 | 82.3% | |
| 5,000 | 86.5% | |
| 8,000 | 90.0% | |
| 10,000 | 91.6% | |
| 15,000 | 94.2% | |
| 20,000 | 95.8% |
Comfortable reading starts around 98%. Everything before that is still worth doing, but the first hundred words buy you a foothold, not comprehension.
The first 100 words of Chinese (simplified)
In order. The percentage is the share of ordinary speech you cover once you know that word and every word above it.
- 1的4.84%
- 2我9.33%
- 3你13.34%
- 4了15.96%
- 5是18.0%
- 6在19.25%
- 7他20.36%
- 8我们21.37%
- 9不22.26%
- 10吗23.06%
- 11好23.74%
- 12什么24.39%
- 13有25.02%
- 14就25.64%
- 15吧26.25%
- 16她26.85%
- 17这27.39%
- 18知道27.92%
- 19说28.43%
- 20都28.93%
- 21去29.39%
- 22很29.83%
- 23啊30.27%
- 24要30.7%
- 25也31.11%
- 26他们31.52%
- 27和31.91%
- 28那32.31%
- 29想32.7%
- 30人33.08%
- 31会33.47%
- 32一个33.82%
- 33来34.16%
- 34把34.47%
- 35对34.79%
- 36没有35.1%
- 37做35.42%
- 38不是35.73%
- 39让36.02%
- 40给36.3%
- 41you36.57%
- 42可以36.83%
- 43i37.09%
- 44你们37.34%
- 45呢37.59%
- 46但37.84%
- 47l38.08%
- 48到38.31%
- 49还38.55%
- 50现在38.78%
- 51the39.01%
- 52这个39.24%
- 53能39.46%
- 54怎么39.67%
- 55就是39.89%
- 56跟40.1%
- 57如果40.31%
- 58没40.52%
- 59它40.73%
- 60走40.93%
- 61上41.13%
- 62得41.33%
- 63被41.52%
- 64这里41.71%
- 65这样41.9%
- 66个42.08%
- 67看42.25%
- 68事42.42%
- 69to42.59%
- 70自己42.75%
- 71着42.91%
- 72谁43.07%
- 73过43.22%
- 74不会43.38%
- 75a43.53%
- 76告诉43.68%
- 77s43.83%
- 78为什么43.97%
- 79不能44.12%
- 80真的44.26%
- 81m44.4%
- 82先生44.55%
- 83只是44.69%
- 84不要44.83%
- 85这是44.97%
- 86已经45.1%
- 87再45.24%
- 88哦45.37%
- 89需要45.49%
- 90it45.62%
- 91因为45.75%
- 92谢谢45.88%
- 93像46.0%
- 94这么46.12%
- 95可能46.24%
- 96但是46.36%
- 97那个46.48%
- 98所以46.6%
- 99喜欢46.72%
- 100又46.83%
Chinese (simplified) does not separate words with spaces, so what is counted here are the segments the corpus splits on, not whole words. The order is still right and the coverage is still real, but read the list as building blocks rather than dictionary entries.
Open all 5,000 words in the codex →
In the codex you tap the ones you already know and the number moves. It stays in your browser, no account.
Where Chinese (simplified) comes from
Sino-Tibetan › Chinese
Tone does the work that endings do in Indo-European.
Same job, opposite tool.
Words Chinese (simplified) has that English never built
- 缘分 yuen-FENThe pull that brings two people together, or does not
Where the numbers come from
Frequencies counted over the OpenSubtitles 2018 dialogue corpus, 81,775,831 tokens of Chinese (simplified) in total. Lists compiled by Hermit Dave, released under CC BY-SA 3.0. Subtitles skew towards conversation, which is the point: this is the language people speak, not the language people publish.