← 漢字星圖 Hanzi Constellation

Credits & data sources

Last reviewed 2026-08-20

Hanzi Constellation is built on open Chinese-language data that other people spent years assembling. The map is ours; most of the facts on it are not. Every source we draw from is listed below with its license and what we did to it.

CNS11643 全字庫

Taiwan’s national character database, published by the Ministry of Digital Affairs. It supplies the reading on every character: 注音 as recorded by Taiwan’s own standard, which is why 危 reads wéi here and 髮 reads , rather than the mainland readings most dictionaries carry. Used under the Open Government Data License, version 1.0, from 全字庫.

Changes we made: we converted 注音 to pinyin. 全字庫 lists every reading a character takes, unordered; we use it to decide which readings are Taiwan’s, and choose the primary one ourselves, from how often a character takes each reading in everyday Taiwan words. Where a character has several readings, the card shows them all, Taiwan’s first, with the mainland reading marked where the two standards differ (期 · ). Those orderings are ours, not the Ministry’s, and a few are hand-set (長 cháng before zhǎng).

Words too: the pinyin on a word comes from CC-CEDICT below, which follows mainland convention. Where CC-CEDICT records a Taiwan reading for a character, we apply it to every word the character appears in (星期 xīngqí, 企業 qìyè); where it records one for the word itself we use that (垃圾 lèsè); and 129 words were additionally reviewed by hand against the Ministry of Education’s dictionary and everyday Taiwan usage (危險 wéixiǎn). Those judgments are ours.

TBCL 漢字表 (NAER)

The character list of the 臺灣華語文能力基準 (Taiwan Benchmarks for the Chinese Language), published by Taiwan’s National Academy for Educational Research (國家教育研究院). It supplies the proficiency band (1–7) shown on characters — the official Taiwan banding that the TOCFL exam is aligned with. Used with the Academy’s written permission (教研語譯字第1150001498號, 27 August 2026), which covers research, educational and commercial use with attribution. The data is provided as-is, with no warranty of completeness or accuracy.

Changes we made — and what they mean: we reuse only the factual character→band mapping, not the document itself, and we changed it in two ways. We ship a subset: only the banded characters that appear in our map, so about a sixth of the Academy’s list is absent. And where the list records glyph variants in one row (裡/裏), we band each variant separately. Characters the list does not teach — mostly transliteration and colloquial characters — carry no band rather than a guessed one. We never change a band itself.

Because the data is modified, this app is not a TBCL-certified product and the bands shown here should not be read as the Academy’s own published classification of our particular selection. The Academy neither published nor endorsed this app.

The Unicode Consortium’s database of Han characters. We take character definitions, traditional/simplified variant links, and readings for the handful of characters 全字庫 does not cover. Used under the Unicode License.

CC-CEDICT

A community-maintained Chinese–English dictionary. It supplies our word list, the English glosses for both words and single characters, and the register tags that drive the family-friendly filter. Used under CC BY-SA 4.0 via MDBG.

Changes we made: we selected a subset of entries, re-keyed them to traditional characters, joined multiple senses into a single gloss string (and replaced the gloss outright on about 300 common characters with our own learner wording), ranked them by corpus frequency, derived per-word register flags from the original (vulgar), (taboo) and (offensive) tags, and turned its “Taiwan pr.” and “also pr.” pronunciation notes into the readings shown on characters and words instead of leaving them in the gloss text.

FrequencyWords (OpenSubtitles)

Word-frequency counts derived from the OpenSubtitles corpus, published by Hermit Dave. They determine which 3,000 characters make up the base map and the order of the star charts. Used under CC BY-SA 4.0 from the FrequencyWords project.

Changes we made: we used the traditional-script (zh_tw) list, summed word counts onto their constituent characters to produce character frequencies, and discarded entries outside our base set. We also reviewed and excluded characters whose counts trace to corrupted subtitle files rather than real usage — every character we teach was verified against Taiwan’s own references, and coverage figures count real text only.

Character decompositions

Which components a character is built from comes from the CJK Decomposition Data file, originally compiled by Gavin Grover, used under the MIT license (one of the six its author offers).

Changes we made: we keep only the top-level components of each character and discard the relation structure around them; unnamed intermediate shapes are resolved to the real components beneath them; simplified-form component spellings are mapped to their traditional forms; and where the Dong Chinese wiki (below) identifies a phonetic component that the visual decomposition renders in a transformed shape, that component is added so the family stays connected. The result is our own arrangement, not a copy of the source file.

Dong Chinese character wiki

Which component carries a character’s sound — the head of each phonetic family on the map — comes, wherever the wiki has an opinion, from the Chinese Character Wiki by Dong Chinese, used under CC BY-SA 4.0. The same wiki also supplies the membership of a phonetic head in its characters where the visual decomposition obscures it, and served as the independent reference when we replaced our decomposition source in August 2026.

Changes we made: we take one fact per character — which component the wiki labels as the sound component — for the characters in our set, and nothing else: no glosses, origins, images or stroke data. Where the wiki names several, we keep the one that is also in our decomposition; where it names a glyph outside our set, or has no entry, our own inference stands. The confidence percentage on every link is our computation.

Our own work

The phonetic-series links are ours: no source we use marks which component supplies a character’s sound, so we infer it by comparing readings — a component counts only when it shares the character’s syllable, or its final together with a related initial (河 hé from 可 kě), never on a rhyme coincidence alone — and some of those inferences will still be wrong. That is why every phonetic link carries a confidence percentage — scored from how closely the readings agree and how reliably the component predicts the sound of its whole family — rather than being presented as settled fact, and why links we cannot score above 50% are not shown at all. Wherever the Dong Chinese character wiki (CC BY-SA 4.0) has an opinion on which component carries a character’s sound, we use its identification instead of our own inference — that single fact per character is the only thing we take from it, its glosses, origins and images are not copied, and the confidence percentage is still ours. Each link records which of the two identified it. The star charts, their ordering, the coverage figures, the application and its design are also ours. Example sentences, when they arrive, will be AI-generated and labeled as such — none ship today.

Sharing our data back

Because our word data and our phonetic-family identifications adapt CC BY-SA 4.0 material, that adapted data is itself offered under CC BY-SA 4.0. Write to support@hanzi-constellation.app and we will send it to you. The application itself is not covered by that license.

Corrections

If we have credited something incorrectly, or used something we should not have, please tell us at support@hanzi-constellation.app and we will fix it.