vernacula-phonemizer data
The data tree read at runtime by vernacula-phonemizer
(the TypeScript, C# and Rust engines): lexicons, rule tables, manifests and int8 ONNX models. A data key
such as languages/english/g2p-dict.tsv is a path in this repo.
This revision is tag 224339e8, built from phonemizer commit 224339e897c3cb5c38dc602bc48840d8546bad38. Pin the same tag as the engine's
git revision.
Licensing
There is no single license. The project's own work is MIT (LICENSE). Third-party-derived
files keep their parent license (CC0, CC-BY, CC-BY-SA, GPL-3.0 and two bespoke licenses), declared per file
in LICENSES/PROVENANCE.md and the *.PROVENANCE.md sidecars. The full license
texts are in LICENSES/, and the attribution roll-up, including attributions that are
mandatory, is NOTICE.md. Redistributing these files means honouring those per-file terms.
Attributions required by name
This package uses the JMdict/EDICT and KANJIDIC dictionary files. These files are the property of the Electronic Dictionary Research and Development Group, and are used in conformance with the Group's licence.
(languages/japanese/readings.tsv, fallback.tsv, adverbs.txt; © EDRDG, CC-BY-SA 4.0,
https://www.edrdg.org/edrdg/licence.html.)
The Sindhi Open Lexicon: SindhiLanguage.org (https://sindhilanguage.org/), prepared and curated by
Amar Fayaz Buriro (امر فياض ٻرڙو). Behind languages/sindhi/sindhi-lexicon.tsv and
languages/sindhi/sd-g2p-tagger.int8.onnx; terms in LICENSES/LicenseRef-SindhiOpenLexicon.txt.