vernacula-phonemizer data

The data tree read at runtime by vernacula-phonemizer (the TypeScript, C# and Rust engines): lexicons, rule tables, manifests and int8 ONNX models. A data key such as languages/english/g2p-dict.tsv is a path in this repo.

This revision is tag 224339e8, built from phonemizer commit 224339e897c3cb5c38dc602bc48840d8546bad38. Pin the same tag as the engine's git revision.

Licensing

There is no single license. The project's own work is MIT (LICENSE). Third-party-derived files keep their parent license (CC0, CC-BY, CC-BY-SA, GPL-3.0 and two bespoke licenses), declared per file in LICENSES/PROVENANCE.md and the *.PROVENANCE.md sidecars. The full license texts are in LICENSES/, and the attribution roll-up, including attributions that are mandatory, is NOTICE.md. Redistributing these files means honouring those per-file terms.

Attributions required by name

This package uses the JMdict/EDICT and KANJIDIC dictionary files. These files are the property of the Electronic Dictionary Research and Development Group, and are used in conformance with the Group's licence.

(languages/japanese/readings.tsv, fallback.tsv, adverbs.txt; © EDRDG, CC-BY-SA 4.0, https://www.edrdg.org/edrdg/licence.html.)

The Sindhi Open Lexicon: SindhiLanguage.org (https://sindhilanguage.org/), prepared and curated by Amar Fayaz Buriro (امر فياض ٻرڙو). Behind languages/sindhi/sindhi-lexicon.tsv and languages/sindhi/sd-g2p-tagger.int8.onnx; terms in LICENSES/LicenseRef-SindhiOpenLexicon.txt.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support