13 KiB
Getting a Persian pipeline into the spaCy ecosystem
Sourced from spaCy's primary docs and repos, URLs inline. Where a contribution goes decides how it has to be shaped, so read this before writing code.
0. Two separate contribution surfaces
| Surface | What it is | Where it lives | How you contribute |
|---|---|---|---|
Language data (fa) |
Hand-written rules: tokenizer exceptions, stop words, LIKE_NUM, punctuation, noun-chunk iterator |
spacy/lang/fa/*.py inside the spaCy repo |
Normal PR to explosion/spaCy |
Trained pipeline (fa_dep_news_sm) |
Statistical weights + config.cfg + meta.json, shipped as a pip wheel |
explosion/spacy-models releases |
You cannot. Publish it yourself (PyPI or HF Hub) and get it listed in spaCy Universe |
spacy/lang/fa already exists upstream. No trained fa pipeline does. So this is a publishing
project, with optional upstream PRs for the language-data gaps found along the way.
1. Policy, verbatim
From https://github.com/explosion/spaCy/blob/master/CONTRIBUTING.md:
- No CLA. There is no contributor licence agreement section; contribution is an ordinary GitHub PR governed by the Contributor Covenant Code of Conduct v1.4.
- Inclusion philosophy:
"Our philosophy is to prefer a smaller core library. […] If you're looking to implement a new spaCy feature, starting with a custom component package is usually the best strategy. […] And if it works well, we can always integrate it into the core library later."
- The sanctioned route for non-trivial additions is the "Publishing spaCy extensions and
plugins" section:
"An extension or plugin should add substantial functionality, be well-documented and open-source. It should be available for users to download and install as a Python package – for example via PyPi." "Once your extension is published, you can open a PR to suggest it for the Universe page."
CONTRIBUTING.mdnever mentions submitting trained models. Neither does https://github.com/explosion/spacy-models, whose README only documentscompatibility.jsonas "the source of spaCy's internal compatibility check, performed when you run the download command". There is no PR template, issue label, or documented process for a third party to land a pipeline inspacy download. So thespacy downloadindex is an Explosion-only release channel. ([INFERENCE]from the absence of any process across all three sources.)https://spacy.io/usage/adding-languagesno longer exists and redirects to https://spacy.io/usage/linguistic-features#language-data. TheBaseDefaultsandLanguagecontract is at https://spacy.io/api/language.
Code rules for a spacy/lang/fa PR
blackformatting,flake8clean, type hints where practical.- Tests go in
spacy/tests/lang/fa/; regression tests use@pytest.mark.issue(N). - Touching
.pyxmeans the reviewer must rebuild:python setup.py build_ext --inplace. - Language data must be data and rules only: no model downloads, no network, no heavy deps.
2. Three publishing routes for the trained pipeline
- Hugging Face Hub, using Explosion's own tool
(https://github.com/explosion/spacy-huggingface-hub):
Users thenpip install spacy-huggingface-hub huggingface-cli login python -m spacy package training/fa_dep_news_sm packages --name dep_news_sm --version 3.8.0 --build wheel python -m spacy huggingface-hub push packages/fa_dep_news_sm-3.8.0/dist/fa_dep_news_sm-3.8.0-py3-none-any.whl --org <org>pip install https://huggingface.co/<org>/fa_dep_news_sm/resolve/main/fa_dep_news_sm-any-py3-none-any.whl. - PyPI or a self-hosted wheel:
spacy package … --build sdist,wheelthentwine upload, or attach the wheel to a GitHub Release. See https://spacy.io/api/cli#package. - spaCy Universe, which lists the package on spacy.io but hosts nothing. Per
https://github.com/explosion/spaCy/blob/master/website/UNIVERSE.md: fork
explosion/spaCy, append an entry towebsite/meta/universe.json, open a PR. Required fields:id,title,slogan,description,github,pip,code_example,code_language,url,thumb,image,author,author_links,category,tags. The package must be open-source with a user-friendly license and "at least somewhat documented".
Route taken here: HF Hub for the artifact, plus a Universe PR once metrics are respectable.
3. Naming and versioning
From https://spacy.io/models#conventions and the spacy-models README:
[lang]_[type]_[genre]_[size] e.g. fa_dep_news_sm
| Slot | Allowed values | Meaning |
|---|---|---|
type |
core |
tagger + parser + lemmatizer + NER |
dep |
tagger + parser + lemmatizer, no NER | |
ent |
NER only | |
sent |
sentence segmentation only | |
genre |
news, web, wiki |
text domain of the training corpus |
size |
sm |
no static vectors |
md |
~20k unique vectors / ~500k keys (or 50k floret rows) | |
lg |
~500k vectors (or 200k floret rows) | |
trf |
transformer, no static vectors |
Version numbers are a.b.c = spaCy-major, spaCy-minor, model-revision. Trained against spaCy
3.8, so the package version starts at 3.8.0; retraining on new data bumps c. Avoid
0.1.0-style versions, which is how hazm's HF pipelines ended up with version: 0.0.0 and
spacy_version: >=3.6.0,<3.7.0.
meta.json fields that matter (https://spacy.io/api/data-formats#meta)
lang, name, version, spacy_version (range), description, author, email, url,
license, sources (list of {name, url, author, license}, which is where treebank
provenance and its CC BY-SA obligation is recorded), requirements (extra pip deps injected
into the generated setup.cfg), vectors, pipeline, labels, performance (auto-filled by
spacy train), speed, spacy_git_version.
In v3, meta.json "isn't used to construct the language class and pipeline anymore".
config.cfg is the single source of truth for loading; meta.json is metadata and packaging.
4. Canonical training workflow
Template: https://github.com/explosion/projects/tree/v3/pipelines/tagger_parser_ud
(project.yml + configs/default.cfg). The directory layout assets / corpus / configs / training / metrics / packages is the convention, which this repo follows.
# 1. assets: clone the UD treebank
git clone https://github.com/UniversalDependencies/UD_Persian-PerDT assets/UD_Persian-PerDT
# 2. corpus: CoNLL-U -> binary DocBin
python -m spacy convert assets/.../fa_perdt-ud-train.conllu corpus/ \
--converter conllu --n-sents 10 --merge-subtokens
# 3. config
python -m spacy init config configs/default.cfg --lang fa \
--pipeline tagger,morphologizer,trainable_lemmatizer,parser --optimize efficiency
python -m spacy init fill-config partial.cfg configs/default.cfg # completes defaults
# 4. (md/lg only) vectors
python -m spacy init vectors fa cc.fa.300.vec vectors/ --mode floret --prune 200000
# 5. train
python -m spacy train configs/default.cfg --output training/ --gpu-id -1 \
--paths.train corpus/train.spacy --paths.dev corpus/dev.spacy
# 6. evaluate on held-out test
python -m spacy benchmark accuracy training/model-best corpus/test.spacy --output metrics/test.json
# ('spacy evaluate' is now an alias for 'benchmark accuracy')
# 7. package
python -m spacy package training/fa_dep_news_sm packages --name dep_news_sm --version 3.8.0 --build sdist,wheel
spacy assemble builds a pipeline from a config without training, which suits a rule-only
artifact (tokenizer + lookup lemmatizer + stop words).
Flags that matter on spacy convert
--converter conllu, explicit rather thanauto.--n-sents N, how many sentences perDoc. 10 is the Explosion default and gives the parser and senter cross-sentence context.--merge-subtokens. spaCy has no multi-word-token layer, and Persian UD treebanks use MWT ranges for clitics (پدرم=پدر+م), so this flag decides whether gold tokens are clitic-split or fused. Seedocs/MODELS.md§5.--morphology, which appends morph features to the tag. Only for pipelines without a morphologizer.
How the size tiers differ mechanically
| Tier | Embedding source | Config difference |
|---|---|---|
sm |
hash embeddings only | spacy.MultiHashEmbed with include_static_vectors = false, spacy.MaxoutWindowEncoder width 96, every other component a spacy.Tok2VecListener on the shared tok2vec |
md / lg |
+ static vectors | same CNN, include_static_vectors = true, [initialize] vectors = <dir built by init vectors>; lg has a bigger table |
trf |
contextual | replace tok2vec with a transformer component; other components listen via TransformerListener instead of Tok2VecListener |
spacy pretrain (Tok2Vec LM-style pretraining on raw text into [initialize] init_tok2vec) is
orthogonal to the tier and optional.
5. Upstream spacy/lang/fa: current state and gaps
Contents of https://github.com/explosion/spaCy/tree/master/spacy/lang/fa:
| File | Size | What it provides |
|---|---|---|
__init__.py |
1.3 KB | PersianDefaults: tokenizer exceptions, TOKENIZER_SUFFIXES, LEX_ATTRS, SYNTAX_ITERATORS, STOP_WORDS, writing_system = {"direction": "rtl", "has_case": False, "has_letters": True}; registers a lemmatizer factory defaulting to mode="rule" |
tokenizer_exceptions.py |
64.9 KB | Generated compound-verb and enclitic exception table |
generate_verbs_exc.py |
14.8 KB | Dev script that generates the above |
stop_words.py |
3.8 KB | ~500 stop words; the comment says "Stop words from HAZM package" |
lex_attrs.py |
1.4 KB | LIKE_NUM only (Persian numerals plus ام and ین suffixes) |
punctuation.py |
508 B | TOKENIZER_SUFFIXES only, no prefixes and no infixes |
syntax_iterators.py |
1.6 KB | noun_chunks(), which needs a trained parser |
Tests: spacy/tests/lang/fa/ has only test_noun_chunks.py. No tokenizer, lemmatizer or
stop-word tests.
spacy-lookups-data (MIT) already ships Persian rule-lemmatizer tables:
| File | Size |
|---|---|
fa_lemma_exc.json |
1.68 MB |
fa_lemma_index.json |
171 KB |
fa_lemma_rules.json |
882 B |
fa_source.txt |
"extracted from Mojgan Seraji's Persian Universal Dependencies Corpus" |
Missing there: fa_lemma_lookup.json (no lookup-mode table), fa_lexeme_norm.json, and
fa_license.txt (Catalan has one; Persian's Seraji-derived provenance is undocumented).
Gaps blocking a full fa_core_news_*:
- No trained artefacts: no weights, no
config.cfg, nometa.json, no entry incompatibility.json.spacy.load("fa_core_news_sm")fails today. - No word vectors for
fa, somdandlgneed vectors built from scratch. - No NER data or labels anywhere in
spacy/lang/fa, socoreneeds an external corpus. - The rule lemmatizer needs
token.posfrom a tagger or morphologizer, which is chicken-and-egg until the tagger exists.trainable_lemmatizeravoids this by learning edit trees and needing no tables. - No Persian-specific prefix or infix rules. ZWNJ (U+200C) is handled only implicitly through the verb-exception table.
- No tokenizer test coverage for a 65 KB exception table.
Items 5 and 6 are upstream PR material, along with the noun_chunks bug in
docs/upstream/fa-noun-chunks.md. Items 1 to 4 are this project's job.
6. Prior art
hazm publishes spaCy-format Persian pipelines on the HF Hub, each a single-task pipeline built
on ParsBERT, with the license field left empty:
| HF repo | Pipeline | Reported | Trained on ([paths] in its config.cfg) |
|---|---|---|---|
roshan-research/hazm-parsbert-postagger |
transformer, tagger |
tag_acc 0.9862 |
data_train_98_rs10.spacy (hazm's own EZ-augmented tagset: NOUN,EZ, ADJ,EZ, …) |
roshan-research/hazm-bert-dependency-parser |
transformer, parser |
dep_uas 0.9246 / dep_las 0.8934 |
modified_fa_perdt-ud-train.spacy, i.e. UD_Persian-PerDT |
roshan-research/hazm-parsbert-chunker |
transformer, tagger (IOB chunk tags) |
tag_acc 0.9618 |
hazm chunk data |
All three use spacy-transformers TransformerModel.v3 on
HooshvareLab/bert-base-parsbert-uncased, strided_spans window 128 and stride 96, with
spacy_version >= 3.6.0,<3.7.0.
Problems this project avoids:
- Three separate pipelines rather than one, so users pay for three BERT forward passes and
cannot share a
Doc. version: 0.0.0, emptylicense,authorandsources, andname: "pipeline", which breaks the conventions in §3 and leaves redistribution rights unclear.- Transformer-only, with no CPU-friendly tier.
- No lemmatizer, morphologizer, NER or vectors.
- Pinned to spaCy 3.6, so unusable on 3.8 without retraining.
The reusable finding is hazm's parser recipe, UD_Persian-PerDT plus spaCy's
TransitionBasedParser, which independently confirms the corpus choice in docs/MODELS.md.