Trains tagger, morphologizer, trainable_lemmatizer, parser and ner against one
fine-tuned HooshvareLab/bert-base-parsbert-uncased through TransformerListener,
rather than the sm/md/lg split where ner is trained separately and sourced in.
Fine-tuning a 162M-parameter encoder twice would double GPU cost, ship two
encoders in one wheel, and collide on the `transformer` component name.
A shared encoder needs one corpus carrying both annotation layers, so
merge_joint_corpus.py fuses them. corpus/perdt-ner/ came from the same
--merge-subtokens CoNLL-U as corpus/merged/ with the same --n-sents, so the
DocBins are token-for-token identical; the script asserts that per document and
copies doc.ents by token index. Char offsets do not work here because the two
converters differ in trailing whitespace, which pushes char_span off the token
grid and returns None.
Trained on a Colab T4 in 1h58m, 3000 steps, no early stop. Test scores against
lg: DEP_LAS 90.79 (+4.19), ENTS_F 82.89 (+6.95, almost all recall), TAG_ACC
97.62 (+1.07). DEP_LAS passes the hazm+ParsBERT reference of 89.34, which no CPU
tier reached. LEMMA_ACC 97.31 and SENTS_F 97.35 regress against lg; the likely
cause of the latter is strided_spans leaving only 32 tokens of overlap.
max_steps and learn_rate.total_steps are held equal on purpose. They are
independent knobs, and a patience stop under a longer total_steps ends training
at a high learning rate, discarding the annealing tail.
finalize_pipeline.py now reads the encoder name out of the trained config
instead of hardcoding it, tracks per-encoder licences, and writes a
redistribution warning into meta.json when the encoder states none. ParsBERT
states none, so that wheel is not redistributable; roberta-fa-zwnj-base
(Apache-2.0) is a one-line change to `name`.
Add benchmark_throughput.py, which times nlp.pipe alone. The words/s from
`spacy benchmark accuracy` includes the Scorer's per-token alignment, which is
why MODELS.md 7 reported ent_lg as faster than sm despite an identical ner
architecture. Remeasured every tier on CPU and the 940MX; a background rsync
halved every figure, so the final numbers are 9-run medians on an idle machine.
Add make_model_card.py, which composes the Hub card from meta.json and the
throughput records. The card spacy package writes has no frontmatter, so the Hub
cannot index the model by language, and no install line or usage.
UD_Persian-PerDT ships entity annotations in not-to-release/Dadegan with NER
tag/ that nothing in this project had looked at: 29,107 sentences, 484,312
tokens, 15,833 entities, under the treebank's own CC BY-SA 4.0. That removes the
reason core was withheld. The previous commit split dep from ent because the only
redistributable Persian NER corpus known then was ParsTwiNER, a Twitter corpus
scoring 67.22 F against 85-98 for the UD components, and hiding that genre and
quality gap behind one package name was not acceptable. A NER layer from the same
corpus has none of those problems: one genre, one tokenization, one licence, one
provenance chain.
Two shipping packages, both from PerDT alone:
fa_dep_news_sm 7.5 MB TAG 95.96 LEMMA 97.91 LAS 85.15
fa_core_news_sm 13 MB the above plus ENTS_P/R/F 77.67 / 66.87 / 71.87
Per label: LOC 80.24, DAT 74.45, MON 73.68, ORG 68.77, TIM 66.67, PER 65.29,
PCT 57.14. PER scoring below LOC on comparable data is the silver labels showing:
PerDT starts 6.24% of PER spans with a title against ParsTwiNER's 1.41%, so the
boundaries are less regular than the count suggests. MON, TIM and PCT rest on 4
to 11 test entities and are indicative only.
The labels are silver, from Beheshti-NER (Taher et al. 2020) with manual
corrections per the treebank README. meta.json notes say so. A human-annotated
test set is the outstanding work, tracked in ../ner_dataset.
Alignment: the NER files use the original Dadegan tokenization, matching the
released UD tokenization in only 57 to 62% of sentences (dropped copulas and
auxiliaries, one honorific corrupted to a comma). scripts/transfer_perdt_ner.py
realigns with difflib at 99.86% train / 99.74% dev / 99.51% test; spans whose
tokens do not all map contiguously are dropped rather than guessed.
Also fixed a silent metadata corruption. Folding both benchmark reports over one
key set gave core tag_acc 0.00, because the NER corpus has no gold tags so its
report carries tag_acc: 0.0, which overwrote the real 95.96; sents_f was wrong
the same way. finalize_pipeline.py now takes --ud-metrics and --ner-metrics
separately, each restricted to the keys its corpus can evidence. Its variant
guards work both directions now: a dep pipeline may not contain ner, a core one
must.
ParsTwiNER moves out of the shipping repo to a future mixed-genre package, with
the ablation that justifies it: prose-trained NER scores 45.72 F on tweets,
tweet-trained scores 55.59 on prose, and mixing gives 72.11 / 66.49, so roughly
20 F of cross-genre robustness for under 1 F on prose.
Adds LICENSE recording the split: MIT for the code, CC BY-SA 4.0 for the trained
pipelines as Adapted Material, with the attribution, modification notice and
warranty disclaimer that CC BY-SA 4.0 section 3 requires. Verified the treebank's
LICENSE.txt is unmodified CC BY-SA 4.0 with no carve-out for not-to-release/.
Author metadata filled in; both packages rebuilt and installed from their wheels.
No trained Persian pipeline exists for spaCy: spacy.load("fa_core_news_sm") has
never worked. This builds one from openly-licensed data so the result can actually
be redistributed.
Test scores (held-out splits): TAG 95.96, POS 96.24, MORPH 96.29, LEMMA 97.91,
UAS 89.69, LAS 85.15, ENTS_F 67.22. ~9,250 words/s, 13 MB wheel. 1h27m on 4 CPU
cores, no GPU.
Corpus choices, with evidence:
- UD_Persian-PerDT (CC BY-SA 4.0) over Seraji: 3.7x more tokens (452k vs 121k) and
Seraji has no PROPN tag at all. hazm's own spaCy parser used PerDT too.
- ParsTwiNER (MIT) for NER. ARMAN/PEYMA/NSURL are research-only; spaCy never
shipped the 2018 Persian models precisely because of corpus licensing. Cost: NER
is trained on tweets, so it is the weak component.
- --merge-subtokens: measured token F 0.9887 vs 0.9823 split. Without it, 1.5% of
gold token boundaries are unreachable by the tokenizer we ship.
- morphologizer + trainable_lemmatizer instead of the English attribute_ruler +
rule lemmatizer, because UD gives gold UPOS/FEATS/lemmas to train and measure on.
The ner component embeds its own tok2vec rather than using a Tok2VecListener, so it
can be sourced into the core pipeline after being trained on a separate corpus.
Also documents an upstream bug: spacy/lang/fa/syntax_iterators.py matches ClearNLP
labels (dobj, pobj, nsubjpass, attr, dative) that do not exist in UD, so
doc.noun_chunks returns bare head nouns (1.31 vs 2.77 tokens/chunk). Patch and
tests in docs/upstream/fa-noun-chunks.md.