diff --git a/README.fa.md b/README.fa.md index 733bc4a..a2417f1 100644 --- a/README.fa.md +++ b/README.fa.md @@ -12,7 +12,7 @@ **بازشناسی موجودیت‌های نام‌دار** را دارد. هر دو تحت لیسانس CC BY-SA ۴٫۰ منتشر شده‌اند. ```bash -pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl +pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl ``` ```python diff --git a/README.md b/README.md index e7d94b5..7e0fe21 100644 --- a/README.md +++ b/README.md @@ -1,9 +1,9 @@ # Persian (Farsi) pipelines for spaCy -Trained spaCy pipelines for Persian, installable now. spaCy has never shipped an official one, and `spacy.blank("fa")` only gives you a tokenizer and stop words. choose between `fa_core_news_sm` (full syntax + NER) or `fa_dep_news_sm` (syntax only). +Trained spaCy pipelines for Persian, installable now. spaCy has never shipped an official one, and `spacy.blank("fa")` only gives you a tokenizer and stop words. Choose between `fa_core_news_sm` (full syntax + NER) or `fa_dep_news_sm` (syntax only). ```bash -pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl +pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl ``` ```python @@ -21,18 +21,9 @@ pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_n [('ایران خودرو', 'ORG'), ('۲۰ درصد', 'PCT')] ``` -## Why spacy-persian? - -- **⚡ Performance** – **96.24%** POS · **97.91%** Lemma · **85.15%** LAS – competitive with English `en_core_web_sm` on syntax. -- **🚀 Speed** – ~9,250 words/sec on a standard CPU. No GPU required. -- **📦 Flexibility** – Choose `fa_core_news_sm` (13MB, syntax + NER) or `fa_dep_news_sm` (7.5MB, syntax-only). -- **🔁 Reproducibility** – Checksummed, versioned builds from UD_Persian-PerDT – no black boxes. -- **🔌 Native spaCy** – Drop-in replacement. `spacy.load()` works instantly with standard `Doc` objects. -- - ## Results -`spacy-persian` delivers production‑ready Persian NLP that stands alongside Hazm—the most popular Persian toolkit—while bringing the full power of the spaCy ecosystem. +Compared against Hazm (the most-used Persian toolkit) and `en_core_web_sm` (English reference). | Metric | **`spacy-persian`**
`fa_core_news_sm` | **Hazm**
(Persian toolkit) | `en_core_web_sm`
(English reference) | |--------|:---:|:---:|:---:| @@ -48,42 +39,63 @@ pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_n > ⚠️ **Note on comparability:** These benchmarks come from *different evaluation sets, treebanks, and test splits*. +From `spacy benchmark accuracy`, stored in `metrics/`. +| Package | Components | Licence | Score | Wheel | +| --- | --- | --- | --- | --- | +| `fa_dep_news_sm` | tok2vec, tagger, morphologizer, trainable_lemmatizer, parser | CC BY-SA 4.0 | LEMMA 97.91 | 7.5 MB | +| `fa_core_news_sm` | the above plus ner | CC BY-SA 4.0 | ENTS_F 71.87 | 13 MB | +| `fa_dep_news_md` | same as `fa_dep_news_sm`, plus floret vectors | CC BY-SA 4.0 | LEMMA 97.96 | 62 MB | +| `fa_core_news_md` | same as `fa_core_news_sm`, plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 68 MB | +| `fa_ent_news_md` | `ner` alone (own embedded tok2vec), plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 58 MB | -| Metric | Score | Reference | -| --- | --- | --- | -| `TOKEN_ACC` / `TOKEN_F` | 99.96 / 99.11 | | -| `TAG_ACC` (XPOS) | 95.96 | | -| `POS_ACC` (UPOS) | 96.24 | | -| `MORPH_ACC` | 96.29 | | -| `LEMMA_ACC` | 97.91 | | -| `SENTS_F` | 99.25 | | -| `DEP_UAS` | 89.69 | hazm+ParsBERT: 92.46 | -| `DEP_LAS` | 85.15 | hazm+ParsBERT: 89.34 | -| Speed | ~9,250 words/s | | +The `md` tier adds a 50k x 300d floret vector table trained on 400k Persian documents. Its +config differs from `sm` by exactly one line (`include_static_vectors`), so the columns below +isolate what the vectors buy. Full breakdown in `docs/MODELS.md` §6. -Entities, `fa_core_news_sm` only, on the PerDT NER test split: `ENTS_P` 77.67, `ENTS_R` 66.87, -`ENTS_F` 71.87. +| Metric | `sm` | `md` | Reference | +| --- | --- | --- | --- | +| `TOKEN_ACC` / `TOKEN_F` | 99.96 / 99.11 | 99.96 / 99.11 | | +| `TAG_ACC` (XPOS) | 95.96 | 96.25 | | +| `POS_ACC` (UPOS) | 96.24 | 96.64 | | +| `MORPH_ACC` | 96.29 | 96.64 | | +| `LEMMA_ACC` | 97.91 | 97.96 | | +| `SENTS_F` | 99.25 | 99.28 | | +| `DEP_UAS` | 89.69 | 90.52 | hazm+ParsBERT: 92.46 | +| `DEP_LAS` | 85.15 | 86.34 | hazm+ParsBERT: 89.34 | +| `ENTS_P` | 77.67 | 76.56 | | +| `ENTS_R` | 66.87 | 72.95 | | +| `ENTS_F` | 71.87 | 74.71 | | +| Speed | ~9,250 words/s | ~7,700 words/s | | -| Label | F | Train examples | -| --- | --- | --- | -| `LOC` | 80.24 | 4,954 | -| `DAT` | 74.45 | 1,323 | -| `MON` | 73.68 | 205 | -| `ORG` | 68.77 | 2,643 | -| `TIM` | 66.67 | 135 | -| `PER` | 65.29 | 4,847 | -| `PCT` | 57.14 | 121 | +Entity scores are `fa_core_news_*` on the PerDT NER test split. The `md` gain is almost +entirely recall (+6.08): static vectors give the model a lexical prior for rare proper nouns +that hash embeddings never had. + +| Label | Gold in test | `sm` F | `md` F | Train examples | +| --- | --- | --- | --- | --- | +| `LOC` | 273 | 80.24 | 84.05 | 4,954 | +| `PER` | 297 | 65.29 | 68.18 | 4,847 | +| `ORG` | 144 | 68.77 | 70.25 | 2,643 | +| `DAT` | 69 | 74.45 | 76.19 | 1,323 | +| `MON` | 10 | 73.68 | 84.21 | 205 | +| `TIM` | 9 | 66.67 | 66.67 | 135 | +| `PCT` | 4 | 57.14 | 33.33 | 121 | + +`MON`, `TIM` and `PCT` have single-digit support in the test split, so their deltas are one +or two entities changing hands, not signal. The three labels that carry the split (`PER`, +`LOC`, `ORG`) all improve. For comparison, `en_core_web_sm` scores TAG 97, LAS 90, ENTS_F 84 on a larger, cleaner corpus. -Trained on a 4-core i5-7200U with no GPU: 1h27m for the syntax components, 17 min for NER. +Trained on a 4-core i5-7200U with no GPU: `sm` 1h27m syntax + 17 min NER, `md` 1h54m syntax ++ 25 min NER (the two `md` runs overlapped, so wall clock overstates each). ## Install ```bash -pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl +pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl # or, without NER: -pip install https://huggingface.co/Phazel/fa_dep_news_sm/resolve/main/fa_dep_news_sm-any-py3-none-any.whl +pip install https://huggingface.co/Phazel/fa_dep_news_sm/resolve/main/fa_dep_news_sm-3.8.0-py3-none-any.whl ``` ## Caveats diff --git a/configs/fa_dep_news_md.cfg b/configs/fa_dep_news_md.cfg new file mode 100644 index 0000000..70cde34 --- /dev/null +++ b/configs/fa_dep_news_md.cfg @@ -0,0 +1,235 @@ +# fa_dep_news_md — tagger, morphologizer, trainable_lemmatizer, parser, WITH static vectors. +# +# Byte-identical to configs/fa_dep_news_sm.cfg except: +# - [components.tok2vec.model.embed] include_static_vectors: false -> true +# +# Everything else (seed, widths, rows, batcher, patience, eval_frequency) is held constant so +# the sm/md delta measures the floret vectors and nothing else. +# +# Vectors are supplied at train time via --paths.vectors, pointing at the fa_floret table +# (50k rows x 300d, floret mode, minn=maxn=5, hash_count=2) trained on 400k Persian documents. +# floret has no OOV: every string hashes into the table, which is the point for Persian, where +# ZWNJ inconsistency (می‌رود / میرود / می رود) would shatter a classic word table. +# +# No `ner` here by design; see configs/fa_ner_md.cfg and project.yml. + +[paths] +train = null +dev = null +vectors = null +init_tok2vec = null + +[system] +gpu_allocator = null +seed = 0 + +[nlp] +lang = "fa" +pipeline = ["tok2vec", "tagger", "morphologizer", "trainable_lemmatizer", "parser"] +batch_size = 1000 +disabled = [] +before_creation = null +after_creation = null +after_pipeline_creation = null + +[corpora] + +[training] +dev_corpus = "corpora.dev" +train_corpus = "corpora.train" +seed = ${system.seed} +gpu_allocator = ${system.gpu_allocator} +dropout = 0.1 +accumulate_gradient = 1 +patience = 1600 +max_epochs = 0 +max_steps = 20000 +eval_frequency = 400 +frozen_components = [] +annotating_components = [] +before_to_disk = null +before_update = null + +[initialize] +vectors = ${paths.vectors} +init_tok2vec = ${paths.init_tok2vec} +vocab_data = null +lookups = null +before_init = null +after_init = null + +[components] + +[pretraining] + +[nlp.tokenizer] +@tokenizers = "spacy.Tokenizer.v1" + +[nlp.vectors] +@vectors = "spacy.Vectors.v1" + +[corpora.train] +@readers = "spacy.Corpus.v1" +path = ${paths.train} +max_length = 0 +gold_preproc = false +limit = 0 +augmenter = null + +[corpora.dev] +@readers = "spacy.Corpus.v1" +path = ${paths.dev} +max_length = 0 +gold_preproc = false +limit = 0 +augmenter = null + +[training.optimizer] +@optimizers = "Adam.v1" +beta1 = 0.9 +beta2 = 0.999 +L2_is_weight_decay = true +L2 = 0.01 +grad_clip = 1.0 +use_averages = false +eps = 1e-08 +learn_rate = 0.001 + +[training.batcher] +@batchers = "spacy.batch_by_words.v1" +discard_oversize = false +tolerance = 0.2 +get_length = null + +[training.logger] +@loggers = "spacy.ConsoleLogger.v1" +progress_bar = false + +[training.score_weights] +tag_acc = 0.25 +pos_acc = 0.12 +tag_micro_p = null +tag_micro_r = null +tag_micro_f = null +morph_acc = 0.12 +morph_per_feat = null +lemma_acc = 0.25 +dep_uas = 0.12 +dep_las = 0.12 +dep_las_per_type = null +sents_p = null +sents_r = null +sents_f = 0.0 + +[initialize.tokenizer] + +[initialize.components] + +[components.tok2vec] +factory = "tok2vec" + +[components.tagger] +factory = "tagger" +label_smoothing = 0.05 +overwrite = false +neg_prefix = "!" + +[components.morphologizer] +factory = "morphologizer" +label_smoothing = 0.05 +overwrite = true +extend = false + +[components.trainable_lemmatizer] +factory = "trainable_lemmatizer" +backoff = "orth" +min_tree_freq = 3 +overwrite = false +top_k = 1 + +[components.parser] +factory = "parser" +moves = null +update_with_oracle_cut_size = 100 +learn_tokens = false +min_action_freq = 30 + +[training.batcher.size] +@schedules = "compounding.v1" +start = 100 +stop = 1000 +compound = 1.001 +t = 0.0 + +[components.tok2vec.model] +@architectures = "spacy.Tok2Vec.v2" + +[components.tagger.model] +@architectures = "spacy.Tagger.v2" +nO = null +normalize = false + +[components.tagger.scorer] +@scorers = "spacy.tagger_scorer.v1" + +[components.morphologizer.model] +@architectures = "spacy.Tagger.v2" +nO = null +normalize = false + +[components.morphologizer.scorer] +@scorers = "spacy.morphologizer_scorer.v1" + +[components.trainable_lemmatizer.model] +@architectures = "spacy.Tagger.v2" +nO = null +normalize = false + +[components.trainable_lemmatizer.scorer] +@scorers = "spacy.lemmatizer_scorer.v1" + +[components.parser.model] +@architectures = "spacy.TransitionBasedParser.v2" +state_type = "parser" +extra_state_tokens = false +hidden_width = 128 +maxout_pieces = 3 +use_upper = true +nO = null + +[components.parser.scorer] +@scorers = "spacy.parser_scorer.v1" + +[components.tok2vec.model.embed] +@architectures = "spacy.MultiHashEmbed.v2" +width = ${components.tok2vec.model.encode.width} +attrs = ["NORM", "PREFIX", "SUFFIX", "SHAPE"] +rows = [5000, 1000, 2500, 2500] +include_static_vectors = true + +[components.tok2vec.model.encode] +@architectures = "spacy.MaxoutWindowEncoder.v2" +width = 96 +depth = 4 +window_size = 1 +maxout_pieces = 3 + +[components.tagger.model.tok2vec] +@architectures = "spacy.Tok2VecListener.v1" +width = ${components.tok2vec.model.encode.width} +upstream = "*" + +[components.morphologizer.model.tok2vec] +@architectures = "spacy.Tok2VecListener.v1" +width = ${components.tok2vec.model.encode.width} +upstream = "*" + +[components.trainable_lemmatizer.model.tok2vec] +@architectures = "spacy.Tok2VecListener.v1" +width = ${components.tok2vec.model.encode.width} +upstream = "*" + +[components.parser.model.tok2vec] +@architectures = "spacy.Tok2VecListener.v1" +width = ${components.tok2vec.model.encode.width} +upstream = "*" \ No newline at end of file diff --git a/configs/fa_ner_md.cfg b/configs/fa_ner_md.cfg new file mode 100644 index 0000000..607dc7a --- /dev/null +++ b/configs/fa_ner_md.cfg @@ -0,0 +1,154 @@ +# fa_ent_news_md — Persian NER with static floret vectors. +# +# Identical to configs/fa_ner_sm.cfg except: +# - [components.ner.model.tok2vec.embed] include_static_vectors: false -> true +# +# Same embedded-tok2vec design as the sm variant (no Tok2VecListener), so the trained +# component stays sourceable into fa_core_news_md via `nlp.add_pipe("ner", source=...)`. +# +# Vectors supplied at train time via --paths.vectors (fa_floret, 50k rows x 300d, +# trained on 400k Persian documents). + +[paths] +train = null +dev = null +vectors = null +init_tok2vec = null + +[system] +gpu_allocator = null +seed = 0 + +[nlp] +lang = "fa" +pipeline = ["ner"] +batch_size = 1000 +disabled = [] +before_creation = null +after_creation = null +after_pipeline_creation = null + +[nlp.tokenizer] +@tokenizers = "spacy.Tokenizer.v1" + +[nlp.vectors] +@vectors = "spacy.Vectors.v1" + +[components] + +[components.ner] +factory = "ner" +moves = null +update_with_oracle_cut_size = 100 +incorrect_spans_key = null + +[components.ner.model] +@architectures = "spacy.TransitionBasedParser.v2" +state_type = "ner" +extra_state_tokens = false +hidden_width = 64 +maxout_pieces = 2 +use_upper = true +nO = null + +[components.ner.model.tok2vec] +@architectures = "spacy.Tok2Vec.v2" + +[components.ner.model.tok2vec.embed] +@architectures = "spacy.MultiHashEmbed.v2" +width = ${components.ner.model.tok2vec.encode.width} +attrs = ["NORM", "PREFIX", "SUFFIX", "SHAPE"] +rows = [5000, 1000, 2500, 2500] +include_static_vectors = true + +[components.ner.model.tok2vec.encode] +@architectures = "spacy.MaxoutWindowEncoder.v2" +width = 96 +depth = 4 +window_size = 1 +maxout_pieces = 3 + +[components.ner.scorer] +@scorers = "spacy.ner_scorer.v1" + +[corpora] + +[corpora.train] +@readers = "spacy.Corpus.v1" +path = ${paths.train} +max_length = 0 +gold_preproc = false +limit = 0 +augmenter = null + +[corpora.dev] +@readers = "spacy.Corpus.v1" +path = ${paths.dev} +max_length = 0 +gold_preproc = false +limit = 0 +augmenter = null + +[training] +dev_corpus = "corpora.dev" +train_corpus = "corpora.train" +seed = ${system.seed} +gpu_allocator = ${system.gpu_allocator} +dropout = 0.1 +accumulate_gradient = 1 +patience = 1600 +max_epochs = 0 +max_steps = 20000 +eval_frequency = 400 +frozen_components = [] +annotating_components = [] +before_to_disk = null +before_update = null + +[training.optimizer] +@optimizers = "Adam.v1" +beta1 = 0.9 +beta2 = 0.999 +L2_is_weight_decay = true +L2 = 0.01 +grad_clip = 1.0 +use_averages = false +eps = 1e-08 +learn_rate = 0.001 + +[training.batcher] +@batchers = "spacy.batch_by_words.v1" +discard_oversize = false +tolerance = 0.2 +get_length = null + +[training.batcher.size] +@schedules = "compounding.v1" +start = 100 +stop = 1000 +compound = 1.001 +t = 0.0 + +[training.logger] +@loggers = "spacy.ConsoleLogger.v1" +progress_bar = false + +[training.score_weights] +ents_f = 1.0 +ents_p = 0.0 +ents_r = 0.0 +ents_per_type = null + +[initialize] +vectors = ${paths.vectors} +init_tok2vec = ${paths.init_tok2vec} +vocab_data = null +lookups = null +before_init = null +after_init = null + +[initialize.tokenizer] + +[initialize.components] + +[pretraining] diff --git a/docs/CONTRIBUTING-GUIDE.md b/docs/CONTRIBUTING-GUIDE.md index de70337..0dc0a14 100644 --- a/docs/CONTRIBUTING-GUIDE.md +++ b/docs/CONTRIBUTING-GUIDE.md @@ -56,7 +56,12 @@ From : python -m spacy package training/fa_dep_news_sm packages --name dep_news_sm --version 3.8.0 --build wheel python -m spacy huggingface-hub push packages/fa_dep_news_sm-3.8.0/dist/fa_dep_news_sm-3.8.0-py3-none-any.whl --org ``` - Users then `pip install https://huggingface.co//fa_dep_news_sm/resolve/main/fa_dep_news_sm-any-py3-none-any.whl`. + Users then `pip install https://huggingface.co//fa_dep_news_sm/resolve/main/fa_dep_news_sm-3.8.0-py3-none-any.whl`. + Note the filename: `spacy huggingface-hub push` uploads the wheel as `-any-py3-none-any.whl`, + but `"any"` is not a valid PEP 440 version and current pip rejects it + (`Invalid wheel filename (invalid version)`). Upload a second copy under its real versioned + filename too (`api.upload_file(path_in_repo=f"{name}-{version}-py3-none-any.whl", ...)`) and + link to that one instead. 2. PyPI or a self-hosted wheel: `spacy package … --build sdist,wheel` then `twine upload`, or attach the wheel to a GitHub Release. See . 3. spaCy Universe, which lists the package on spacy.io but hosts nothing. Per diff --git a/docs/MODELS.md b/docs/MODELS.md index d34ebe7..a56f071 100644 --- a/docs/MODELS.md +++ b/docs/MODELS.md @@ -21,6 +21,13 @@ requirement. That ordering sets the roadmap below. Source: . +That last point does **not** transfer to Persian. The `sm` -> `md` step measured on this +project buys +1.19 LAS and +2.85 NER F (§6), where English gets ~0.00 LAS. Two reasons: PerDT +is roughly a tenth the size of OntoNotes, so hash embeddings have far less signal to learn a +lexicon from, and floret's subword hashing gives 0% OOV on a language whose ZWNJ variation +(می‌رود / میرود / می رود) fragments any fixed word-key table. English `md` uses 20k classic +word vectors and hits OOV constantly. Do not use the English row as the Persian prior. + ## 2. Target: the Persian pipelines Naming follows `[lang]_[type]_[genre]_[size]` (). The @@ -35,8 +42,9 @@ pipelines such as `de_core_news_sm` as `news`. | `fa_core_news_sm` | the above plus ner | hash embeddings | built, shipping | | `fa_ent_news_sm` | ner (own internal tok2vec) | hash embeddings | built, optional | | `fa_core_web_sm` | same as core, mixed-genre training data | hash embeddings | not built; would add ParsTwiNER to cover social media | -| `fa_core_news_md` | + static vectors | floret, 50k rows | vectors must be trained first (CPU-days on fa Wikipedia + OSCAR) | -| `fa_core_news_lg` | same | floret, 200k rows | same as md, bigger table | +| `fa_dep_news_md` | same as `fa_dep_news_sm` | floret, 50k rows / 300d | built, shipping | +| `fa_core_news_md` | same as `fa_core_news_sm` | floret, 50k rows / 300d | built, shipping | +| `fa_core_news_lg` | same | floret, 200k rows | not built; bigger table, same recipe as md | | `fa_core_news_trf` | transformer instead of tok2vec | `HooshvareLab/roberta-fa-zwnj-base` (Apache-2.0) | not on this hardware; 2 GB VRAM cannot fine-tune a 125M-param encoder | ### Why `core` is honest here @@ -276,3 +284,68 @@ Sources are recorded in each `meta.json` with their licences, per crediting Beheshti-NER. CC BY-SA 4.0 on PerDT means every package carries attribution and a share-alike notice, handled in `scripts/finalize_pipeline.py`, which also enforces the shape of each variant: it refuses to publish a `dep` pipeline containing `ner`, or a `core` one without it. + +## 6. The `md` tier: floret static vectors + +Built after the `sm` tier, from `fa_floret` — 50,000 rows x 300d, floret mode, `minn=maxn=5`, +`hash_count=2`, trained on 400,000 Persian documents. The wheel is a vectors-only pipeline; +`scripts/unpack_vectors.py` unwraps it into a directory `--paths.vectors` can read, so nothing +needs pip-installing to train against it. + +`configs/fa_dep_news_md.cfg` and `configs/fa_ner_md.cfg` are their `sm` counterparts with one +line changed, `include_static_vectors = false -> true`. Same seed, same corpus, same widths, +same batcher, same patience. The deltas below are therefore attributable to the vector table +and nothing else. Reproduce with `spacy project run md`, or the table alone with +`python scripts/compare_tiers.py`. + +### UD_Persian-PerDT test split + +| Metric | `sm` | `md` | Delta | +| --- | --- | --- | --- | +| `TAG_ACC` | 95.96 | 96.25 | +0.29 | +| `POS_ACC` | 96.24 | 96.64 | +0.40 | +| `MORPH_ACC` | 96.29 | 96.64 | +0.35 | +| `LEMMA_ACC` | 97.91 | 97.96 | +0.05 | +| `SENTS_F` | 99.25 | 99.28 | +0.03 | +| `DEP_UAS` | 89.69 | 90.52 | +0.83 | +| `DEP_LAS` | 85.15 | 86.34 | +1.19 | +| Speed (dep) | 12,505 w/s | 10,493 w/s | -16.1% | + +### PerDT NER test split, `fa_core_news_md` + +| Metric | `sm` | `md` | Delta | +| --- | --- | --- | --- | +| `ENTS_P` | 77.67 | 76.56 | -1.10 | +| `ENTS_R` | 66.87 | 72.95 | +6.08 | +| `ENTS_F` | 71.87 | 74.71 | +2.85 | + +Almost all of the NER gain is recall. That is the expected shape of a fix for a coverage +problem: hash embeddings had no lexical prior for rare proper nouns, so the `sm` model +declined to tag them. Precision slips ~1 point because the model now guesses more. + +| Label | Gold in test | `sm` F | `md` F | Delta | +| --- | --- | --- | --- | --- | +| `PER` | 297 | 65.29 | 68.18 | +2.89 | +| `LOC` | 273 | 80.24 | 84.05 | +3.81 | +| `ORG` | 144 | 68.77 | 70.25 | +1.48 | +| `DAT` | 69 | 74.45 | 76.19 | +1.74 | +| `MON` | 10 | 73.68 | 84.21 | +10.53 | +| `TIM` | 9 | 66.67 | 66.67 | +0.00 | +| `PCT` | 4 | 57.14 | 33.33 | -23.81 | + +Read the bottom three rows as noise, not signal. `PCT` has four gold entities in the whole +test split, so its -23.81 F is one entity changing hands; `MON`'s +10.53 is likewise one of +ten. The three labels with real support (`PER`, `LOC`, `ORG`, 714 entities between them) all +improve, which is the finding. + +### Cost + +The vectors dominate the artifact: `fa_dep_news_md` is a 62 MB wheel against 7.5 MB for `sm`, +`fa_core_news_md` 68 MB against 13 MB. Inference is ~16% slower across all three pipelines, +a uniform hit consistent with the extra 300d concatenation per token rather than anything +component-specific. Training cost was comparable to `sm` (early stop at step 12,400 of 20,000, +best checkpoint near 10,800). + +Whether that trade is worth it depends on deployment. For a 1.19 LAS and 2.85 NER F gain, a +9x larger download and 16% slower parse is a good deal on a server and a bad one in a browser +or a Lambda cold start. Both tiers ship; pick per target. diff --git a/project.yml b/project.yml index 62f7ac8..797973a 100644 --- a/project.yml +++ b/project.yml @@ -29,6 +29,11 @@ vars: # -1 = CPU. A GTX 940MX (2 GB) is not worth the transfer overhead for an sm pipeline. gpu: -1 n_sents: 10 + # md tier. Same architecture as sm plus the fa_floret static vector table. + dep_md_package_name: "dep_news_md" + core_md_package_name: "core_news_md" + floret_wheel: "fa_floret-0.1.0-py3-none-any-400k-documents.whl" + vectors_dir: "assets/vectors/fa_floret_400k" directories: - "assets" @@ -88,6 +93,18 @@ workflows: - finalize-ent - evaluate-ent - package-ent + # The md tier: same corpus and architecture, plus the fa_floret static vectors. + md: + - vectors-md + - train-dep-md + - train-ner-md + - finalize-dep-md + - assemble-core-md + - evaluate-md + - finalize-meta-md + - compare-md + - package-md + - smoke-md commands: - name: "inspect" @@ -297,6 +314,122 @@ commands: outputs: - "packages/${vars.lang}_${vars.ent_package_name}-${vars.package_version}" + # ---------------------------------------------------------------- md tier + + - name: "vectors-md" + help: > + Unpack the fa_floret wheel into a plain spaCy model directory that + `--paths.vectors` can point at. The wheel is a vectors-only pipeline + (empty `pipeline: []`), 50k rows x 300d in floret mode, trained on 400k + Persian documents, so no `spacy init vectors` step is needed. + script: + - "python scripts/unpack_vectors.py ${vars.floret_wheel} ${vars.vectors_dir}" + deps: + - "${vars.floret_wheel}" + - "scripts/unpack_vectors.py" + outputs: + - "${vars.vectors_dir}" + + - name: "train-dep-md" + help: "Train the dep pipeline with static floret vectors" + script: + - "python -m spacy train configs/fa_dep_news_md.cfg --output training/dep-md --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy --paths.vectors ${vars.vectors_dir} --gpu-id ${vars.gpu}" + deps: + - "corpus/merged/${vars.treebank}-ud-train.spacy" + - "corpus/merged/${vars.treebank}-ud-dev.spacy" + - "configs/fa_dep_news_md.cfg" + - "${vars.vectors_dir}" + outputs: + - "training/dep-md/model-best" + + - name: "train-ner-md" + help: "Train the NER component with static floret vectors" + script: + - "python -m spacy train configs/fa_ner_md.cfg --output training/perdt-ner-md --paths.train corpus/perdt-ner/train.spacy --paths.dev corpus/perdt-ner/dev.spacy --paths.vectors ${vars.vectors_dir} --gpu-id ${vars.gpu}" + deps: + - "corpus/perdt-ner/train.spacy" + - "corpus/perdt-ner/dev.spacy" + - "configs/fa_ner_md.cfg" + - "${vars.vectors_dir}" + outputs: + - "training/perdt-ner-md/model-best" + + - name: "finalize-dep-md" + help: "Write fa_dep_news_md metadata onto the trained md model" + script: + - "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_dep_news_md --variant dep --size md --version ${vars.package_version}" + deps: + - "training/dep-md/model-best" + - "scripts/finalize_pipeline.py" + outputs: + - "training/fa_dep_news_md" + + - name: "assemble-core-md" + help: "Source the md ner into the md dep pipeline to produce fa_core_news_md" + script: + - "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_core_news_md --variant core --size md --version ${vars.package_version} --add-ner training/perdt-ner-md/model-best" + deps: + - "training/dep-md/model-best" + - "training/perdt-ner-md/model-best" + - "scripts/finalize_pipeline.py" + outputs: + - "training/fa_core_news_md" + + - name: "evaluate-md" + help: "Score both md packages on the held-out test splits" + script: + - "python -m spacy benchmark accuracy training/fa_dep_news_md corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/md-ud-test.json --gpu-id ${vars.gpu}" + - "python -m spacy benchmark accuracy training/fa_core_news_md corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/md-core-ud-test.json --gpu-id ${vars.gpu}" + - "python -m spacy benchmark accuracy training/fa_core_news_md corpus/perdt-ner/test.spacy --output metrics/md-perdt-ner-test.json --gpu-id ${vars.gpu}" + deps: + - "training/fa_dep_news_md" + - "training/fa_core_news_md" + outputs: + - "metrics/md-ud-test.json" + - "metrics/md-core-ud-test.json" + - "metrics/md-perdt-ner-test.json" + + - name: "finalize-meta-md" + help: "Fold the md test scores into both md meta.json files" + script: + - "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_dep_news_md --variant dep --size md --version ${vars.package_version} --ud-metrics metrics/md-ud-test.json" + - "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_core_news_md --variant core --size md --version ${vars.package_version} --add-ner training/perdt-ner-md/model-best --ud-metrics metrics/md-core-ud-test.json --ner-metrics metrics/md-perdt-ner-test.json" + deps: + - "metrics/md-ud-test.json" + - "metrics/md-perdt-ner-test.json" + - "scripts/finalize_pipeline.py" + + - name: "compare-md" + help: "Table the sm vs md deltas from the metrics/ JSON reports" + script: + - "python scripts/compare_tiers.py" + deps: + - "metrics/md-ud-test.json" + - "metrics/md-perdt-ner-test.json" + - "scripts/compare_tiers.py" + + - name: "package-md" + help: "Build installable wheels + sdists for both md packages" + script: + - "python -m spacy package training/fa_dep_news_md packages --name ${vars.dep_md_package_name} --version ${vars.package_version} --build sdist,wheel --force" + - "python -m spacy package training/fa_core_news_md packages --name ${vars.core_md_package_name} --version ${vars.package_version} --build sdist,wheel --force" + deps: + - "training/fa_dep_news_md" + - "training/fa_core_news_md" + outputs: + - "packages/${vars.lang}_${vars.dep_md_package_name}-${vars.package_version}" + - "packages/${vars.lang}_${vars.core_md_package_name}-${vars.package_version}" + + - name: "smoke-md" + help: "Load both md pipelines and run them over real Persian text" + script: + - "python scripts/smoke_test.py training/fa_dep_news_md" + - "python scripts/smoke_test.py training/fa_core_news_md" + deps: + - "training/fa_dep_news_md" + - "training/fa_core_news_md" + + - name: "clean" help: "Drop corpora, training runs and metrics (keeps downloaded assets)" script: diff --git a/scripts/compare_tiers.py b/scripts/compare_tiers.py new file mode 100644 index 0000000..9bbc727 --- /dev/null +++ b/scripts/compare_tiers.py @@ -0,0 +1,101 @@ +"""Table the sm vs md test-set deltas. + +Both tiers are trained from the same corpus, the same seed and the same architecture; the +only difference is `include_static_vectors`. So the delta printed here is attributable to the +fa_floret vector table and nothing else. + +Reads the `spacy benchmark accuracy` reports written by the `evaluate-*` targets. Missing +files are reported rather than fatal, so this is runnable mid-build. + +Usage: + python scripts/compare_tiers.py [--metrics-dir metrics] +""" + +import argparse +import json +from pathlib import Path + +# (label, sm report, md report) +PAIRS = [ + ("dep pipeline, UD test", "ud-test.json", "md-ud-test.json"), + ("core pipeline, UD test", "core-ud-test.json", "md-core-ud-test.json"), + ("core pipeline, NER test", "perdt-ner-test.json", "md-perdt-ner-test.json"), +] + +SCALARS = [ + ("tag_acc", "TAG_ACC"), + ("pos_acc", "POS_ACC"), + ("morph_acc", "MORPH_ACC"), + ("lemma_acc", "LEMMA_ACC"), + ("dep_uas", "DEP_UAS"), + ("dep_las", "DEP_LAS"), + ("sents_f", "SENTS_F"), + ("ents_p", "ENTS_P"), + ("ents_r", "ENTS_R"), + ("ents_f", "ENTS_F"), +] + + +def load(path): + return json.loads(path.read_text(encoding="utf8")) if path.exists() else None + + +def table(title, sm, md, rows): + print(f"\n## {title}\n") + print(f"| {'metric':<12} | {'sm':>7} | {'md':>7} | {'delta':>7} |") + print(f"| {'-' * 12} | {'-' * 7} | {'-' * 7} | {'-' * 7} |") + for key, label in rows: + a, b = sm.get(key), md.get(key) + if a is None and b is None: + continue + # The NER report scores tag_acc 0.0 because its corpus has no gold tags. + if a == 0.0 and b == 0.0: + continue + cells = [f"{v * 100:.2f}" if isinstance(v, float) else "-" for v in (a, b)] + delta = f"{(b - a) * 100:+.2f}" if isinstance(a, float) and isinstance(b, float) else "-" + print(f"| {label:<12} | {cells[0]:>7} | {cells[1]:>7} | {delta:>7} |") + for key, label in (("speed", "words/s"),): + a, b = sm.get(key), md.get(key) + if isinstance(a, float) and isinstance(b, float): + print(f"| {label:<12} | {a:>7.0f} | {b:>7.0f} | {b / a - 1:>+6.1%} |") + + +def per_type(title, sm, md): + a, b = sm.get("ents_per_type"), md.get("ents_per_type") + if not a or not b: + return + + def pct(v): + return f"{v * 100:.2f}" if v is not None else "-" + + print(f"\n### {title}, per label\n") + print(f"| {'label':<6} | {'sm F':>7} | {'md F':>7} | {'delta':>7} |") + print(f"| {'-' * 6} | {'-' * 7} | {'-' * 7} | {'-' * 7} |") + for label in sorted(set(a) | set(b), key=lambda k: -b.get(k, {}).get("f", 0)): + fa, fb = a.get(label, {}).get("f"), b.get(label, {}).get("f") + delta = f"{(fb - fa) * 100:+.2f}" if fa is not None and fb is not None else "-" + print(f"| {label:<6} | {pct(fa):>7} | {pct(fb):>7} | {delta:>7} |") + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--metrics-dir", type=Path, default=Path("metrics")) + args = ap.parse_args() + + print("# sm vs md (fa_floret 400k static vectors)") + print("\nSame corpus, same seed, same architecture. Only difference:") + print("`include_static_vectors = false -> true`.") + + for title, sm_name, md_name in PAIRS: + sm = load(args.metrics_dir / sm_name) + md = load(args.metrics_dir / md_name) + if sm is None or md is None: + missing = [n for n, d in ((sm_name, sm), (md_name, md)) if d is None] + print(f"\n## {title}\n\n (skipped, missing {', '.join(missing)})") + continue + table(title, sm, md, SCALARS) + per_type(title, sm, md) + + +if __name__ == "__main__": + main() diff --git a/scripts/finalize_pipeline.py b/scripts/finalize_pipeline.py index 8e2c4a4..89c5b84 100644 --- a/scripts/finalize_pipeline.py +++ b/scripts/finalize_pipeline.py @@ -3,9 +3,13 @@ Three variants, following spaCy's `[lang]_[type]_[genre]_[size]` naming (https://spacy.io/models#conventions): - dep -> fa_dep_news_sm tagger + morphologizer + trainable_lemmatizer + parser - core -> fa_core_news_sm the above plus ner - ent -> fa_ent_news_sm ner only + dep -> fa_dep_news_ tagger + morphologizer + trainable_lemmatizer + parser + core -> fa_core_news_ the above plus ner + ent -> fa_ent_news_ ner only + +`--size` fills the size slot: `sm` (hash embeddings only, the default) or `md` (the same +architecture plus the fa_floret static vector table). It is metadata only; which vectors a +model actually carries is decided at train time by `--paths.vectors`. All three are built from UD_Persian-PerDT alone, including the NER, which comes from that treebank's own `not-to-release/Dadegan with NER tag/` layer. That is what makes `core` @@ -56,6 +60,25 @@ LANG_DATA = { "author": "Explosion and spaCy contributors", "license": "MIT", } +FLORET = { + "name": "fa_floret static vectors (50k rows x 300d, floret mode, 400k Persian documents)", + "url": PROJECT_URL, + "author": "Kiyarash Fazeli", + "license": "CC BY-SA 4.0", +} +FLORET_LG = { + "name": "fa_floret static vectors (lg tier: larger floret table trained on fa Wikipedia + " + "OSCAR via spacy-vectors-builder)", + "url": PROJECT_URL, + "author": "Kiyarash Fazeli", + "license": "CC BY-SA 4.0", +} +TRANSFORMER = { + "name": "HooshvareLab/roberta-fa-zwnj-base", + "url": "https://huggingface.co/HooshvareLab/roberta-fa-zwnj-base", + "author": "Hooshvare Team", + "license": "Apache-2.0", +} NER_NOTE = ( "The ner component is trained on the NER layer shipped in UD_Persian-PerDT's " @@ -78,6 +101,36 @@ CHUNK_NOTE = ( "ClearNLP labels that do not exist in Universal Dependencies, see " "docs/upstream/fa-noun-chunks.md." ) +VECTORS_NOTE = ( + "This is the `md` tier: identical architecture to the `sm` pipeline plus static floret " + "vectors (50,000 rows x 300 dimensions, minn=maxn=5, hash_count=2) trained on 400,000 " + "Persian documents. floret hashes subwords into a fixed table, so there are no " + "out-of-vocabulary tokens and `token.has_vector` is always True. That matters for " + "Persian, where inconsistent ZWNJ (U+200C) usage splits one word across several surface " + "forms (mi-ravad written joined, with ZWNJ, or with a space) that a classic word-vector " + "table would miss." +) + + +def vectors_note_lg(nlp): + """Row/dim counts come from the trained model, not a hardcoded description, because the + lg-tier floret table is still being iterated on (unlike md's fixed, shipped table).""" + rows, dim = nlp.vocab.vectors.shape + return ( + f"This is the `lg` tier: identical architecture to `sm`/`md` but a larger static " + f"floret vector table ({rows:,} rows x {dim} dimensions, minn=maxn=5, hash_count=2) " + f"trained on Persian Wikipedia + OSCAR via spacy-vectors-builder. Same zero-OOV " + f"rationale as `md` (see docs/MODELS.md): floret hashes subwords into a fixed table, " + f"so `token.has_vector` is always True despite Persian's ZWNJ (U+200C) inconsistency." + ) + + +TRANSFORMER_NOTE = ( + "This is the `trf` tier: no static vectors; contextual embeddings instead come from a " + "fine-tuned HooshvareLab/roberta-fa-zwnj-base (Apache-2.0) transformer via " + "spacy-transformers. Not ParsBERT: its model card carries no licence. GPU is recommended " + "for both training and inference." +) # CC BY-SA 4.0 on the treebank propagates to anything derived from it. PERDT_LICENSE = "CC BY-SA 4.0" ATTRIBUTION = ( @@ -92,10 +145,10 @@ NER_KEYS = ("ents_p", "ents_r", "ents_f", "ents_per_type") VARIANTS = { "dep": { - "name": "dep_news_sm", + "name": "dep_news_{size}", "description": ( "Persian dependency pipeline optimized for CPU. Components: tok2vec, tagger, " - "morphologizer, trainable_lemmatizer, parser. No NER, see fa_core_news_sm." + "morphologizer, trainable_lemmatizer, parser. No NER, see fa_core_news_{size}." ), "license": PERDT_LICENSE, "sources": [PERDT, LANG_DATA], @@ -105,7 +158,7 @@ VARIANTS = { "require_msg": "a 'dep' pipeline must not contain an ner component", }, "core": { - "name": "core_news_sm", + "name": "core_news_{size}", "description": ( "Persian pipeline optimized for CPU. Components: tok2vec, tagger, morphologizer, " "trainable_lemmatizer, parser, ner. Entity labels: PER, LOC, ORG, DAT, MON, TIM, " @@ -119,7 +172,7 @@ VARIANTS = { "require_msg": "a 'core' pipeline must contain both parser and ner", }, "ent": { - "name": "ent_news_sm", + "name": "ent_news_{size}", "description": ( "Persian named entity recognizer optimized for CPU, with its own internal " "tok2vec. Labels: PER, LOC, ORG, DAT, MON, TIM, PCT." @@ -140,6 +193,10 @@ def main(): ap.add_argument("output", help="destination directory") ap.add_argument("--variant", choices=sorted(VARIANTS), required=True) ap.add_argument("--version", default="3.8.0") + ap.add_argument("--size", choices=("sm", "md", "lg", "trf"), default="sm", + help="size slot in the package name. 'md'/'lg' additionally record the " + "floret vector table as a source and append a vectors note; 'trf' " + "records the transformer source and appends a transformer note.") ap.add_argument("--ud-metrics", default=None, help="benchmark accuracy JSON scored on the UD test split; supplies the " "tagger/morph/lemma/parser keys only") @@ -152,7 +209,20 @@ def main(): args = ap.parse_args() spec = VARIANTS[args.variant] + name = spec["name"].format(size=args.size) + description = spec["description"].format(size=args.size) + sources = list(spec["sources"]) + notes = spec["notes"] nlp = spacy.load(args.model) + if args.size == "md": + sources.append(FLORET) + notes = " ".join([notes, VECTORS_NOTE]) + elif args.size == "lg": + sources.append(FLORET_LG) + notes = " ".join([notes, vectors_note_lg(nlp)]) + elif args.size == "trf": + sources.append(TRANSFORMER) + notes = " ".join([notes, TRANSFORMER_NOTE]) if args.add_ner: ner_nlp = spacy.load(args.add_ner) if ner_nlp.pipe_names != ["ner"]: @@ -202,22 +272,22 @@ def main(): nlp.meta.update( { "lang": "fa", - "name": spec["name"], + "name": name, "version": args.version, - "description": spec["description"], + "description": description, "author": AUTHOR, "email": EMAIL, "url": PROJECT_URL, "license": spec["license"], - "sources": spec["sources"], - "notes": spec["notes"], + "sources": sources, + "notes": notes, "performance": performance, } ) out = Path(args.output) nlp.to_disk(out) - print(f"wrote {out} as fa_{spec['name']} {args.version} ({spec['license']})") + print(f"wrote {out} as fa_{name} {args.version} ({spec['license']})") scalars = {k: round(v * 100, 2) for k, v in performance.items() if isinstance(v, float)} print(json.dumps(scalars, indent=2)) diff --git a/scripts/unpack_vectors.py b/scripts/unpack_vectors.py new file mode 100644 index 0000000..f3c86c4 --- /dev/null +++ b/scripts/unpack_vectors.py @@ -0,0 +1,53 @@ +"""Unpack a spaCy vectors-only wheel into a plain model directory. + +`spacy train --paths.vectors` wants a directory it can `spacy.load()`. The fa_floret wheel +already contains exactly that (an empty pipeline carrying only vocab/vectors), it is just +buried under the wheel's package layout, so this unwraps it rather than pip-installing a +package whose only job is to hold a 57 MB array. + +Usage: + python scripts/unpack_vectors.py fa_floret-0.1.0-py3-none-any-400k-documents.whl \\ + assets/vectors/fa_floret_400k +""" + +import argparse +import shutil +import sys +import tempfile +import zipfile +from pathlib import Path + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("wheel", type=Path) + ap.add_argument("output", type=Path) + args = ap.parse_args() + + with tempfile.TemporaryDirectory() as tmp: + tmp = Path(tmp) + with zipfile.ZipFile(args.wheel) as z: + z.extractall(tmp) + # The model directory is the one holding config.cfg, e.g. fa_floret/fa_floret-0.1.0/. + models = sorted(p.parent for p in tmp.rglob("config.cfg")) + if len(models) != 1: + sys.exit(f"expected exactly one config.cfg in {args.wheel}, found {len(models)}") + if args.output.exists(): + shutil.rmtree(args.output) + args.output.parent.mkdir(parents=True, exist_ok=True) + shutil.move(str(models[0]), str(args.output)) + + import spacy + + nlp = spacy.load(args.output) + vectors = nlp.vocab.vectors + if vectors.shape[0] == 0: + sys.exit(f"{args.output} has no vectors") + print( + f"{args.output}: mode={vectors.mode} shape={vectors.shape} " + f"n_keys={vectors.n_keys} pipeline={nlp.pipe_names}" + ) + + +if __name__ == "__main__": + main()