Split the pipeline into fa_dep_news_sm + fa_ent_news_sm, drop core
spaCy's naming scheme encodes contents: dep = tagger+parser+lemmatizer,
ent = NER only, core = both. Shipping a single fa_core_news_sm implied the NER
was held to the same standard as the rest of the pipeline. It is not, and the
numbers are not close:
UD components (PerDT, edited prose): TAG 95.96 LEMMA 97.91 LAS 85.15
ner (ParsTwiNER, tweets): ENTS_F 67.22
One package name and one version number would paper over that. So the treebank
components ship as fa_dep_news_sm (CC BY-SA 4.0, 7.5 MB) and NER ships as an
opt-in fa_ent_news_sm (MIT, 5.6 MB). Users now choose the weak component
deliberately instead of inheriting it.
Counting ParsTwiNER settles what "weak" means. It is not a small corpus --
232,917 tokens and 16,250 entities, the same order as the restricted ARMAN and
PEYMA. But the labels are skewed: PER 6258, LOC 5478, ORG 2694, NAT 939,
EVE 482, POG 399. The two starved labels are exactly the two scoring worst
(EVE 30.0, POG 41.2). Head labels suffer genre mismatch, tail labels suffer data
starvation -- two problems needing two different fixes. Recorded in README and
docs/MODELS.md.
fa_core_news_sm is reserved, not abandoned. ../ner_dataset/PLAN.md targets
prose-genre NER data that must beat ParsTwiNER on a human-annotated test set
before the name gets used. Nothing technical blocks the merge: the ner component
already embeds its own tok2vec instead of a Tok2VecListener, so
dep.add_pipe("ner", source=spacy.load("fa_ent_news_sm"))
reassembles a core-equivalent pipeline at runtime -- verified, not assumed.
Mechanics:
- scripts/assemble_core.py -> scripts/finalize_pipeline.py, now variant-aware
(dep|ent) and refusing to publish a dep pipeline containing an ner component,
so the split cannot silently regress.
- published meta.json["performance"] now comes from a strict whitelist. It was
inheriting raw *_loss values and a bogus tag_micro_f: 0.0 from training meta.
- project.yml split into two independent workflows, `all` and `ner`; all 16
commands dry-run clean.
- configs/fa_core_news_sm.cfg -> configs/fa_dep_news_sm.cfg.
- smoke_test.py degrades gracefully on pipelines lacking DEP/MORPH/ner.
Also folded into ../ner_dataset/PLAN.md: PerDT's XPOS encodes animacy on proper
nouns (N_ANM 6752, N_IANM 12682). Animate PROPN is a strong free prior for PER,
shrinking the annotation task to splitting inanimate PROPN into LOC/ORG, and
giving a gold-grounded cross-check that beats the model's self-reported
confidence for routing items to human review.
Committed unsigned: the OpenPGP smartcard holding 05E227BF4D6736DE is not
present (gpg: selecting card failed: No such device).
This commit is contained in:
parent
92fc1c3002
commit
41d5a46d96
141
README.md
141
README.md
|
|
@ -1,4 +1,4 @@
|
||||||
# fa_core_news_sm — a Persian pipeline for spaCy
|
# fa_dep_news_sm — a Persian pipeline for spaCy
|
||||||
|
|
||||||
There is no trained Persian pipeline for spaCy. `spacy.load("fa_core_news_sm")` has never
|
There is no trained Persian pipeline for spaCy. `spacy.load("fa_core_news_sm")` has never
|
||||||
worked; `spacy.blank("fa")` gives you a tokenizer and stop words and nothing else. This
|
worked; `spacy.blank("fa")` gives you a tokenizer and stop words and nothing else. This
|
||||||
|
|
@ -11,16 +11,36 @@ the result can actually be redistributed.
|
||||||
already has:** [`docs/CONTRIBUTING-GUIDE.md`](docs/CONTRIBUTING-GUIDE.md)
|
already has:** [`docs/CONTRIBUTING-GUIDE.md`](docs/CONTRIBUTING-GUIDE.md)
|
||||||
- **The build itself:** [`project.yml`](project.yml)
|
- **The build itself:** [`project.yml`](project.yml)
|
||||||
|
|
||||||
## The short version
|
## Two packages, not one
|
||||||
|
|
||||||
| | |
|
spaCy's naming scheme encodes what a pipeline contains: `dep` = tagger + parser + lemmatizer,
|
||||||
| --- | --- |
|
`ent` = NER only, `core` = both. This project ships the first two separately and deliberately
|
||||||
| Pipeline | `fa_core_news_sm` — tok2vec, tagger, morphologizer, trainable_lemmatizer, parser, ner |
|
does **not** ship a `core`:
|
||||||
| Syntax/morphology data | [UD_Persian-PerDT](https://github.com/UniversalDependencies/UD_Persian-PerDT) (PerUDT v1.0) — 29,107 sentences, **CC BY-SA 4.0** |
|
|
||||||
| NER data | [ParsTwiNER](https://github.com/overfit-ir/parstwiner) — 7,667 tweets, **MIT** |
|
| Package | Components | Trained on | Licence | Headline |
|
||||||
| Language data | `spacy/lang/fa` upstream (its stop word list comes from hazm) |
|
| --- | --- | --- | --- | --- |
|
||||||
| Licence of the result | CC BY-SA 4.0 (inherited from the treebank) |
|
| **`fa_dep_news_sm`** | tok2vec, tagger, morphologizer, trainable_lemmatizer, parser | [UD_Persian-PerDT](https://github.com/UniversalDependencies/UD_Persian-PerDT), 29,107 sentences of edited prose | CC BY-SA 4.0 | **LAS 85.15**, LEMMA 97.91 |
|
||||||
| Hardware | 4-core CPU. No GPU needed. |
|
| `fa_ent_news_sm` (optional) | ner | [ParsTwiNER](https://github.com/overfit-ir/parstwiner), 7,667 **tweets**, 16,250 entities | MIT | **ENTS_F 67.22** |
|
||||||
|
|
||||||
|
Those two headline numbers are the whole argument. The UD components score 85–98 on edited
|
||||||
|
prose; the NER manages 67 F on a different genre entirely, because the good Persian NER
|
||||||
|
corpora (ARMAN, PEYMA, NSURL) are research-use-only and cannot be redistributed. Folding both
|
||||||
|
into one `fa_core_news_sm` would hide that gap behind a single package name and a single
|
||||||
|
version number — users would reasonably assume the NER is held to the same standard as the
|
||||||
|
parser. It is not.
|
||||||
|
|
||||||
|
So NER ships as its own opt-in package, and **`fa_core_news_sm` is reserved** for when
|
||||||
|
[`../ner_dataset`](../ner_dataset/PLAN.md) delivers prose-genre NER data that beats ParsTwiNER
|
||||||
|
on a human-annotated test set. The `ner` component already embeds its own tok2vec rather than
|
||||||
|
a `Tok2VecListener` precisely so that merge is a one-liner when the data arrives:
|
||||||
|
|
||||||
|
```python
|
||||||
|
dep = spacy.load("fa_dep_news_sm")
|
||||||
|
dep.add_pipe("ner", source=spacy.load("fa_ent_news_sm")) # verified working
|
||||||
|
```
|
||||||
|
|
||||||
|
Language data comes from `spacy/lang/fa` upstream (its stop word list originally from hazm).
|
||||||
|
Everything trains on 4 CPU cores with no GPU.
|
||||||
|
|
||||||
### Results
|
### Results
|
||||||
|
|
||||||
|
|
@ -28,7 +48,9 @@ Trained and evaluated on this laptop (4-core i5-7200U, CPU only, 1h27m for the U
|
||||||
components, ~25 min for NER). Scores are on the **held-out test splits**, produced by
|
components, ~25 min for NER). Scores are on the **held-out test splits**, produced by
|
||||||
`spacy benchmark accuracy` and stored in `metrics/`.
|
`spacy benchmark accuracy` and stored in `metrics/`.
|
||||||
|
|
||||||
| Metric | `fa_core_news_sm` | reference |
|
**`fa_dep_news_sm`** (UD_Persian-PerDT test split):
|
||||||
|
|
||||||
|
| Metric | Score | reference |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `TOKEN_ACC` / `TOKEN_F` | 99.96 / 99.11 | |
|
| `TOKEN_ACC` / `TOKEN_F` | 99.96 / 99.11 | |
|
||||||
| `TAG_ACC` (XPOS) | **95.96** | |
|
| `TAG_ACC` (XPOS) | **95.96** | |
|
||||||
|
|
@ -38,39 +60,47 @@ components, ~25 min for NER). Scores are on the **held-out test splits**, produc
|
||||||
| `SENTS_F` | 99.25 | |
|
| `SENTS_F` | 99.25 | |
|
||||||
| `DEP_UAS` | **89.69** | hazm+ParsBERT: 92.46 |
|
| `DEP_UAS` | **89.69** | hazm+ParsBERT: 92.46 |
|
||||||
| `DEP_LAS` | **85.15** | hazm+ParsBERT: 89.34 |
|
| `DEP_LAS` | **85.15** | hazm+ParsBERT: 89.34 |
|
||||||
| `ENTS_P` / `ENTS_R` / `ENTS_F` | 74.77 / 61.06 / **67.22** | |
|
|
||||||
| Speed | ~9,250 words/s (CPU) | |
|
| Speed | ~9,250 words/s (CPU) | |
|
||||||
| Wheel size | 13 MB | `en_core_web_sm`: 12 MB |
|
| Wheel | 7.5 MB | `en_core_web_sm`: 12 MB |
|
||||||
|
|
||||||
Per-entity F: `LOC` 73.9, `PER` 69.1, `NAT` 63.2, `ORG` 59.3, `POG` 41.2, `EVE` 30.0.
|
**`fa_ent_news_sm`** (ParsTwiNER test split): `ENTS_P` 74.77 / `ENTS_R` 61.06 /
|
||||||
|
**`ENTS_F` 67.22**, 5.6 MB wheel. Per label:
|
||||||
|
`LOC` 73.9, `PER` 69.1, `NAT` 63.2, `ORG` 59.3, `POG` 41.2, `EVE` 30.0.
|
||||||
|
|
||||||
Read these honestly:
|
Read these honestly:
|
||||||
|
|
||||||
- **Parsing is 4.2 LAS behind hazm's parser**, which is the expected gap between a 13 MB
|
- **Parsing is 4.2 LAS behind hazm's parser**, which is the expected gap between a 7.5 MB
|
||||||
CPU model with hash embeddings and a fine-tuned ParsBERT. It is the same corpus and the
|
CPU model with hash embeddings and a fine-tuned ParsBERT. Same corpus, same spaCy parser
|
||||||
same spaCy parser architecture, so the comparison is fair, and it sets the target for the
|
architecture, so the comparison is fair — and it sets the target for a future `trf` tier.
|
||||||
future `trf` tier.
|
- **NER is the weak artifact, and the per-label numbers say why.** ParsTwiNER is not small
|
||||||
- **NER is the weak component.** 67 F reflects three compounding handicaps: an `sm` model
|
(232,917 tokens, 16,250 entities, 7.0% density — the same order as the restricted ARMAN and
|
||||||
with no static vectors, a 233k-token training corpus, and a genre mismatch (trained on
|
PEYMA). But its label distribution is brutally skewed: `PER` 6258, `LOC` 5478, `ORG` 2694,
|
||||||
tweets, most users will run it on prose). `EVE` and `POG` are near-useless. This is the
|
`NAT` 939, **`EVE` 482, `POG` 399**. The two starved labels are exactly the two that score
|
||||||
price of using the only MIT-licensed Persian NER corpus that exists.
|
30.0 and 41.2. So the head labels suffer from genre mismatch and the tail labels from raw
|
||||||
|
data starvation — two different problems needing two different fixes.
|
||||||
- **Everything else is competitive with the English `sm` pipeline** (`en_core_web_sm`:
|
- **Everything else is competitive with the English `sm` pipeline** (`en_core_web_sm`:
|
||||||
TAG 97, LAS 90, ENTS_F 84 — on a much larger and cleaner corpus).
|
TAG 97, LAS 90, ENTS_F 84 — on a much larger and cleaner corpus).
|
||||||
|
|
||||||
Reproduce: `.venv/bin/python -m spacy project run all`.
|
Reproduce: `.venv/bin/python -m spacy project run all` (add `run ner` for the NER package).
|
||||||
|
|
||||||
### Install the built pipeline
|
### Install
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
.venv/bin/python -m pip install packages/fa_core_news_sm-3.8.0/dist/fa_core_news_sm-3.8.0-py3-none-any.whl
|
.venv/bin/python -m pip install packages/fa_dep_news_sm-3.8.0/dist/fa_dep_news_sm-3.8.0-py3-none-any.whl
|
||||||
|
# optional, separate package:
|
||||||
|
.venv/bin/python -m pip install packages/fa_ent_news_sm-3.8.0/dist/fa_ent_news_sm-3.8.0-py3-none-any.whl
|
||||||
```
|
```
|
||||||
|
|
||||||
```python
|
```python
|
||||||
import spacy
|
import spacy
|
||||||
nlp = spacy.load("fa_core_news_sm")
|
nlp = spacy.load("fa_dep_news_sm")
|
||||||
doc = nlp("دانشگاه تهران در سال ۱۳۱۳ تأسیس شد.")
|
doc = nlp("دانشگاه تهران در سال ۱۳۱۳ تأسیس شد.")
|
||||||
print([(t.text, t.pos_, t.lemma_, t.dep_) for t in doc])
|
print([(t.text, t.pos_, t.lemma_, t.dep_) for t in doc])
|
||||||
print(doc.ents) # (دانشگاه تهران, ORG)
|
# ('دانشگاه', 'PROPN', 'دانشگاه', 'nsubj') ('تهران', 'PROPN', 'تهران', 'flat:name') ...
|
||||||
|
|
||||||
|
# Want entities too? Attach the NER package yourself, eyes open about its 67 F:
|
||||||
|
nlp.add_pipe("ner", source=spacy.load("fa_ent_news_sm"))
|
||||||
|
print(nlp(doc.text).ents) # (دانشگاه تهران,) -> ORG
|
||||||
```
|
```
|
||||||
|
|
||||||
### Why not hazm's own models
|
### Why not hazm's own models
|
||||||
|
|
@ -115,30 +145,36 @@ python -m venv .venv
|
||||||
|
|
||||||
## Build
|
## Build
|
||||||
|
|
||||||
Everything is driven by [`project.yml`](project.yml):
|
Everything is driven by [`project.yml`](project.yml), which has **two independent workflows**:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
.venv/bin/python -m spacy project assets # download + checksum the corpora
|
.venv/bin/python -m spacy project assets # download + checksum the corpora
|
||||||
.venv/bin/python -m spacy project run all # inspect -> convert -> train -> assemble -> evaluate -> package
|
.venv/bin/python -m spacy project run all # -> fa_dep_news_sm (the shipping artifact)
|
||||||
|
.venv/bin/python -m spacy project run ner # -> fa_ent_news_sm (optional)
|
||||||
```
|
```
|
||||||
|
|
||||||
Individual steps:
|
| Workflow | Command | What it does |
|
||||||
|
| --- | --- | --- |
|
||||||
| Command | What it does |
|
| `all` | `inspect` | annotation coverage of the treebanks (`scripts/inspect_treebanks.py`) |
|
||||||
| --- | --- |
|
| | `convert-ud` | CoNLL-U → `DocBin` with `--merge-subtokens`, plus the tokenizer-agreement report |
|
||||||
| `inspect` | annotation coverage of the treebanks (`scripts/inspect_treebanks.py`) |
|
| | `debug-data` | `spacy debug data` before spending CPU |
|
||||||
| `convert-ud` | CoNLL-U → `DocBin` with `--merge-subtokens`, plus the tokenizer-agreement report |
|
| | `train-core` | tagger + morphologizer + trainable_lemmatizer + parser on PerDT |
|
||||||
| `convert-ner` | unpack ParsTwiNER, IOB2 → `DocBin` |
|
| | `finalize` | write `fa_dep_news_sm` metadata: sources, licence, notes (`scripts/finalize_pipeline.py`) |
|
||||||
| `debug-data` | `spacy debug data` on both corpora before spending CPU |
|
| | `evaluate` | `spacy benchmark accuracy` on the held-out UD test split |
|
||||||
| `train-core` | tagger + morphologizer + trainable_lemmatizer + parser on PerDT |
|
| | `finalize-meta` | re-run `finalize`, folding test scores into `meta.json["performance"]` |
|
||||||
| `train-ner` | standalone `ner` with its own embedded tok2vec on ParsTwiNER |
|
| | `package` | build the wheel + sdist |
|
||||||
| `assemble` | source `ner` into the core pipeline, write full `meta.json` (`scripts/assemble_core.py`) |
|
| | `smoke` | run the pipeline over real Persian text and print every annotation layer |
|
||||||
| `evaluate` | `spacy benchmark accuracy` on both held-out test sets |
|
| `ner` | `convert-ner` | unpack ParsTwiNER, IOB2 → `DocBin` |
|
||||||
| `package` | build the wheel + sdist |
|
| | `debug-data-ner`, `train-ner`, `finalize-ner`, `evaluate-ner`, `package-ner` | the same sequence for `fa_ent_news_sm` |
|
||||||
| `smoke` | run the pipeline over real Persian text and print every annotation layer |
|
|
||||||
|
|
||||||
The two training runs are independent and can run concurrently — each is single-threaded.
|
The two training runs are independent and can run concurrently — each is single-threaded.
|
||||||
|
|
||||||
|
`finalize` and `finalize-meta` are separate steps for an unavoidable ordering reason: test
|
||||||
|
scores can only exist after `evaluate`, and `evaluate` needs a finalized pipeline to score.
|
||||||
|
Re-running `finalize` afterwards is cheap (it only copies models).
|
||||||
|
`scripts/finalize_pipeline.py` **refuses** to publish a `dep` pipeline containing an `ner`
|
||||||
|
component, so the split cannot silently regress.
|
||||||
|
|
||||||
## Design decisions worth knowing before you touch anything
|
## Design decisions worth knowing before you touch anything
|
||||||
|
|
||||||
1. **`--merge-subtokens`.** spaCy has no multiword-token layer, and PerDT splits pronominal
|
1. **`--merge-subtokens`.** spaCy has no multiword-token layer, and PerDT splits pronominal
|
||||||
|
|
@ -157,8 +193,17 @@ The two training runs are independent and can run concurrently — each is singl
|
||||||
|
|
||||||
## Roadmap
|
## Roadmap
|
||||||
|
|
||||||
`md`/`lg` need floret vectors trained on Persian Wikipedia + OSCAR (see
|
In value order, not difficulty order:
|
||||||
`spacy-vectors-builder`); floret rather than classic fastText because Persian's ZWNJ usage is
|
|
||||||
inconsistent and explodes the surface vocabulary. `trf` needs a rented GPU and should use
|
1. **Prose-genre NER** → [`../ner_dataset`](../ner_dataset/PLAN.md). This is what unlocks a real
|
||||||
`HooshvareLab/roberta-fa-zwnj-base` (Apache-2.0) rather than ParsBERT, whose model card
|
`fa_core_news_sm`. Note that PerDT's XPOS already encodes animacy on proper nouns
|
||||||
carries no licence. `senter` is one extra training run away. Details in `docs/MODELS.md` §2.
|
(`N_ANM` 6,752 vs `N_IANM` 12,682), which is a strong free prior for PER vs LOC/ORG.
|
||||||
|
2. **`md`/`lg`** need floret vectors trained on Persian Wikipedia + OSCAR (see
|
||||||
|
`spacy-vectors-builder`); floret rather than classic fastText because Persian's ZWNJ usage
|
||||||
|
is inconsistent and explodes the surface vocabulary.
|
||||||
|
3. **`trf`** needs a rented GPU and should use `HooshvareLab/roberta-fa-zwnj-base`
|
||||||
|
(Apache-2.0) rather than ParsBERT, whose model card carries no licence.
|
||||||
|
4. **`senter`** is one extra training run away.
|
||||||
|
5. **Upstream PRs** to `spacy/lang/fa` — see [`docs/upstream/fa-noun-chunks.md`](docs/upstream/fa-noun-chunks.md).
|
||||||
|
|
||||||
|
Details in [`docs/MODELS.md`](docs/MODELS.md) §2.
|
||||||
|
|
|
||||||
|
|
@ -1,13 +1,14 @@
|
||||||
# fa_core_news_sm — UD components (tagger, morphologizer, trainable_lemmatizer, parser).
|
# fa_dep_news_sm — tagger, morphologizer, trainable_lemmatizer, parser.
|
||||||
#
|
#
|
||||||
# Generated with:
|
# Generated with:
|
||||||
# spacy init config configs/fa_core_news_sm.cfg --lang fa \
|
# spacy init config configs/fa_dep_news_sm.cfg --lang fa \
|
||||||
# --pipeline tagger,morphologizer,trainable_lemmatizer,parser --optimize efficiency
|
# --pipeline tagger,morphologizer,trainable_lemmatizer,parser --optimize efficiency
|
||||||
#
|
#
|
||||||
# Deviations from the generated defaults:
|
# Deviations from the generated defaults:
|
||||||
# - eval_frequency 200 -> 400 (dev set is only 146 docs; 200 wastes CPU on a 4-core box)
|
# - eval_frequency 200 -> 400 (dev set is only 146 docs; 200 wastes CPU on a 4-core box)
|
||||||
# The `ner` component is trained separately (configs/fa_ner_sm.cfg) from a different
|
#
|
||||||
# corpus and merged in by scripts/assemble_core.py.
|
# No `ner` here by design. NER is a separate artifact built from a different corpus by the
|
||||||
|
# `ner` workflow (configs/fa_ner_sm.cfg -> fa_ent_news_sm); see project.yml for why.
|
||||||
|
|
||||||
[paths]
|
[paths]
|
||||||
train = null
|
train = null
|
||||||
|
|
@ -1,15 +1,19 @@
|
||||||
# fa NER component, CPU size (sm).
|
# fa_ent_news_sm — Persian NER, CPU size (sm).
|
||||||
#
|
#
|
||||||
# Trained on ParsTwiNER (MIT) — a different corpus from the UD treebank that trains
|
# A standalone artifact, not part of the shipping `dep` pipeline. It is trained on
|
||||||
# the rest of fa_core_news_sm, so this is a standalone run whose `ner` component is
|
# ParsTwiNER (MIT, tweets) while the rest of the project trains on UD_Persian-PerDT
|
||||||
# later sourced into the core pipeline by scripts/assemble_core.py.
|
# (CC BY-SA 4.0, edited prose), and it scores 67.22 F against 85-98 for the UD
|
||||||
|
# components — different corpus, different genre, different quality tier, so it gets its
|
||||||
|
# own package rather than being hidden inside a `core` one. Built by `spacy project run ner`.
|
||||||
#
|
#
|
||||||
# Deliberate deviation from `spacy init config --pipeline ner`: the tok2vec is
|
# Deliberate deviation from `spacy init config --pipeline ner`: the tok2vec is
|
||||||
# EMBEDDED inside components.ner.model instead of being a separate `tok2vec`
|
# EMBEDDED inside components.ner.model instead of being a separate `tok2vec`
|
||||||
# component with a Tok2VecListener. A listener can only resolve inside the pipeline
|
# component with a Tok2VecListener. A listener can only resolve inside the pipeline
|
||||||
# it was trained in; embedding makes the component self-contained and therefore
|
# it was trained in; embedding makes the component self-contained and therefore
|
||||||
# sourceable. This is the same design as en_core_web_sm, whose `ner` "has its own
|
# sourceable into another pipeline via `nlp.add_pipe("ner", source=...)`. This is the same
|
||||||
# independent internal tok2vec" (https://spacy.io/models#design).
|
# design as en_core_web_sm, whose `ner` "has its own independent internal tok2vec"
|
||||||
|
# (https://spacy.io/models#design). That is what will let a future fa_core_news_sm combine
|
||||||
|
# this component (or its ../ner_dataset replacement) with the dep pipeline.
|
||||||
|
|
||||||
[paths]
|
[paths]
|
||||||
train = null
|
train = null
|
||||||
|
|
|
||||||
|
|
@ -10,7 +10,7 @@ spaCy splits Persian support into two *completely different* contribution surfac
|
||||||
| Surface | What it is | Where it lives | How you contribute |
|
| Surface | What it is | Where it lives | How you contribute |
|
||||||
| --- | --- | --- | --- |
|
| --- | --- | --- | --- |
|
||||||
| **Language data** (`fa`) | Hand-written rules: tokenizer exceptions, stop words, `LIKE_NUM`, punctuation, noun-chunk iterator | `spacy/lang/fa/*.py` inside the spaCy repo | Normal PR to `explosion/spaCy` |
|
| **Language data** (`fa`) | Hand-written rules: tokenizer exceptions, stop words, `LIKE_NUM`, punctuation, noun-chunk iterator | `spacy/lang/fa/*.py` inside the spaCy repo | Normal PR to `explosion/spaCy` |
|
||||||
| **Trained pipeline** (`fa_core_news_sm`) | Statistical weights + `config.cfg` + `meta.json`, shipped as a pip wheel | `explosion/spacy-models` releases | **You cannot.** Publish it yourself (PyPI / HF Hub) and get it listed in spaCy Universe |
|
| **Trained pipeline** (`fa_dep_news_sm`) | Statistical weights + `config.cfg` + `meta.json`, shipped as a pip wheel | `explosion/spacy-models` releases | **You cannot.** Publish it yourself (PyPI / HF Hub) and get it listed in spaCy Universe |
|
||||||
|
|
||||||
`spacy/lang/fa` **already exists upstream**. What does not exist is any trained `fa` pipeline.
|
`spacy/lang/fa` **already exists upstream**. What does not exist is any trained `fa` pipeline.
|
||||||
So this project is a *publishing* project, not an upstream-PR project — with optional
|
So this project is a *publishing* project, not an upstream-PR project — with optional
|
||||||
|
|
@ -57,10 +57,10 @@ From <https://github.com/explosion/spaCy/blob/master/CONTRIBUTING.md>:
|
||||||
```bash
|
```bash
|
||||||
pip install spacy-huggingface-hub
|
pip install spacy-huggingface-hub
|
||||||
huggingface-cli login
|
huggingface-cli login
|
||||||
python -m spacy package training/model-best packages --name core_news_sm --version 3.8.0 --build wheel
|
python -m spacy package training/fa_dep_news_sm packages --name dep_news_sm --version 3.8.0 --build wheel
|
||||||
python -m spacy huggingface-hub push packages/fa_core_news_sm-3.8.0/dist/fa_core_news_sm-3.8.0-py3-none-any.whl --org <org>
|
python -m spacy huggingface-hub push packages/fa_dep_news_sm-3.8.0/dist/fa_dep_news_sm-3.8.0-py3-none-any.whl --org <org>
|
||||||
```
|
```
|
||||||
Users then `pip install https://huggingface.co/<org>/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl`.
|
Users then `pip install https://huggingface.co/<org>/fa_dep_news_sm/resolve/main/fa_dep_news_sm-any-py3-none-any.whl`.
|
||||||
2. **PyPI / self-hosted wheel** — `spacy package … --build sdist,wheel` then `twine upload`,
|
2. **PyPI / self-hosted wheel** — `spacy package … --build sdist,wheel` then `twine upload`,
|
||||||
or attach the wheel to a GitHub Release. See <https://spacy.io/api/cli#package>.
|
or attach the wheel to a GitHub Release. See <https://spacy.io/api/cli#package>.
|
||||||
3. **spaCy Universe listing** — lists the package on spacy.io but hosts nothing. Per
|
3. **spaCy Universe listing** — lists the package on spacy.io but hosts nothing. Per
|
||||||
|
|
@ -77,7 +77,7 @@ From <https://github.com/explosion/spaCy/blob/master/CONTRIBUTING.md>:
|
||||||
From <https://spacy.io/models#conventions> and the `spacy-models` README:
|
From <https://spacy.io/models#conventions> and the `spacy-models` README:
|
||||||
|
|
||||||
```
|
```
|
||||||
[lang]_[type]_[genre]_[size] e.g. fa_core_news_sm
|
[lang]_[type]_[genre]_[size] e.g. fa_dep_news_sm
|
||||||
```
|
```
|
||||||
|
|
||||||
| Slot | Allowed values | Meaning |
|
| Slot | Allowed values | Meaning |
|
||||||
|
|
@ -139,7 +139,7 @@ python -m spacy benchmark accuracy training/model-best corpus/test.spacy --outpu
|
||||||
# ('spacy evaluate' is now just an alias for 'benchmark accuracy')
|
# ('spacy evaluate' is now just an alias for 'benchmark accuracy')
|
||||||
|
|
||||||
# 7. package
|
# 7. package
|
||||||
python -m spacy package training/model-best packages --name core_news_sm --version 3.8.0 --build sdist,wheel
|
python -m spacy package training/fa_dep_news_sm packages --name dep_news_sm --version 3.8.0 --build sdist,wheel
|
||||||
```
|
```
|
||||||
|
|
||||||
`spacy assemble` builds a pipeline from a config **without training** — useful for a
|
`spacy assemble` builds a pipeline from a config **without training** — useful for a
|
||||||
|
|
|
||||||
|
|
@ -22,19 +22,31 @@ parsing** (identical to two decimals) and ~1–2 F on NER. The transformer buys
|
||||||
|
|
||||||
Sources: <https://spacy.io/models/en>.
|
Sources: <https://spacy.io/models/en>.
|
||||||
|
|
||||||
## 2. Target: the four Persian pipelines
|
## 2. Target: the Persian pipelines
|
||||||
|
|
||||||
Naming follows `[lang]_[type]_[genre]_[size]` (<https://spacy.io/models#conventions>).
|
Naming follows `[lang]_[type]_[genre]_[size]` (<https://spacy.io/models#conventions>), and the
|
||||||
`core` = tagger + parser + lemmatizer + NER. Genre is `news`, after the dominant genre of
|
`type` slot is load-bearing: `dep` = tagger + parser + lemmatizer, `ent` = NER only,
|
||||||
UD_Persian-PerDT (its README lists "news fiction nonfiction academic web blog", and spaCy
|
`core` = both. Genre is `news`, after the dominant genre of UD_Persian-PerDT (its README lists
|
||||||
labels comparable treebank-trained pipelines such as `de_core_news_sm` as `news`).
|
"news fiction nonfiction academic web blog", and spaCy labels comparable treebank-trained
|
||||||
|
pipelines such as `de_core_news_sm` as `news`).
|
||||||
|
|
||||||
| Pipeline | Components | Embeddings | Trainable on this hardware? |
|
| Pipeline | Components | Embeddings | Status |
|
||||||
| --- | --- | --- | --- |
|
| --- | --- | --- | --- |
|
||||||
| **`fa_core_news_sm`** | tok2vec, tagger, morphologizer, trainable_lemmatizer, parser, ner | hash embeddings | **yes — this is what we train now** |
|
| **`fa_dep_news_sm`** | tok2vec, tagger, morphologizer, trainable_lemmatizer, parser | hash embeddings | **built — the shipping artifact** |
|
||||||
| `fa_core_news_md` | same | floret vectors, 50k rows | yes, but vectors must be trained first (CPU-days on fa Wikipedia + OSCAR) |
|
| `fa_ent_news_sm` | ner (own internal tok2vec) | hash embeddings | **built — optional, separate package** |
|
||||||
| `fa_core_news_lg` | same | floret vectors, 200k rows | same as md, bigger table |
|
| `fa_core_news_sm` | the two above, merged | hash embeddings | **reserved.** Blocked on prose-genre NER data from `../ner_dataset` |
|
||||||
| `fa_core_news_trf` | transformer, tagger, morphologizer, trainable_lemmatizer, parser, ner | `HooshvareLab/roberta-fa-zwnj-base` (Apache-2.0) | **no** — 2 GB VRAM (GTX 940MX) cannot fine-tune a 125M-param encoder; needs rented GPU |
|
| `fa_core_news_md` | + static vectors | floret, 50k rows | vectors must be trained first (CPU-days on fa Wikipedia + OSCAR) |
|
||||||
|
| `fa_core_news_lg` | same | floret, 200k rows | same as md, bigger table |
|
||||||
|
| `fa_core_news_trf` | transformer instead of tok2vec | `HooshvareLab/roberta-fa-zwnj-base` (Apache-2.0) | **not on this hardware** — 2 GB VRAM cannot fine-tune a 125M-param encoder |
|
||||||
|
|
||||||
|
**Why `dep` + `ent` rather than a single `core`.** `core` is a promise that the NER is part of
|
||||||
|
the same pipeline, built to the same standard, versioned together. Ours is not: the UD
|
||||||
|
components score 85–98 on edited prose, while the NER scores 67.22 F and is trained on tweets
|
||||||
|
because the good Persian NER corpora are research-use-only. One package name and one version
|
||||||
|
number would paper over a gap of that size. Shipping two packages makes the user opt into the
|
||||||
|
weak component knowingly, and costs nothing technically — `fa_ent_news_sm` embeds its own
|
||||||
|
tok2vec, so `nlp.add_pipe("ner", source=...)` reassembles a `core`-equivalent pipeline at
|
||||||
|
runtime (verified). `fa_core_news_sm` gets published when the NER earns the name.
|
||||||
|
|
||||||
One deviation from the English design, deliberate: Persian gets a **`morphologizer`**
|
One deviation from the English design, deliberate: Persian gets a **`morphologizer`**
|
||||||
(UPOS + morphological features) and a **`trainable_lemmatizer`** instead of English's
|
(UPOS + morphological features) and a **`trainable_lemmatizer`** instead of English's
|
||||||
|
|
@ -163,7 +175,7 @@ So what hazm genuinely contributes to this project is **one validated design dec
|
||||||
(PerDT is the corpus) and **the stop-word list already vendored into `spacy/lang/fa`**.
|
(PerDT is the corpus) and **the stop-word list already vendored into `spacy/lang/fa`**.
|
||||||
Everything else is built with spaCy-native tooling.
|
Everything else is built with spaCy-native tooling.
|
||||||
|
|
||||||
## 5. Decision: train `fa_core_news_sm` first
|
## 5. Decision: train the `sm` tier first
|
||||||
|
|
||||||
Not md/lg: those need floret vectors trained from scratch on Wikipedia+OSCAR (CPU-days) and
|
Not md/lg: those need floret vectors trained from scratch on Wikipedia+OSCAR (CPU-days) and
|
||||||
buy ~0.00 tag/dep accuracy in the English reference numbers.
|
buy ~0.00 tag/dep accuracy in the English reference numbers.
|
||||||
|
|
@ -194,7 +206,9 @@ lemmas that come out as two words.
|
||||||
The better long-term fix is Persian clitic-splitting suffix rules in `spacy/lang/fa`, which
|
The better long-term fix is Persian clitic-splitting suffix rules in `spacy/lang/fa`, which
|
||||||
would be an upstream PR, not a model change. Recorded in `docs/CONTRIBUTING-GUIDE.md` §5.
|
would be an upstream PR, not a model change. Recorded in `docs/CONTRIBUTING-GUIDE.md` §5.
|
||||||
|
|
||||||
### Final composition of `fa_core_news_sm`
|
### Final composition
|
||||||
|
|
||||||
|
**`fa_dep_news_sm`** — CC BY-SA 4.0, 7.5 MB wheel:
|
||||||
|
|
||||||
| Component | Trained on | Metric | Test score |
|
| Component | Trained on | Metric | Test score |
|
||||||
| --- | --- | --- | --- |
|
| --- | --- | --- | --- |
|
||||||
|
|
@ -203,17 +217,30 @@ would be an upstream PR, not a model change. Recorded in `docs/CONTRIBUTING-GUID
|
||||||
| `morphologizer` (UPOS + FEATS) | PerDT, 298 labels | `pos_acc` / `morph_acc` | 96.24 / 96.29 |
|
| `morphologizer` (UPOS + FEATS) | PerDT, 298 labels | `pos_acc` / `morph_acc` | 96.24 / 96.29 |
|
||||||
| `trainable_lemmatizer` | PerDT, 1,908 edit trees | `lemma_acc` | 97.91 |
|
| `trainable_lemmatizer` | PerDT, 1,908 edit trees | `lemma_acc` | 97.91 |
|
||||||
| `parser` | PerDT, 34 deprels | `dep_uas` / `dep_las` | 89.69 / 85.15 |
|
| `parser` | PerDT, 34 deprels | `dep_uas` / `dep_las` | 89.69 / 85.15 |
|
||||||
| `ner` (own internal tok2vec) | ParsTwiNER, 6 labels | `ents_p/r/f` | 74.77 / 61.06 / 67.22 |
|
|
||||||
|
**`fa_ent_news_sm`** — MIT, 5.6 MB wheel, separate package:
|
||||||
|
|
||||||
|
| Component | Trained on | Metric | Test score |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `ner` (own internal tok2vec) | ParsTwiNER, 6 labels | `ents_p/r/f` | 74.77 / 61.06 / **67.22** |
|
||||||
|
|
||||||
|
ParsTwiNER's label counts explain that last row better than any prose can:
|
||||||
|
`PER` 6258, `LOC` 5478, `ORG` 2694, `NAT` 939, **`EVE` 482, `POG` 399** over 232,917 tokens
|
||||||
|
(16,250 entities, 7.0% density). The two starved labels are exactly the two that score worst
|
||||||
|
(`EVE` 30.0, `POG` 41.2). So the corpus is not small — it is skewed and out of genre, which
|
||||||
|
are two different problems: the head labels need in-genre data, the tail labels need more data
|
||||||
|
of any kind.
|
||||||
|
|
||||||
Training cost on the target hardware (4-core i5-7200U, no GPU): **1h27m** for the UD
|
Training cost on the target hardware (4-core i5-7200U, no GPU): **1h27m** for the UD
|
||||||
components (early-stopped at step 10,800; best checkpoint step 9,200) and ~25 min for NER
|
components (early-stopped at step 10,800; best checkpoint step 9,200) and ~25 min for NER
|
||||||
(best at step 4,000). Both runs are single-threaded, so they were run concurrently.
|
(best at step 4,000). Both runs are single-threaded, so they were run concurrently.
|
||||||
Inference: ~9,250 words/s. Wheel: 13 MB.
|
Inference: ~9,250 words/s.
|
||||||
|
|
||||||
The `--merge-subtokens` artefacts predicted above are visible in the shipped model exactly as
|
The `--merge-subtokens` artefacts predicted above are visible in the shipped model exactly as
|
||||||
expected — `کتابهایش` ("his/her books") comes out as one token tagged `N_IANM_PR_JOPER` with
|
expected — `کتابهایش` ("his/her books") comes out as one token tagged `N_IANM_PR_JOPER` with
|
||||||
lemma `کتاب او`. Worth knowing before you consume `token.lemma_` downstream.
|
lemma `کتاب او`. Worth knowing before you consume `token.lemma_` downstream.
|
||||||
|
|
||||||
Sources recorded in `meta.json`, both with their licenses, per
|
Sources recorded in each `meta.json` with their licences, per
|
||||||
<https://spacy.io/api/data-formats#meta>. CC BY-SA 4.0 on PerDT means the packaged pipeline
|
<https://spacy.io/api/data-formats#meta>. CC BY-SA 4.0 on PerDT means `fa_dep_news_sm` must
|
||||||
must carry attribution and share-alike notice — handled in `scripts/assemble_core.py`.
|
carry attribution and share-alike notice — handled in `scripts/finalize_pipeline.py`, which
|
||||||
|
also refuses to publish a `dep` pipeline that contains an `ner` component.
|
||||||
|
|
|
||||||
131
project.yml
131
project.yml
|
|
@ -1,19 +1,27 @@
|
||||||
title: "fa_core_news_sm"
|
title: "fa_dep_news_sm"
|
||||||
description: >
|
description: >
|
||||||
A CPU-sized Persian (fa) core pipeline for spaCy 3.8: tagger (XPOS), morphologizer
|
A CPU-sized Persian (fa) dependency pipeline for spaCy 3.8: tagger (XPOS),
|
||||||
(UPOS + FEATS), trainable lemmatizer, dependency parser and NER.
|
morphologizer (UPOS + FEATS), trainable lemmatizer and dependency parser. Trained on
|
||||||
|
UD_Persian-PerDT (PerUDT v1.0, CC BY-SA 4.0).
|
||||||
|
|
||||||
UD components are trained on UD_Persian-PerDT (PerUDT v1.0, CC BY-SA 4.0). The NER
|
This repo deliberately ships `dep`, not `core`. In spaCy's naming scheme `core` means
|
||||||
component is trained separately on ParsTwiNER (MIT) and merged in, because no
|
"tagger + parser + lemmatizer + NER" in one package, and the only redistributably-licensed
|
||||||
redistributably-licensed Persian NER corpus shares a genre with the treebank. See
|
Persian NER corpus (ParsTwiNER, MIT) is a Twitter corpus — it scores 67.22 F against
|
||||||
docs/MODELS.md for the full source/licence analysis and docs/CONTRIBUTING-GUIDE.md
|
85-98 for the UD components, and folding it into a single `core` artifact would hide that
|
||||||
for how this gets published.
|
behind one package name. So NER ships separately and optionally as `fa_ent_news_sm`
|
||||||
|
(workflow: `ner`), and `fa_core_news_sm` is reserved for when ../ner_dataset delivers
|
||||||
|
prose-genre NER data that beats ParsTwiNER on a human-annotated test set.
|
||||||
|
|
||||||
Run everything with: `spacy project run all`
|
See docs/MODELS.md for the source/licence analysis and docs/CONTRIBUTING-GUIDE.md for
|
||||||
|
how this gets published.
|
||||||
|
|
||||||
|
Run the shipping pipeline with: `spacy project run all`
|
||||||
|
Optionally build the standalone NER with: `spacy project run ner`
|
||||||
|
|
||||||
vars:
|
vars:
|
||||||
lang: "fa"
|
lang: "fa"
|
||||||
package_name: "core_news_sm"
|
package_name: "dep_news_sm"
|
||||||
|
ner_package_name: "ent_news_sm"
|
||||||
package_version: "3.8.0"
|
package_version: "3.8.0"
|
||||||
treebank: "fa_perdt"
|
treebank: "fa_perdt"
|
||||||
# -1 = CPU. A GTX 940MX (2 GB) is not worth the transfer overhead for an sm pipeline.
|
# -1 = CPU. A GTX 940MX (2 GB) is not worth the transfer overhead for an sm pipeline.
|
||||||
|
|
@ -48,18 +56,25 @@ assets:
|
||||||
description: "ParsTwiNER Persian Twitter NER corpus, IOB2 (MIT)"
|
description: "ParsTwiNER Persian Twitter NER corpus, IOB2 (MIT)"
|
||||||
|
|
||||||
workflows:
|
workflows:
|
||||||
|
# The shipping artifact: fa_dep_news_sm.
|
||||||
all:
|
all:
|
||||||
- inspect
|
- inspect
|
||||||
- convert-ud
|
- convert-ud
|
||||||
- convert-ner
|
|
||||||
- debug-data
|
- debug-data
|
||||||
- train-core
|
- train-core
|
||||||
- train-ner
|
- finalize
|
||||||
- assemble
|
|
||||||
- evaluate
|
- evaluate
|
||||||
- finalize-meta
|
- finalize-meta
|
||||||
- package
|
- package
|
||||||
- smoke
|
- smoke
|
||||||
|
# Optional, separate artifact: fa_ent_news_sm. Not part of `all` — see the note above.
|
||||||
|
ner:
|
||||||
|
- convert-ner
|
||||||
|
- debug-data-ner
|
||||||
|
- train-ner
|
||||||
|
- finalize-ner
|
||||||
|
- evaluate-ner
|
||||||
|
- package-ner
|
||||||
|
|
||||||
commands:
|
commands:
|
||||||
- name: "inspect"
|
- name: "inspect"
|
||||||
|
|
@ -107,24 +122,29 @@ commands:
|
||||||
- "corpus/ner/test.spacy"
|
- "corpus/ner/test.spacy"
|
||||||
|
|
||||||
- name: "debug-data"
|
- name: "debug-data"
|
||||||
help: "Validate both corpora against their configs before burning CPU on training"
|
help: "Validate the treebank against the config before burning CPU on training"
|
||||||
script:
|
script:
|
||||||
- "python -m spacy debug data configs/fa_core_news_sm.cfg --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy"
|
- "python -m spacy debug data configs/fa_dep_news_sm.cfg --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy"
|
||||||
- "python -m spacy debug data configs/fa_ner_sm.cfg --paths.train corpus/ner/train.spacy --paths.dev corpus/ner/dev.spacy"
|
|
||||||
deps:
|
deps:
|
||||||
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
||||||
|
- "configs/fa_dep_news_sm.cfg"
|
||||||
|
|
||||||
|
- name: "debug-data-ner"
|
||||||
|
help: "Validate the ParsTwiNER corpus against the NER config"
|
||||||
|
script:
|
||||||
|
- "python -m spacy debug data configs/fa_ner_sm.cfg --paths.train corpus/ner/train.spacy --paths.dev corpus/ner/dev.spacy"
|
||||||
|
deps:
|
||||||
- "corpus/ner/train.spacy"
|
- "corpus/ner/train.spacy"
|
||||||
- "configs/fa_core_news_sm.cfg"
|
|
||||||
- "configs/fa_ner_sm.cfg"
|
- "configs/fa_ner_sm.cfg"
|
||||||
|
|
||||||
- name: "train-core"
|
- name: "train-core"
|
||||||
help: "Train tok2vec + tagger + morphologizer + trainable_lemmatizer + parser on PerDT"
|
help: "Train tok2vec + tagger + morphologizer + trainable_lemmatizer + parser on PerDT"
|
||||||
script:
|
script:
|
||||||
- "python -m spacy train configs/fa_core_news_sm.cfg --output training/core --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy --gpu-id ${vars.gpu}"
|
- "python -m spacy train configs/fa_dep_news_sm.cfg --output training/core --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy --gpu-id ${vars.gpu}"
|
||||||
deps:
|
deps:
|
||||||
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
||||||
- "corpus/merged/${vars.treebank}-ud-dev.spacy"
|
- "corpus/merged/${vars.treebank}-ud-dev.spacy"
|
||||||
- "configs/fa_core_news_sm.cfg"
|
- "configs/fa_dep_news_sm.cfg"
|
||||||
outputs:
|
outputs:
|
||||||
- "training/core/model-best"
|
- "training/core/model-best"
|
||||||
|
|
||||||
|
|
@ -139,57 +159,82 @@ commands:
|
||||||
outputs:
|
outputs:
|
||||||
- "training/ner/model-best"
|
- "training/ner/model-best"
|
||||||
|
|
||||||
- name: "assemble"
|
- name: "finalize"
|
||||||
help: "Source the trained ner into the core pipeline and write full meta.json"
|
help: "Write fa_dep_news_sm metadata (sources, licence, notes) onto the trained model"
|
||||||
script:
|
script:
|
||||||
- "python scripts/assemble_core.py training/core/model-best training/ner/model-best training/fa_core_news_sm --version ${vars.package_version}"
|
- "python scripts/finalize_pipeline.py training/core/model-best training/fa_dep_news_sm --variant dep --version ${vars.package_version}"
|
||||||
deps:
|
deps:
|
||||||
- "training/core/model-best"
|
- "training/core/model-best"
|
||||||
- "training/ner/model-best"
|
- "scripts/finalize_pipeline.py"
|
||||||
- "scripts/assemble_core.py"
|
|
||||||
outputs:
|
outputs:
|
||||||
- "training/fa_core_news_sm"
|
- "training/fa_dep_news_sm"
|
||||||
|
|
||||||
- name: "evaluate"
|
- name: "evaluate"
|
||||||
help: "Score the assembled pipeline on both held-out test sets"
|
help: "Score fa_dep_news_sm on the held-out UD test split"
|
||||||
script:
|
script:
|
||||||
- "python -m spacy benchmark accuracy training/fa_core_news_sm corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/ud-test.json --gpu-id ${vars.gpu}"
|
- "python -m spacy benchmark accuracy training/fa_dep_news_sm corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/ud-test.json --gpu-id ${vars.gpu}"
|
||||||
- "python -m spacy benchmark accuracy training/fa_core_news_sm corpus/ner/test.spacy --output metrics/ner-test.json --gpu-id ${vars.gpu}"
|
|
||||||
deps:
|
deps:
|
||||||
- "training/fa_core_news_sm"
|
- "training/fa_dep_news_sm"
|
||||||
- "corpus/merged/${vars.treebank}-ud-test.spacy"
|
- "corpus/merged/${vars.treebank}-ud-test.spacy"
|
||||||
- "corpus/ner/test.spacy"
|
|
||||||
outputs:
|
outputs:
|
||||||
- "metrics/ud-test.json"
|
- "metrics/ud-test.json"
|
||||||
- "metrics/ner-test.json"
|
|
||||||
|
|
||||||
- name: "finalize-meta"
|
- name: "finalize-meta"
|
||||||
help: >
|
help: >
|
||||||
Re-assemble, this time folding the test scores into meta.json["performance"].
|
Re-run finalize, this time folding the test scores into meta.json["performance"].
|
||||||
Separate from `assemble` because the scores can only exist after `evaluate`, and
|
Separate from `finalize` because the scores can only exist after `evaluate`, and
|
||||||
`evaluate` needs an assembled pipeline to score. Cheap: it only copies models.
|
`evaluate` needs a finalized pipeline to score. Cheap: it only copies models.
|
||||||
script:
|
script:
|
||||||
- "python scripts/assemble_core.py training/core/model-best training/ner/model-best training/fa_core_news_sm --version ${vars.package_version} --ud-metrics metrics/ud-test.json --ner-metrics metrics/ner-test.json"
|
- "python scripts/finalize_pipeline.py training/core/model-best training/fa_dep_news_sm --variant dep --version ${vars.package_version} --metrics metrics/ud-test.json"
|
||||||
deps:
|
deps:
|
||||||
- "metrics/ud-test.json"
|
- "metrics/ud-test.json"
|
||||||
- "metrics/ner-test.json"
|
- "scripts/finalize_pipeline.py"
|
||||||
- "scripts/assemble_core.py"
|
|
||||||
|
|
||||||
- name: "package"
|
- name: "package"
|
||||||
help: "Build the installable wheel + sdist"
|
help: "Build the installable fa_dep_news_sm wheel + sdist"
|
||||||
script:
|
script:
|
||||||
- "python -m spacy package training/fa_core_news_sm packages --name ${vars.package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
- "python -m spacy package training/fa_dep_news_sm packages --name ${vars.package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||||
deps:
|
deps:
|
||||||
- "training/fa_core_news_sm"
|
- "training/fa_dep_news_sm"
|
||||||
outputs:
|
outputs:
|
||||||
- "packages/${vars.lang}_${vars.package_name}-${vars.package_version}"
|
- "packages/${vars.lang}_${vars.package_name}-${vars.package_version}"
|
||||||
|
|
||||||
- name: "smoke"
|
- name: "smoke"
|
||||||
help: "Load the packaged pipeline and run it over real Persian text"
|
help: "Load the pipeline and run it over real Persian text"
|
||||||
script:
|
script:
|
||||||
- "python scripts/smoke_test.py training/fa_core_news_sm"
|
- "python scripts/smoke_test.py training/fa_dep_news_sm"
|
||||||
deps:
|
deps:
|
||||||
- "training/fa_core_news_sm"
|
- "training/fa_dep_news_sm"
|
||||||
|
|
||||||
|
- name: "finalize-ner"
|
||||||
|
help: "Write fa_ent_news_sm metadata onto the trained NER model"
|
||||||
|
script:
|
||||||
|
- "python scripts/finalize_pipeline.py training/ner/model-best training/fa_ent_news_sm --variant ent --version ${vars.package_version} --metrics metrics/ner-test.json"
|
||||||
|
deps:
|
||||||
|
- "training/ner/model-best"
|
||||||
|
- "scripts/finalize_pipeline.py"
|
||||||
|
outputs:
|
||||||
|
- "training/fa_ent_news_sm"
|
||||||
|
|
||||||
|
- name: "evaluate-ner"
|
||||||
|
help: "Score fa_ent_news_sm on the held-out ParsTwiNER test split"
|
||||||
|
script:
|
||||||
|
- "python -m spacy benchmark accuracy training/fa_ent_news_sm corpus/ner/test.spacy --output metrics/ner-test.json --gpu-id ${vars.gpu}"
|
||||||
|
- "python scripts/finalize_pipeline.py training/ner/model-best training/fa_ent_news_sm --variant ent --version ${vars.package_version} --metrics metrics/ner-test.json"
|
||||||
|
deps:
|
||||||
|
- "training/fa_ent_news_sm"
|
||||||
|
- "corpus/ner/test.spacy"
|
||||||
|
outputs:
|
||||||
|
- "metrics/ner-test.json"
|
||||||
|
|
||||||
|
- name: "package-ner"
|
||||||
|
help: "Build the installable fa_ent_news_sm wheel + sdist"
|
||||||
|
script:
|
||||||
|
- "python -m spacy package training/fa_ent_news_sm packages --name ${vars.ner_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||||
|
deps:
|
||||||
|
- "training/fa_ent_news_sm"
|
||||||
|
outputs:
|
||||||
|
- "packages/${vars.lang}_${vars.ner_package_name}-${vars.package_version}"
|
||||||
|
|
||||||
- name: "clean"
|
- name: "clean"
|
||||||
help: "Drop corpora, training runs and metrics (keeps downloaded assets)"
|
help: "Drop corpora, training runs and metrics (keeps downloaded assets)"
|
||||||
|
|
|
||||||
|
|
@ -1,133 +0,0 @@
|
||||||
"""Merge the separately-trained NER component into the UD pipeline and write full metadata.
|
|
||||||
|
|
||||||
The two components cannot be trained together: the UD components come from
|
|
||||||
UD_Persian-PerDT (CC BY-SA 4.0, edited prose) and `ner` comes from ParsTwiNER (MIT,
|
|
||||||
tweets). `ner` was therefore configured with its own embedded tok2vec (see
|
|
||||||
configs/fa_ner_sm.cfg) so it can be sourced into another pipeline without a dangling
|
|
||||||
Tok2VecListener — the same design as en_core_web_sm.
|
|
||||||
|
|
||||||
Run it twice: once right after training (no metrics yet), and again after
|
|
||||||
`spacy benchmark accuracy` has produced metrics/*.json so `meta.json["performance"]`
|
|
||||||
reflects the assembled pipeline on the held-out test sets.
|
|
||||||
|
|
||||||
Usage:
|
|
||||||
.venv/bin/python scripts/assemble_core.py training/core/model-best training/ner/model-best \\
|
|
||||||
training/fa_core_news_sm --version 3.8.0 \\
|
|
||||||
[--ud-metrics metrics/ud-test.json] [--ner-metrics metrics/ner-test.json]
|
|
||||||
"""
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import json
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
import spacy
|
|
||||||
|
|
||||||
DESCRIPTION = (
|
|
||||||
"Persian pipeline optimized for CPU. Components: tok2vec, tagger, morphologizer, "
|
|
||||||
"trainable_lemmatizer, parser, ner."
|
|
||||||
)
|
|
||||||
|
|
||||||
SOURCES = [
|
|
||||||
{
|
|
||||||
"name": "UD_Persian-PerDT (PerUDT v1.0)",
|
|
||||||
"url": "https://github.com/UniversalDependencies/UD_Persian-PerDT",
|
|
||||||
"author": (
|
|
||||||
"Mohammad Sadegh Rasooli, Pegah Safari, Amirsaeid Moloodi, Alireza Nourian"
|
|
||||||
),
|
|
||||||
"license": "CC BY-SA 4.0",
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "ParsTwiNER",
|
|
||||||
"url": "https://github.com/overfit-ir/parstwiner",
|
|
||||||
"author": "MohammadMahdi Aghajani, AliAkbar Badri, Hamid Beigy et al. (Overfit-IR)",
|
|
||||||
"license": "MIT",
|
|
||||||
},
|
|
||||||
{
|
|
||||||
"name": "spaCy lang/fa language data (stop words originally from HAZM)",
|
|
||||||
"url": "https://github.com/explosion/spaCy/tree/master/spacy/lang/fa",
|
|
||||||
"author": "Explosion and spaCy contributors",
|
|
||||||
"license": "MIT",
|
|
||||||
},
|
|
||||||
]
|
|
||||||
|
|
||||||
# CC BY-SA 4.0 on the treebank propagates to anything derived from it.
|
|
||||||
NOTES = (
|
|
||||||
"The tagger, morphologizer, trainable_lemmatizer and parser are trained on "
|
|
||||||
"UD_Persian-PerDT, which is licensed CC BY-SA 4.0; this pipeline is therefore "
|
|
||||||
"distributed under CC BY-SA 4.0 with attribution to the treebank authors. "
|
|
||||||
"The ner component is trained on ParsTwiNER (MIT), a Twitter corpus, so entity "
|
|
||||||
"recognition is weaker on formal/edited prose than on social media text. "
|
|
||||||
"Multiword tokens in the treebank (pronominal clitics, enclitic copulas) were merged "
|
|
||||||
"with `spacy convert --merge-subtokens`, so a small number of XPOS tags are composite "
|
|
||||||
"(e.g. N_IANM_PR_JOPER) and ~1.5% of lemmas contain a space."
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def load_metrics(path):
|
|
||||||
if not path:
|
|
||||||
return {}
|
|
||||||
p = Path(path)
|
|
||||||
if not p.exists():
|
|
||||||
print(f" (no metrics at {p}, skipping)")
|
|
||||||
return {}
|
|
||||||
return json.loads(p.read_text(encoding="utf8"))
|
|
||||||
|
|
||||||
|
|
||||||
def main():
|
|
||||||
ap = argparse.ArgumentParser()
|
|
||||||
ap.add_argument("core", help="trained UD pipeline (training/core/model-best)")
|
|
||||||
ap.add_argument("ner", help="trained NER pipeline (training/ner/model-best)")
|
|
||||||
ap.add_argument("output", help="destination directory for the assembled pipeline")
|
|
||||||
ap.add_argument("--version", default="3.8.0")
|
|
||||||
ap.add_argument("--ud-metrics", default=None)
|
|
||||||
ap.add_argument("--ner-metrics", default=None)
|
|
||||||
args = ap.parse_args()
|
|
||||||
|
|
||||||
nlp = spacy.load(args.core)
|
|
||||||
ner_nlp = spacy.load(args.ner)
|
|
||||||
if "ner" in nlp.pipe_names:
|
|
||||||
nlp.remove_pipe("ner")
|
|
||||||
nlp.add_pipe("ner", source=ner_nlp)
|
|
||||||
print(f"pipeline: {nlp.pipe_names}")
|
|
||||||
print(f"ner labels: {sorted(nlp.get_pipe('ner').labels)}")
|
|
||||||
|
|
||||||
ud = load_metrics(args.ud_metrics)
|
|
||||||
ner = load_metrics(args.ner_metrics)
|
|
||||||
performance = dict(nlp.meta.get("performance", {}))
|
|
||||||
for key in ("token_acc", "token_p", "token_r", "token_f", "tag_acc", "pos_acc",
|
|
||||||
"morph_acc", "lemma_acc", "dep_uas", "dep_las", "sents_p", "sents_r",
|
|
||||||
"sents_f"):
|
|
||||||
if key in ud:
|
|
||||||
performance[key] = ud[key]
|
|
||||||
for key in ("ents_p", "ents_r", "ents_f"):
|
|
||||||
if key in ner:
|
|
||||||
performance[key] = ner[key]
|
|
||||||
if "ents_per_type" in ner:
|
|
||||||
performance["ents_per_type"] = ner["ents_per_type"]
|
|
||||||
if "dep_las_per_type" in ud:
|
|
||||||
performance["dep_las_per_type"] = ud["dep_las_per_type"]
|
|
||||||
|
|
||||||
nlp.meta.update(
|
|
||||||
{
|
|
||||||
"lang": "fa",
|
|
||||||
"name": "core_news_sm",
|
|
||||||
"version": args.version,
|
|
||||||
"description": DESCRIPTION,
|
|
||||||
"author": "",
|
|
||||||
"email": "",
|
|
||||||
"url": "",
|
|
||||||
"license": "CC BY-SA 4.0",
|
|
||||||
"sources": SOURCES,
|
|
||||||
"notes": NOTES,
|
|
||||||
"performance": performance,
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
out = Path(args.output)
|
|
||||||
nlp.to_disk(out)
|
|
||||||
print(f"wrote {out}")
|
|
||||||
print(json.dumps(performance, indent=2, ensure_ascii=False)[:1200])
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
|
|
@ -0,0 +1,159 @@
|
||||||
|
"""Write complete, convention-compliant metadata onto a trained pipeline.
|
||||||
|
|
||||||
|
Two variants, following spaCy's `[lang]_[type]_[genre]_[size]` naming
|
||||||
|
(https://spacy.io/models#conventions):
|
||||||
|
|
||||||
|
dep -> fa_dep_news_sm tagger + morphologizer + trainable_lemmatizer + parser
|
||||||
|
ent -> fa_ent_news_sm ner only
|
||||||
|
|
||||||
|
`core` is deliberately NOT produced. `core` means "tagger + parser + lemmatizer + NER" in
|
||||||
|
one package, and shipping one would imply the NER is of the same standard as the rest of
|
||||||
|
the pipeline. It is not: the UD components score 85-98 on edited prose while the ner
|
||||||
|
component manages 67.22 F and is trained on tweets, because no redistributably-licensed
|
||||||
|
Persian NER corpus shares a genre with the treebank. Merging them into a single `core`
|
||||||
|
artifact would hide that behind one package name.
|
||||||
|
|
||||||
|
`fa_core_news_sm` is reserved for when ../ner_dataset delivers prose-genre NER data that
|
||||||
|
survives its own ablation (see ../ner_dataset/PLAN.md §5).
|
||||||
|
|
||||||
|
Run twice: once after training (no metrics yet), then again after
|
||||||
|
`spacy benchmark accuracy` so meta.json["performance"] carries held-out test scores.
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
.venv/bin/python scripts/finalize_pipeline.py training/core/model-best training/fa_dep_news_sm \\
|
||||||
|
--variant dep --version 3.8.0 [--metrics metrics/ud-test.json]
|
||||||
|
"""
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import spacy
|
||||||
|
|
||||||
|
PERDT = {
|
||||||
|
"name": "UD_Persian-PerDT (PerUDT v1.0)",
|
||||||
|
"url": "https://github.com/UniversalDependencies/UD_Persian-PerDT",
|
||||||
|
"author": "Mohammad Sadegh Rasooli, Pegah Safari, Amirsaeid Moloodi, Alireza Nourian",
|
||||||
|
"license": "CC BY-SA 4.0",
|
||||||
|
}
|
||||||
|
PARSTWINER = {
|
||||||
|
"name": "ParsTwiNER",
|
||||||
|
"url": "https://github.com/overfit-ir/parstwiner",
|
||||||
|
"author": "MohammadMahdi Aghajani, AliAkbar Badri, Hamid Beigy et al. (Overfit-IR)",
|
||||||
|
"license": "MIT",
|
||||||
|
}
|
||||||
|
LANG_DATA = {
|
||||||
|
"name": "spaCy lang/fa language data (stop words originally from HAZM)",
|
||||||
|
"url": "https://github.com/explosion/spaCy/tree/master/spacy/lang/fa",
|
||||||
|
"author": "Explosion and spaCy contributors",
|
||||||
|
"license": "MIT",
|
||||||
|
}
|
||||||
|
|
||||||
|
VARIANTS = {
|
||||||
|
"dep": {
|
||||||
|
"name": "dep_news_sm",
|
||||||
|
"description": (
|
||||||
|
"Persian dependency pipeline optimized for CPU. Components: tok2vec, tagger, "
|
||||||
|
"morphologizer, trainable_lemmatizer, parser. No NER — see fa_ent_news_sm."
|
||||||
|
),
|
||||||
|
# CC BY-SA 4.0 on the treebank propagates to anything derived from it.
|
||||||
|
"license": "CC BY-SA 4.0",
|
||||||
|
"sources": [PERDT, LANG_DATA],
|
||||||
|
"notes": (
|
||||||
|
"Trained on UD_Persian-PerDT, licensed CC BY-SA 4.0; this pipeline is therefore "
|
||||||
|
"distributed under CC BY-SA 4.0 with attribution to the treebank authors. "
|
||||||
|
"Multiword tokens (pronominal clitics, enclitic copulas) were merged with "
|
||||||
|
"`spacy convert --merge-subtokens`, so a small number of XPOS tags are composite "
|
||||||
|
"(e.g. N_IANM_PR_JOPER) and ~1.5% of lemmas contain a space. "
|
||||||
|
"doc.noun_chunks under-fires on this pipeline: spacy/lang/fa/syntax_iterators.py "
|
||||||
|
"matches ClearNLP labels that do not exist in Universal Dependencies — see "
|
||||||
|
"docs/upstream/fa-noun-chunks.md."
|
||||||
|
),
|
||||||
|
"keys": ("token_acc", "token_p", "token_r", "token_f", "tag_acc", "pos_acc",
|
||||||
|
"morph_acc", "lemma_acc", "dep_uas", "dep_las", "sents_p", "sents_r",
|
||||||
|
"sents_f", "dep_las_per_type"),
|
||||||
|
},
|
||||||
|
"ent": {
|
||||||
|
"name": "ent_news_sm",
|
||||||
|
"description": (
|
||||||
|
"Persian named entity recognizer optimized for CPU, with its own internal "
|
||||||
|
"tok2vec. Labels: PER, ORG, LOC, NAT, POG, EVE."
|
||||||
|
),
|
||||||
|
"license": "MIT",
|
||||||
|
"sources": [PARSTWINER, LANG_DATA],
|
||||||
|
"notes": (
|
||||||
|
"Trained on ParsTwiNER (MIT), a Persian Twitter corpus, because the standard "
|
||||||
|
"Persian NER corpora (ARMAN, PEYMA, NSURL) are research-use-only and cannot be "
|
||||||
|
"redistributed. Expect degraded accuracy on formal or edited prose: this scores "
|
||||||
|
"67.22 F on its own in-genre test set, and EVE (30.0) and POG (41.2) are weak "
|
||||||
|
"enough to be treated as unreliable. The component embeds its own tok2vec rather "
|
||||||
|
"than using a Tok2VecListener, so it can be sourced into another pipeline with "
|
||||||
|
"nlp.add_pipe('ner', source=...). A prose-genre replacement is being built in "
|
||||||
|
"../ner_dataset."
|
||||||
|
),
|
||||||
|
"keys": ("token_acc", "ents_p", "ents_r", "ents_f", "ents_per_type"),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("model", help="trained pipeline, e.g. training/core/model-best")
|
||||||
|
ap.add_argument("output", help="destination directory")
|
||||||
|
ap.add_argument("--variant", choices=sorted(VARIANTS), required=True)
|
||||||
|
ap.add_argument("--version", default="3.8.0")
|
||||||
|
ap.add_argument("--metrics", nargs="*", default=[],
|
||||||
|
help="benchmark accuracy JSON files to fold into performance")
|
||||||
|
args = ap.parse_args()
|
||||||
|
|
||||||
|
spec = VARIANTS[args.variant]
|
||||||
|
nlp = spacy.load(args.model)
|
||||||
|
print(f"pipeline: {nlp.pipe_names}")
|
||||||
|
if args.variant == "dep" and "ner" in nlp.pipe_names:
|
||||||
|
raise SystemExit("refusing to publish a 'dep' pipeline that contains an ner component")
|
||||||
|
if args.variant == "ent" and nlp.pipe_names != ["ner"]:
|
||||||
|
raise SystemExit(f"expected exactly ['ner'], got {nlp.pipe_names}")
|
||||||
|
|
||||||
|
# Build from a strict whitelist rather than inheriting nlp.meta["performance"], which
|
||||||
|
# after training also carries raw *_loss values and scorer keys that do not apply to
|
||||||
|
# this pipeline (e.g. tag_micro_f: 0.0). Published metadata should be the held-out
|
||||||
|
# scores and nothing else.
|
||||||
|
trained = nlp.meta.get("performance", {})
|
||||||
|
performance = {k: trained[k] for k in spec["keys"] if trained.get(k) is not None}
|
||||||
|
for path in args.metrics:
|
||||||
|
p = Path(path)
|
||||||
|
if not p.exists():
|
||||||
|
print(f" (no metrics at {p}, skipping — run `evaluate` first)")
|
||||||
|
continue
|
||||||
|
scored = json.loads(p.read_text(encoding="utf8"))
|
||||||
|
for key in spec["keys"]:
|
||||||
|
if scored.get(key) is not None:
|
||||||
|
performance[key] = scored[key]
|
||||||
|
if scored.get("speed") is not None:
|
||||||
|
performance["speed"] = scored["speed"]
|
||||||
|
|
||||||
|
nlp.meta.update(
|
||||||
|
{
|
||||||
|
"lang": "fa",
|
||||||
|
"name": spec["name"],
|
||||||
|
"version": args.version,
|
||||||
|
"description": spec["description"],
|
||||||
|
"author": "",
|
||||||
|
"email": "",
|
||||||
|
"url": "",
|
||||||
|
"license": spec["license"],
|
||||||
|
"sources": spec["sources"],
|
||||||
|
"notes": spec["notes"],
|
||||||
|
"performance": performance,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
out = Path(args.output)
|
||||||
|
nlp.to_disk(out)
|
||||||
|
print(f"wrote {out} as fa_{spec['name']} {args.version} ({spec['license']})")
|
||||||
|
scalars = {k: round(v * 100, 2) for k, v in performance.items() if isinstance(v, float)}
|
||||||
|
print(json.dumps(scalars, indent=2))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
|
|
@ -1,9 +1,10 @@
|
||||||
"""Run the assembled pipeline over real Persian text and print every annotation layer.
|
"""Run a trained pipeline over real Persian text and print every annotation layer.
|
||||||
|
|
||||||
This is the end-to-end check that the artifact actually works: tokenizer -> tagger ->
|
This is the end-to-end check that the artifact actually works: tokenizer -> tagger ->
|
||||||
morphologizer -> lemmatizer -> parser -> ner -> noun_chunks.
|
morphologizer -> lemmatizer -> parser -> noun_chunks, plus ner when the pipeline has one.
|
||||||
|
Works on fa_dep_news_sm, fa_ent_news_sm, or the two combined.
|
||||||
|
|
||||||
Usage: .venv/bin/python scripts/smoke_test.py training/fa_core_news_sm
|
Usage: .venv/bin/python scripts/smoke_test.py training/fa_dep_news_sm
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import sys
|
import sys
|
||||||
|
|
@ -20,7 +21,7 @@ EXTRA = [
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
path = sys.argv[1] if len(sys.argv) > 1 else "training/fa_core_news_sm"
|
path = sys.argv[1] if len(sys.argv) > 1 else "training/fa_dep_news_sm"
|
||||||
nlp = spacy.load(path)
|
nlp = spacy.load(path)
|
||||||
print(f"loaded {nlp.meta['lang']}_{nlp.meta['name']} {nlp.meta['version']}")
|
print(f"loaded {nlp.meta['lang']}_{nlp.meta['name']} {nlp.meta['version']}")
|
||||||
print(f"pipeline: {nlp.pipe_names}")
|
print(f"pipeline: {nlp.pipe_names}")
|
||||||
|
|
@ -35,10 +36,13 @@ def main():
|
||||||
f"{t.text:<16}{t.lemma_:<16}{t.pos_:<7}{t.tag_:<14}"
|
f"{t.text:<16}{t.lemma_:<16}{t.pos_:<7}{t.tag_:<14}"
|
||||||
f"{t.dep_:<14}{t.head.text}"
|
f"{t.dep_:<14}{t.head.text}"
|
||||||
)
|
)
|
||||||
|
if doc.has_annotation("MORPH"):
|
||||||
print(f"morph[0]: {doc[0].morph}")
|
print(f"morph[0]: {doc[0].morph}")
|
||||||
|
if doc.has_annotation("DEP"):
|
||||||
print(f"sents: {[s.text for s in doc.sents]}")
|
print(f"sents: {[s.text for s in doc.sents]}")
|
||||||
print(f"ents: {[(e.text, e.label_) for e in doc.ents]}")
|
|
||||||
print(f"noun_chunks: {[c.text for c in doc.noun_chunks]}")
|
print(f"noun_chunks: {[c.text for c in doc.noun_chunks]}")
|
||||||
|
if "ner" in nlp.pipe_names:
|
||||||
|
print(f"ents: {[(e.text, e.label_) for e in doc.ents]}")
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue