Compare commits
18 Commits
colab-lg-t
...
main
| Author | SHA1 | Date |
|---|---|---|
|
|
0626a12352 | |
|
|
37c85d0186 | |
|
|
e37d052435 | |
|
|
c80bcb27f3 | |
|
|
b26b194d9f | |
|
|
f1a95d92c3 | |
|
|
e7871a29a1 | |
|
|
4135142e72 | |
|
|
6fa8aa70b8 | |
|
|
5da9dd1524 | |
|
|
f45d0db643 | |
|
|
9e8ed06361 | |
|
|
b89b01ceb6 | |
|
|
c3cb02d9c3 | |
|
|
d39c09daca | |
|
|
8e42c38ed6 | |
|
|
fe6c81272f | |
|
|
60516f6d35 |
|
|
@ -9,6 +9,10 @@ packages/
|
|||
# Separate env for `spacy huggingface-hub push`: it caps typer<0.8, which breaks the
|
||||
# spaCy CLI in the training venv. See .omp/AGENTS.md.
|
||||
.venv-publish/
|
||||
# Local envs for verifying and benchmarking the trf wheel: CPU-only torch, and a cu126 build
|
||||
# for the 940MX. Kept out of .venv so a CUDA-lib downgrade cannot reach the training env.
|
||||
.venv-trf/
|
||||
.venv-trf-gpu/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
# Personal scratch list, not part of the project
|
||||
|
|
|
|||
74
README.fa.md
74
README.fa.md
|
|
@ -12,7 +12,7 @@
|
|||
**بازشناسی موجودیتهای نامدار** را دارد. هر دو تحت لیسانس CC BY-SA ۴٫۰ منتشر شدهاند.
|
||||
|
||||
```bash
|
||||
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl
|
||||
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl
|
||||
```
|
||||
|
||||
```python
|
||||
|
|
@ -22,22 +22,51 @@ doc = nlp("محمدرضا شجریان در مشهد به دنیا آمد.")
|
|||
print(doc.ents) # (محمدرضا شجریان, مشهد)
|
||||
```
|
||||
|
||||
بستههای منتشرشده روی Hugging Face:
|
||||
[`fa_core_news_sm`](https://huggingface.co/Phazel/fa_core_news_sm) ·
|
||||
[`fa_dep_news_sm`](https://huggingface.co/Phazel/fa_dep_news_sm) ·
|
||||
[`fa_ent_news_sm`](https://huggingface.co/Phazel/fa_ent_news_sm) ·
|
||||
[`fa_core_news_md`](https://huggingface.co/Phazel/fa_core_news_md) ·
|
||||
[`fa_dep_news_md`](https://huggingface.co/Phazel/fa_dep_news_md) ·
|
||||
[`fa_ent_news_md`](https://huggingface.co/Phazel/fa_ent_news_md) ·
|
||||
[`fa_core_news_lg`](https://huggingface.co/Phazel/fa_core_news_lg) ·
|
||||
[`fa_dep_news_lg`](https://huggingface.co/Phazel/fa_dep_news_lg) ·
|
||||
[`fa_ent_news_lg`](https://huggingface.co/Phazel/fa_ent_news_lg) ·
|
||||
[`fa_core_news_trf`](https://huggingface.co/Phazel/fa_core_news_trf).
|
||||
جدولهای بردار floret جداگانه (فقط بردار، بدون هیچ مؤلفهای):
|
||||
|
||||
```bash
|
||||
# ۵۰ هزار سطر × ۳۰۰ بعد، ۴۰۰ هزار سند فارسی (جدول ردهٔ md)
|
||||
pip install https://huggingface.co/Phazel/fa_floret_400k/resolve/main/fa_floret_400k-0.1.0-py3-none-any.whl
|
||||
# ۵۰ هزار سطر × ۳۰۰ بعد، کل دامپ ویکیپدیای فارسی
|
||||
pip install https://huggingface.co/Phazel/fa_floret_full_wiki/resolve/main/fa_floret_full_wiki-0.1.0-py3-none-any.whl
|
||||
# ۲۰۰ هزار سطر × ۳۰۰ بعد، کل دامپ ویکیپدیای فارسی، ۵ دوره (جدول ردهٔ lg)
|
||||
pip install https://huggingface.co/Phazel/fa-floret-wiki-vectors/resolve/main/fa_floret_wiki_200k-0.1.0-py3-none-any.whl
|
||||
```
|
||||
|
||||
## کارایی
|
||||
|
||||
ارزیابی با `spacy benchmark accuracy` روی بخش آزمون همان پیکره انجام شده است:
|
||||
|
||||
| سنجه | امتیاز | مرجع |
|
||||
| --- | --- | --- |
|
||||
| `TOKEN_ACC` / `TOKEN_F` | ۹۹٫۹۶ / ۹۹٫۱۱ | |
|
||||
| `TAG_ACC` (XPOS) | ۹۵٫۹۶ | |
|
||||
| `POS_ACC` (UPOS) | ۹۶٫۲۴ | |
|
||||
| `MORPH_ACC` | ۹۶٫۲۹ | |
|
||||
| `LEMMA_ACC` | ۹۷٫۹۱ | |
|
||||
| `SENTS_F` | ۹۹٫۲۵ | |
|
||||
| `DEP_UAS` | ۸۹٫۶۹ | hazm+ParsBERT: ۹۲٫۴۶ |
|
||||
| `DEP_LAS` | ۸۵٫۱۵ | hazm+ParsBERT: ۸۹٫۳۴ |
|
||||
| `ENTS_F` | ۷۱٫۸۷ | تنها در `fa_core_news_sm` |
|
||||
| سرعت | حدود ۹٬۲۵۰ واژه بر ثانیه | |
|
||||
| سنجه | `sm` | `md` | `lg` | `trf` | مرجع |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `TOKEN_ACC` / `TOKEN_F` | ۹۹٫۹۶ / ۹۹٫۱۱ | ۹۹٫۹۶ / ۹۹٫۱۱ | ۹۹٫۹۶ / ۹۹٫۱۱ | ۹۹٫۹۶ / ۹۹٫۱۱ | |
|
||||
| `TAG_ACC` (XPOS) | ۹۵٫۹۶ | ۹۶٫۲۵ | ۹۶٫۵۵ | **۹۷٫۶۲** | |
|
||||
| `POS_ACC` (UPOS) | ۹۶٫۲۴ | ۹۶٫۶۴ | ۹۶٫۶۸ | **۹۷٫۶۳** | |
|
||||
| `MORPH_ACC` | ۹۶٫۲۹ | ۹۶٫۶۴ | ۹۶٫۷۰ | **۹۷٫۸۲** | |
|
||||
| `LEMMA_ACC` | ۹۷٫۹۱ | ۹۷٫۹۶ | **۹۸٫۰۸** | ۹۷٫۳۱ | |
|
||||
| `SENTS_F` | ۹۹٫۲۵ | **۹۹٫۲۸** | ۹۹٫۱۸ | ۹۷٫۳۵ | |
|
||||
| `DEP_UAS` | ۸۹٫۶۹ | ۹۰٫۵۲ | ۹۰٫۹۶ | **۹۳٫۸۷** | hazm+ParsBERT: ۹۲٫۴۶ |
|
||||
| `DEP_LAS` | ۸۵٫۱۵ | ۸۶٫۳۴ | ۸۶٫۶۰ | **۹۰٫۷۹** | hazm+ParsBERT: ۸۹٫۳۴ |
|
||||
| `ENTS_P` | ۷۷٫۶۷ | ۷۶٫۵۶ | ۸۱٫۵۱ | **۸۴٫۰۶** | |
|
||||
| `ENTS_R` | ۶۶٫۸۷ | ۷۲٫۹۵ | ۷۱٫۰۹ | **۸۱٫۷۶** | |
|
||||
| `ENTS_F` | ۷۱٫۸۷ | ۷۴٫۷۱ | ۷۵٫۹۴ | **۸۲٫۸۹** | |
|
||||
| سرعت (940MX، دستهٔ ۳۲) | ۱۰٬۲۳۵ | ۹٬۰۵۸ | ۹٬۲۱۵ | بخش توان عملیاتی | |
|
||||
| حجم بستهٔ نصب | ۱۳٫۵ مگابایت | ۶۸٫۵ مگابایت | ۲۳۵ مگابایت | ۶۰۸ مگابایت | |
|
||||
|
||||
ردهٔ `trf` در همهجا جلو است مگر در واژهیابی و مرزبندی جمله، و تنها ردهٔای است که از مرجع
|
||||
`DEP_LAS` برابر ۸۹٫۳۴ عبور میکند. به کارت گرافیک نیاز دارد و مدل پایهٔ آن پروانهٔ مشخصی ندارد،
|
||||
پس قابل بازانتشار نیست (`docs/MODELS.md` بخش ۸).
|
||||
|
||||
برچسبهای موجودیت «نقرهای» هستند: از لایهای در خود پیکره میآیند که با برچسبزن Beheshti-NER
|
||||
تولید و سپس دستی اصلاح شده است. بنابراین `ENTS_F` تا اندازهای همخوانی با آن برچسبزن را
|
||||
|
|
@ -46,6 +75,25 @@ print(doc.ents) # (محمدرضا شجریان, مشهد)
|
|||
آموزش روی یک پردازندهٔ چهارهستهای i5-7200U و بدون کارت گرافیک انجام شده است: ۱ ساعت و ۲۷ دقیقه
|
||||
برای اجزای نحوی و ۱۷ دقیقه برای NER. این دو اجرا مستقلاند و میتوانند همزمان انجام شوند.
|
||||
|
||||
## توان عملیاتی
|
||||
|
||||
میانهٔ چند اجرای پیاپی `nlp.pipe` روی ۱۴۶ سند بخش آزمون PerDT (۲۳٬۸۲۵ توکن). تنها زمان خودِ
|
||||
`pipe` اندازهگیری شده و اجرای گرمکردن کنار گذاشته میشود. برای بازتولید:
|
||||
`python scripts/benchmark_throughput.py <model> --gpu-id <n>`؛ دادهٔ خام در
|
||||
`metrics/throughput-*.json` است.
|
||||
|
||||
| رده | پردازنده i5-7200U | کارت 940MX | کارت Tesla T4 |
|
||||
| --- | ---: | ---: | ---: |
|
||||
| `sm` | ۵٬۴۸۴ | ۱۰٬۲۳۵ | |
|
||||
| `md` | ۵٬۴۰۸ | ۹٬۰۵۸ | |
|
||||
| `lg` | ۴٬۷۱۵ | ۹٬۲۱۵ | |
|
||||
| `trf` | ۱۸۷ | ۱٬۱۰۶ | ۸٬۳۲۰ |
|
||||
|
||||
ردهٔ `trf` روی یک پردازنده ۲۹ برابر کندتر از `sm` است. عددهای T4 و Xeon از یک ماشین Colab
|
||||
میآیند، یعنی شتاب ۲۵ برابری. فاصلهٔ ردههای پردازندهای کمتر از ۱۵ درصد است، پس گلوگاه
|
||||
تجزیهگر و واژهیاب است نه جستوجوی tok2vec. پراکندگی اجراها روی لپتاپ حدود ۱۰± درصد است.
|
||||
اجرای `trf` روی 940MX به نسخهٔ مشخصی از torch نیاز دارد؛ بخش ۹ از `docs/MODELS.md` را ببینید.
|
||||
|
||||
گامهای تبدیل پیکره، آموزش، ارزیابی و بستهبندی در [`project.yml`](project.yml) تعریف شدهاند.
|
||||
توضیح بیشتر دربارهٔ گزینش پیکره و پروانهها در [`docs/MODELS.md`](docs/MODELS.md) و شرح انگلیسی
|
||||
پروژه در [`README.md`](README.md) آمده است.
|
||||
|
|
|
|||
188
README.md
188
README.md
|
|
@ -1,9 +1,9 @@
|
|||
# Persian (Farsi) pipelines for spaCy
|
||||
|
||||
Trained spaCy pipelines for Persian, installable now. spaCy has never shipped an official one, and `spacy.blank("fa")` only gives you a tokenizer and stop words. choose between `fa_core_news_sm` (full syntax + NER) or `fa_dep_news_sm` (syntax only).
|
||||
Trained spaCy pipelines for Persian, installable with pip. spaCy has never shipped an official one, and `spacy.blank("fa")` only gives you a tokenizer and stop words. Choose between `fa_core_news_sm` (full syntax + NER) or `fa_dep_news_sm` (syntax only).
|
||||
|
||||
```bash
|
||||
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl
|
||||
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl
|
||||
```
|
||||
|
||||
```python
|
||||
|
|
@ -21,84 +21,153 @@ pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_n
|
|||
[('ایران خودرو', 'ORG'), ('۲۰ درصد', 'PCT')]
|
||||
```
|
||||
|
||||
## Why spacy-persian?
|
||||
|
||||
- **⚡ Performance** – **96.24%** POS · **97.91%** Lemma · **85.15%** LAS – competitive with English `en_core_web_sm` on syntax.
|
||||
- **🚀 Speed** – ~9,250 words/sec on a standard CPU. No GPU required.
|
||||
- **📦 Flexibility** – Choose `fa_core_news_sm` (13MB, syntax + NER) or `fa_dep_news_sm` (7.5MB, syntax-only).
|
||||
- **🔁 Reproducibility** – Checksummed, versioned builds from UD_Persian-PerDT – no black boxes.
|
||||
- **🔌 Native spaCy** – Drop-in replacement. `spacy.load()` works instantly with standard `Doc` objects.
|
||||
-
|
||||
|
||||
## Results
|
||||
|
||||
`spacy-persian` delivers production‑ready Persian NLP that stands alongside Hazm—the most popular Persian toolkit—while bringing the full power of the spaCy ecosystem.
|
||||
Compared against Hazm (the most-used Persian toolkit) and `en_core_web_sm` (English reference).
|
||||
|
||||
| Metric | **`spacy-persian`**<br>`fa_core_news_sm` | **Hazm**<br>(Persian toolkit) | `en_core_web_sm`<br>(English reference) |
|
||||
| Metric | **`spacy-persian`**<br>`fa_core_news_trf` | **Hazm**<br>(Persian toolkit) | `en_core_web_sm`<br>(English reference) |
|
||||
|--------|:---:|:---:|:---:|
|
||||
| **POS Accuracy (UPOS)** | **96.24%** | ~95.69%¹ | 97.21%² |
|
||||
| **Lemma Accuracy** | **97.91%** | 89.9%¹ | — |
|
||||
| **Dependency LAS** | 85.15% | 85.6%¹ | 91.85%² |
|
||||
| **NER F-score** | 71.87% | — | 83.80%² |
|
||||
| **Package Size** | **13 MB** (syntax+NER)<br>**7.5 MB** (syntax-only) | ~7 MB | 12 MB |
|
||||
| **POS Accuracy (UPOS)** | **97.63%** | ~95.69%¹ | 97.21%² |
|
||||
| **Lemma Accuracy** | **97.31%** | 89.9%¹ | — |
|
||||
| **Dependency LAS** | **90.79%** | 85.6%¹ | 91.85%² |
|
||||
| **NER F-score** | **82.89%** | — | 83.80%² |
|
||||
|
||||
> **¹** Hazm scores from its official README
|
||||
> **²** `en_core_web_sm` scores from spaCy's official model card
|
||||
|
||||
> ⚠️ **Note on comparability:** These benchmarks come from *different evaluation sets, treebanks, and test splits*.
|
||||
> **Note on comparability:** these benchmarks come from different evaluation sets, treebanks, and test splits.
|
||||
|
||||
### Packages
|
||||
|
||||
| Package | Components | Licence | Score | Wheel |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| [`fa_dep_news_sm`](https://huggingface.co/Phazel/fa_dep_news_sm) | tok2vec, tagger, morphologizer, trainable_lemmatizer, parser | CC BY-SA 4.0 | LEMMA 97.91 | 7.9 MB |
|
||||
| [`fa_core_news_sm`](https://huggingface.co/Phazel/fa_core_news_sm) | the above plus ner | CC BY-SA 4.0 | ENTS_F 71.87 | 13.5 MB |
|
||||
| [`fa_ent_news_sm`](https://huggingface.co/Phazel/fa_ent_news_sm) | `ner` alone (own embedded tok2vec) | CC BY-SA 4.0 | ENTS_F 71.87 | 5.9 MB |
|
||||
| [`fa_dep_news_md`](https://huggingface.co/Phazel/fa_dep_news_md) | same as `fa_dep_news_sm`, plus floret vectors | CC BY-SA 4.0 | LEMMA 97.96 | 62.6 MB |
|
||||
| [`fa_core_news_md`](https://huggingface.co/Phazel/fa_core_news_md) | same as `fa_core_news_sm`, plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 68.5 MB |
|
||||
| [`fa_ent_news_md`](https://huggingface.co/Phazel/fa_ent_news_md) | `ner` alone (own embedded tok2vec), plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 60.6 MB |
|
||||
| [`fa_dep_news_lg`](https://huggingface.co/Phazel/fa_dep_news_lg) | same as `fa_dep_news_sm`, plus full-wiki floret vectors | CC BY-SA 4.0 | LEMMA 98.08 | 229.3 MB |
|
||||
| [`fa_core_news_lg`](https://huggingface.co/Phazel/fa_core_news_lg) | same as `fa_core_news_sm`, plus full-wiki floret vectors | CC BY-SA 4.0 | ENTS_F 75.94 | 235.2 MB |
|
||||
| [`fa_ent_news_lg`](https://huggingface.co/Phazel/fa_ent_news_lg) | `ner` alone (own embedded tok2vec), plus full-wiki floret vectors | CC BY-SA 4.0 | ENTS_F 75.94 | 227.3 MB |
|
||||
| [`fa_core_news_trf`](https://huggingface.co/Phazel/fa_core_news_trf) | transformer, tagger, morphologizer, trainable_lemmatizer, parser, ner | see §8, encoder unlicensed | ENTS_F 82.89, LAS 90.79 | 608.2 MB |
|
||||
|
||||
| Metric | Score | Reference |
|
||||
| --- | --- | --- |
|
||||
| `TOKEN_ACC` / `TOKEN_F` | 99.96 / 99.11 | |
|
||||
| `TAG_ACC` (XPOS) | 95.96 | |
|
||||
| `POS_ACC` (UPOS) | 96.24 | |
|
||||
| `MORPH_ACC` | 96.29 | |
|
||||
| `LEMMA_ACC` | 97.91 | |
|
||||
| `SENTS_F` | 99.25 | |
|
||||
| `DEP_UAS` | 89.69 | hazm+ParsBERT: 92.46 |
|
||||
| `DEP_LAS` | 85.15 | hazm+ParsBERT: 89.34 |
|
||||
| Speed | ~9,250 words/s | |
|
||||
These scores are from `spacy benchmark accuracy`, stored in `metrics/`.
|
||||
|
||||
Entities, `fa_core_news_sm` only, on the PerDT NER test split: `ENTS_P` 77.67, `ENTS_R` 66.87,
|
||||
`ENTS_F` 71.87.
|
||||
Raw `fa.floret` and `fa.vec` exports of the `lg` tier's 200k-row table are in
|
||||
[`fa-floret-wiki-vectors`](https://huggingface.co/Phazel/fa-floret-wiki-vectors).
|
||||
|
||||
| Label | F | Train examples |
|
||||
| --- | --- | --- |
|
||||
| `LOC` | 80.24 | 4,954 |
|
||||
| `DAT` | 74.45 | 1,323 |
|
||||
| `MON` | 73.68 | 205 |
|
||||
| `ORG` | 68.77 | 2,643 |
|
||||
| `TIM` | 66.67 | 135 |
|
||||
| `PER` | 65.29 | 4,847 |
|
||||
| `PCT` | 57.14 | 121 |
|
||||
### Tier comparison
|
||||
|
||||
The `md` tier adds a 50k x 300d floret vector table trained on 400k Persian documents. Its
|
||||
config differs from `sm` by exactly one line (`include_static_vectors`), so the columns below
|
||||
isolate what the vectors buy. Full breakdown in `docs/MODELS.md` §6.
|
||||
|
||||
For comparison, `en_core_web_sm` scores TAG 97, LAS 90, ENTS_F 84 on a larger, cleaner corpus.
|
||||
Trained on a 4-core i5-7200U with no GPU: 1h27m for the syntax components, 17 min for NER.
|
||||
| Metric | `sm` | `md` | `lg` | `trf` | Reference |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `TOKEN_ACC` / `TOKEN_F` | 99.96 / 99.11 | 99.96 / 99.11 | 99.96 / 99.11 | 99.96 / 99.11 | |
|
||||
| `TAG_ACC` (XPOS) | 95.96 | 96.25 | 96.55 | **97.62** | |
|
||||
| `POS_ACC` (UPOS) | 96.24 | 96.64 | 96.68 | **97.63** | |
|
||||
| `MORPH_ACC` | 96.29 | 96.64 | 96.70 | **97.82** | |
|
||||
| `LEMMA_ACC` | 97.91 | 97.96 | **98.08** | 97.31 | |
|
||||
| `SENTS_F` | 99.25 | **99.28** | 99.18 | 97.35 | |
|
||||
| `DEP_UAS` | 89.69 | 90.52 | 90.96 | **93.87** | Hazm+ParsBERT: 92.46 |
|
||||
| `DEP_LAS` | 85.15 | 86.34 | 86.60 | **90.79** | Hazm+ParsBERT: 89.34 |
|
||||
| `ENTS_P` | 77.67 | 76.56 | 81.51 | **84.06** | |
|
||||
| `ENTS_R` | 66.87 | 72.95 | 71.09 | **81.76** | |
|
||||
| `ENTS_F` | 71.87 | 74.71 | 75.94 | **82.89** | |
|
||||
| Speed (940MX, batch 32) | 10,235 words/s | 9,058 words/s | 9,215 words/s | 1,106 words/s | |
|
||||
| Wheel size | 13.5 MB | 68.5 MB | 235 MB | 608 MB | |
|
||||
|
||||
`trf` leads on every metric except lemmatization and sentence segmentation. It is also the only
|
||||
tier to clear the Hazm+ParsBERT `DEP_LAS` reference of 89.34. It needs a GPU, and its encoder
|
||||
states no licence, so it is not redistributable (`docs/MODELS.md` §8).
|
||||
|
||||
Entity scores are `fa_core_news_*` on the PerDT NER test split; per-label breakdown and
|
||||
caveats are in [Named entity recognition](#named-entity-recognition).
|
||||
|
||||
Trained on a 4-core i5-7200U with no GPU: `sm` took 1h27m for syntax plus 17 min for NER,
|
||||
`md` 1h54m plus 25 min (the two `md` runs overlapped, so wall clock overstates each).
|
||||
|
||||
### Vector packages
|
||||
|
||||
Standalone floret vector packages (vectors only, `pipeline: []`), usable as
|
||||
`--paths.vectors` for your own training or as a plain embedding table:
|
||||
|
||||
```bash
|
||||
# 50k rows x 300d, 400k Persian documents (the md tier's table)
|
||||
pip install https://huggingface.co/Phazel/fa_floret_400k/resolve/main/fa_floret_400k-0.1.0-py3-none-any.whl
|
||||
# 50k rows x 300d, full Persian Wikipedia dump
|
||||
pip install https://huggingface.co/Phazel/fa_floret_full_wiki/resolve/main/fa_floret_full_wiki-0.1.0-py3-none-any.whl
|
||||
# 200k rows x 300d, full Persian Wikipedia dump, 5 epochs (the lg tier's table)
|
||||
pip install https://huggingface.co/Phazel/fa-floret-wiki-vectors/resolve/main/fa_floret_wiki_200k-0.1.0-py3-none-any.whl
|
||||
```
|
||||
|
||||
## Throughput
|
||||
|
||||
Median of repeated `nlp.pipe` passes over the 146-document PerDT test split (23,825 tokens),
|
||||
timing the pipe only, warmup discarded. Reproduce with
|
||||
`python scripts/benchmark_throughput.py <model> --gpu-id <n>`; raw records are in
|
||||
`metrics/throughput-*.json`.
|
||||
|
||||
| Tier | CPU, i5-7200U | GPU, GeForce 940MX | CPU, Xeon @ 2.00GHz | GPU, Tesla T4 |
|
||||
| --- | ---: | ---: | ---: | ---: |
|
||||
| `sm` | 5,484 | 10,235 | | |
|
||||
| `md` | 5,408 | 9,058 | | |
|
||||
| `lg` | 4,715 | 9,215 | | |
|
||||
| `trf` | 187 | 1,106 | 336 | 8,320 |
|
||||
|
||||
`trf` is 29x slower than `sm` on the same CPU. The Xeon and T4 columns come from one Colab VM,
|
||||
a 25x GPU speedup. The CPU tiers sit within 15% of each other, so the bottleneck is the parser
|
||||
and lemmatizer, not the tok2vec lookup. Laptop spread is about 10% with thermal state. Running
|
||||
`trf` on the 940MX needs a `cu126` torch build, see `docs/MODELS.md` §9.
|
||||
|
||||
## Named entity recognition
|
||||
|
||||
Seven labels: `LOC`, `PER`, `ORG`, `DAT`, `MON`, `TIM`, `PCT`. They come from PerDT's own
|
||||
`not-to-release/Dadegan with NER tag/` layer, transferred onto this pipeline's tokenization
|
||||
by difflib at a 99.86% alignment rate (`scripts/transfer_perdt_ner.py`). Spans that could not
|
||||
be aligned exactly were dropped rather than guessed. That layer is silver: PerDT's README
|
||||
states it was produced by the BERT-based Beheshti-NER tagger with manual corrections for
|
||||
recall, so the `ENTS_F` numbers below partly reflect agreement with that tagger, not with
|
||||
human annotation.
|
||||
|
||||
`ner` runs standalone with its own embedded tok2vec (`fa_ent_news_sm`, `fa_ent_news_md`), or
|
||||
bundled into `fa_core_news_sm`/`fa_core_news_md` alongside the syntax pipeline. In `trf` it is
|
||||
trained jointly against the shared transformer instead, so there is no standalone trf variant.
|
||||
|
||||
| Label | `sm` F | `md` F | `lg` F | `trf` F | Train examples |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `LOC` | 80.24 | 84.05 | 83.66 | **87.78** | 4,954 |
|
||||
| `PER` | 65.29 | 68.18 | 72.63 | **81.88** | 4,847 |
|
||||
| `ORG` | 68.77 | 70.25 | 71.01 | **78.50** | 2,643 |
|
||||
| `DAT` | 74.45 | 76.19 | 70.83 | **82.52** | 1,323 |
|
||||
| `MON` | 73.68 | 84.21 | 88.89 | 88.89 | 205 |
|
||||
| `TIM` | 66.67 | 66.67 | 61.54 | 50.00 | 135 |
|
||||
| `PCT` | 57.14 | 33.33 | 57.14 | 33.33 | 121 |
|
||||
|
||||
`MON`, `TIM` and `PCT` have single-digit support in the test split, so their deltas are one or
|
||||
two entities changing hands, not signal. `PER`, `LOC` and `ORG` carry the split. The `md` gain
|
||||
over `sm` (`ENTS_F` 71.87 to 74.71) is almost entirely recall (+6.08), the lexical prior static
|
||||
vectors give rare proper nouns that hash embeddings never had. `trf` adds another +6.95 F over
|
||||
`lg`, again mostly recall (71.09 to 81.76), and its largest per-label gains are `PER` (+9.25)
|
||||
and `DAT` (+11.69).
|
||||
|
||||
## Install
|
||||
|
||||
```bash
|
||||
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-any-py3-none-any.whl
|
||||
pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_news_sm-3.8.0-py3-none-any.whl
|
||||
# or, without NER:
|
||||
pip install https://huggingface.co/Phazel/fa_dep_news_sm/resolve/main/fa_dep_news_sm-any-py3-none-any.whl
|
||||
pip install https://huggingface.co/Phazel/fa_dep_news_sm/resolve/main/fa_dep_news_sm-3.8.0-py3-none-any.whl
|
||||
```
|
||||
|
||||
## Caveats
|
||||
|
||||
- **The entity labels are silver.** They come from the treebank's own
|
||||
`not-to-release/Dadegan with NER tag/` layer, which its README states was produced by the
|
||||
BERT-based Beheshti-NER tagger with manual corrections for recall. `ENTS_F 71.87` is measured
|
||||
against a silver test split and partly reflects agreement with that tagger.
|
||||
- **Three entity labels are thin.** `MON` (205 training examples), `TIM` (135) and `PCT` (121)
|
||||
rest on 4 to 11 test entities each. `PER`, `LOC`, `ORG` and `DAT` have 1,300 or more.
|
||||
- **Some lemmas contain a space.** Multiword tokens were merged, so `کتابهایش` is one token
|
||||
tagged `N_IANM_PR_JOPER` with lemma `کتاب او`. This affects about 1.5% of tokens.
|
||||
- **`doc.noun_chunks` under-fires.** `spacy/lang/fa/syntax_iterators.py` upstream matches
|
||||
ClearNLP labels that do not exist in Universal Dependencies. Patch in
|
||||
[`docs/upstream/fa-noun-chunks.md`](docs/upstream/fa-noun-chunks.md).
|
||||
ClearNLP labels that do not exist in Universal Dependencies. Bug analysis and proposed
|
||||
upstream patch in [`docs/upstream/fa-noun-chunks.md`](docs/upstream/fa-noun-chunks.md).
|
||||
|
||||
## Build
|
||||
|
||||
|
|
@ -147,14 +216,11 @@ The two training runs are single-threaded and independent, so they can run concu
|
|||
English pipelines derive UPOS from PTB tags by rule because OntoNotes has no UPOS. UD gives
|
||||
gold UPOS, FEATS and lemmas, which yields real `pos_acc`, `morph_acc` and `lemma_acc` numbers
|
||||
instead of unmeasurable rule coverage.
|
||||
4. PerDT, not Seraji: 3.7x more tokens, and Seraji has no `PROPN` tag. Entity spans were
|
||||
transferred onto this pipeline's tokenization by difflib at a 99.86% rate, and spans that
|
||||
could not be aligned exactly were dropped rather than guessed
|
||||
(`scripts/transfer_perdt_ner.py`).
|
||||
4. PerDT, not Seraji: 3.7x more tokens, and Seraji has no `PROPN` tag.
|
||||
|
||||
## Why not hazm's own models
|
||||
## Why not Hazm's own models
|
||||
|
||||
hazm is the reference Persian NLP toolkit and publishes spaCy-format pipelines on the HF Hub,
|
||||
Hazm is the reference Persian NLP toolkit and publishes spaCy-format pipelines on the HF Hub,
|
||||
so it was the obvious starting point. Four problems:
|
||||
|
||||
- Its trainable models are pycrfsuite CRFs (`hazm/sequence_tagger.py`). The repo contains no
|
||||
|
|
@ -167,7 +233,7 @@ so it was the obvious starting point. Four problems:
|
|||
- Most corpora it reads (Bijankhan, Peykare, Hamshahri, raw PerDT) sit behind `peykaregan.ir`
|
||||
or `dadegan.ir` under research-only terms.
|
||||
|
||||
It did confirm the corpus choice. hazm's own spaCy parser was trained on
|
||||
It did confirm the corpus choice. Hazm's own spaCy parser was trained on
|
||||
`modified_fa_perdt-ud-train.spacy`, the same treebank used here.
|
||||
|
||||
## More
|
||||
|
|
@ -176,5 +242,5 @@ It did confirm the corpus choice. hazm's own spaCy parser was trained on
|
|||
- How spaCy models get published, and what upstream `fa` already has:
|
||||
[`docs/CONTRIBUTING-GUIDE.md`](docs/CONTRIBUTING-GUIDE.md)
|
||||
- The build: [`project.yml`](project.yml)
|
||||
- Language data comes from `spacy/lang/fa` upstream, whose stop word list came from Hazm.
|
||||
- خلاصهٔ فارسی: [`README.fa.md`](README.fa.md)
|
||||
- Language data comes from `spacy/lang/fa` upstream, whose stop word list came from hazm.
|
||||
|
|
|
|||
|
|
@ -0,0 +1,291 @@
|
|||
# fa_core_news_trf: the whole pipeline on one fine-tuned ParsBERT encoder.
|
||||
#
|
||||
# Differences from the sm/md/lg tiers, all forced by the transformer:
|
||||
#
|
||||
# * One corpus, not two. sm/md/lg train `ner` separately (own embedded tok2vec) and source
|
||||
# it into the dep model. Fine-tuning a 162M-parameter encoder twice would double GPU cost
|
||||
# and ship two encoders in one wheel, and the second would collide on the `transformer`
|
||||
# component name. So every component listens to a single shared transformer and trains
|
||||
# against corpus/joint/, built by scripts/merge_joint_corpus.py (UD layer + the
|
||||
# difflib-transferred NER layer on identical tokenization).
|
||||
# * `use_upper = false` on both transition-based parsers: with a transformer upstream the
|
||||
# extra maxout layer is redundant, and this matches the upstream *_trf configs.
|
||||
# * Adam + warmup_linear and accumulate_gradient=3, not the flat 0.001 the CPU tiers use.
|
||||
# Fine-tuning a pretrained encoder at 1e-3 diverges.
|
||||
# * gpu_allocator = "pytorch" so thinc and torch share one CUDA memory pool.
|
||||
#
|
||||
# Encoder: HooshvareLab/bert-base-parsbert-uncased. NOTE the licence caveat in
|
||||
# docs/MODELS.md §3.4 - ParsBERT's model card carries no licence statement, so this wheel is
|
||||
# NOT redistributable on those grounds; HooshvareLab/roberta-fa-zwnj-base (Apache-2.0) is the
|
||||
# publishable alternative and drops in by changing `name` below.
|
||||
|
||||
[paths]
|
||||
train = null
|
||||
dev = null
|
||||
vectors = null
|
||||
init_tok2vec = null
|
||||
|
||||
[system]
|
||||
gpu_allocator = "pytorch"
|
||||
seed = 0
|
||||
|
||||
[nlp]
|
||||
lang = "fa"
|
||||
pipeline = ["transformer","tagger","morphologizer","trainable_lemmatizer","parser","ner"]
|
||||
batch_size = 128
|
||||
disabled = []
|
||||
before_creation = null
|
||||
after_creation = null
|
||||
after_pipeline_creation = null
|
||||
|
||||
[nlp.tokenizer]
|
||||
@tokenizers = "spacy.Tokenizer.v1"
|
||||
|
||||
[nlp.vectors]
|
||||
@vectors = "spacy.Vectors.v1"
|
||||
|
||||
[components]
|
||||
|
||||
[components.transformer]
|
||||
factory = "transformer"
|
||||
max_batch_items = 4096
|
||||
|
||||
[components.transformer.set_extra_annotations]
|
||||
@annotation_setters = "spacy-transformers.null_annotation_setter.v1"
|
||||
|
||||
[components.transformer.model]
|
||||
@architectures = "spacy-transformers.TransformerModel.v3"
|
||||
name = "HooshvareLab/bert-base-parsbert-uncased"
|
||||
mixed_precision = false
|
||||
|
||||
[components.transformer.model.get_spans]
|
||||
@span_getters = "spacy-transformers.strided_spans.v1"
|
||||
window = 128
|
||||
stride = 96
|
||||
|
||||
[components.transformer.model.tokenizer_config]
|
||||
use_fast = true
|
||||
|
||||
[components.transformer.model.transformer_config]
|
||||
|
||||
[components.transformer.model.grad_scaler_config]
|
||||
|
||||
[components.tagger]
|
||||
factory = "tagger"
|
||||
label_smoothing = 0.05
|
||||
overwrite = false
|
||||
neg_prefix = "!"
|
||||
|
||||
[components.tagger.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.tagger.model.tok2vec]
|
||||
@architectures = "spacy-transformers.TransformerListener.v1"
|
||||
grad_factor = 1.0
|
||||
upstream = "*"
|
||||
|
||||
[components.tagger.model.tok2vec.pooling]
|
||||
@layers = "reduce_mean.v1"
|
||||
|
||||
[components.tagger.scorer]
|
||||
@scorers = "spacy.tagger_scorer.v1"
|
||||
|
||||
[components.morphologizer]
|
||||
factory = "morphologizer"
|
||||
label_smoothing = 0.05
|
||||
overwrite = true
|
||||
extend = false
|
||||
|
||||
[components.morphologizer.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.morphologizer.model.tok2vec]
|
||||
@architectures = "spacy-transformers.TransformerListener.v1"
|
||||
grad_factor = 1.0
|
||||
upstream = "*"
|
||||
|
||||
[components.morphologizer.model.tok2vec.pooling]
|
||||
@layers = "reduce_mean.v1"
|
||||
|
||||
[components.morphologizer.scorer]
|
||||
@scorers = "spacy.morphologizer_scorer.v1"
|
||||
|
||||
[components.trainable_lemmatizer]
|
||||
factory = "trainable_lemmatizer"
|
||||
backoff = "orth"
|
||||
min_tree_freq = 3
|
||||
overwrite = false
|
||||
top_k = 1
|
||||
|
||||
[components.trainable_lemmatizer.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.trainable_lemmatizer.model.tok2vec]
|
||||
@architectures = "spacy-transformers.TransformerListener.v1"
|
||||
grad_factor = 1.0
|
||||
upstream = "*"
|
||||
|
||||
[components.trainable_lemmatizer.model.tok2vec.pooling]
|
||||
@layers = "reduce_mean.v1"
|
||||
|
||||
[components.trainable_lemmatizer.scorer]
|
||||
@scorers = "spacy.lemmatizer_scorer.v1"
|
||||
|
||||
[components.parser]
|
||||
factory = "parser"
|
||||
moves = null
|
||||
update_with_oracle_cut_size = 100
|
||||
learn_tokens = false
|
||||
min_action_freq = 30
|
||||
|
||||
[components.parser.model]
|
||||
@architectures = "spacy.TransitionBasedParser.v2"
|
||||
state_type = "parser"
|
||||
extra_state_tokens = false
|
||||
hidden_width = 64
|
||||
maxout_pieces = 2
|
||||
use_upper = false
|
||||
nO = null
|
||||
|
||||
[components.parser.model.tok2vec]
|
||||
@architectures = "spacy-transformers.TransformerListener.v1"
|
||||
grad_factor = 1.0
|
||||
upstream = "*"
|
||||
|
||||
[components.parser.model.tok2vec.pooling]
|
||||
@layers = "reduce_mean.v1"
|
||||
|
||||
[components.parser.scorer]
|
||||
@scorers = "spacy.parser_scorer.v1"
|
||||
|
||||
[components.ner]
|
||||
factory = "ner"
|
||||
moves = null
|
||||
update_with_oracle_cut_size = 100
|
||||
incorrect_spans_key = null
|
||||
|
||||
[components.ner.model]
|
||||
@architectures = "spacy.TransitionBasedParser.v2"
|
||||
state_type = "ner"
|
||||
extra_state_tokens = false
|
||||
hidden_width = 64
|
||||
maxout_pieces = 2
|
||||
use_upper = false
|
||||
nO = null
|
||||
|
||||
[components.ner.model.tok2vec]
|
||||
@architectures = "spacy-transformers.TransformerListener.v1"
|
||||
grad_factor = 1.0
|
||||
upstream = "*"
|
||||
|
||||
[components.ner.model.tok2vec.pooling]
|
||||
@layers = "reduce_mean.v1"
|
||||
|
||||
[components.ner.scorer]
|
||||
@scorers = "spacy.ner_scorer.v1"
|
||||
|
||||
[corpora]
|
||||
|
||||
[corpora.train]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.train}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[corpora.dev]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.dev}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[training]
|
||||
dev_corpus = "corpora.dev"
|
||||
train_corpus = "corpora.train"
|
||||
seed = ${system.seed}
|
||||
gpu_allocator = ${system.gpu_allocator}
|
||||
dropout = 0.1
|
||||
accumulate_gradient = 3
|
||||
# 3000 steps is ~40 epochs over this 445k-token corpus, measured at ~29 steps/min on a T4
|
||||
# (~1.8h). The CPU tiers' 20000/1600 would be ~270 epochs and ~12h here, and worse than
|
||||
# wasteful: warmup_linear anneals against `total_steps`, so a run stopped early by patience
|
||||
# never leaves the peak learning rate. Budget and schedule are kept equal on purpose:
|
||||
# training.optimizer.learn_rate.total_steps must track any change to max_steps.
|
||||
patience = 600
|
||||
max_epochs = 0
|
||||
max_steps = 3000
|
||||
eval_frequency = 100
|
||||
frozen_components = []
|
||||
annotating_components = []
|
||||
before_to_disk = null
|
||||
before_update = null
|
||||
|
||||
[training.optimizer]
|
||||
@optimizers = "Adam.v1"
|
||||
beta1 = 0.9
|
||||
beta2 = 0.999
|
||||
L2_is_weight_decay = true
|
||||
L2 = 0.01
|
||||
grad_clip = 1.0
|
||||
use_averages = false
|
||||
eps = 1e-08
|
||||
|
||||
[training.optimizer.learn_rate]
|
||||
@schedules = "warmup_linear.v1"
|
||||
warmup_steps = 250
|
||||
total_steps = 3000
|
||||
initial_rate = 5e-5
|
||||
|
||||
[training.batcher]
|
||||
@batchers = "spacy.batch_by_padded.v1"
|
||||
discard_oversize = true
|
||||
size = 2000
|
||||
buffer = 256
|
||||
get_length = null
|
||||
|
||||
[training.logger]
|
||||
@loggers = "spacy.ConsoleLogger.v1"
|
||||
progress_bar = false
|
||||
|
||||
[training.score_weights]
|
||||
tag_acc = 0.16
|
||||
pos_acc = 0.08
|
||||
tag_micro_p = null
|
||||
tag_micro_r = null
|
||||
tag_micro_f = null
|
||||
morph_acc = 0.08
|
||||
morph_per_feat = null
|
||||
lemma_acc = 0.16
|
||||
dep_uas = 0.08
|
||||
dep_las = 0.16
|
||||
dep_las_per_type = null
|
||||
sents_p = null
|
||||
sents_r = null
|
||||
sents_f = 0.0
|
||||
ents_f = 0.28
|
||||
ents_p = 0.0
|
||||
ents_r = 0.0
|
||||
ents_per_type = null
|
||||
|
||||
[initialize]
|
||||
vectors = ${paths.vectors}
|
||||
init_tok2vec = ${paths.init_tok2vec}
|
||||
vocab_data = null
|
||||
lookups = null
|
||||
before_init = null
|
||||
after_init = null
|
||||
|
||||
[initialize.tokenizer]
|
||||
|
||||
[initialize.components]
|
||||
|
||||
[pretraining]
|
||||
|
|
@ -0,0 +1,231 @@
|
|||
# fa_dep_news_lg: tagger, morphologizer, trainable_lemmatizer, parser, WITH the lg-tier
|
||||
# static vectors.
|
||||
#
|
||||
# Byte-identical to configs/fa_dep_news_md.cfg. Only the vector table supplied at train time
|
||||
# via --paths.vectors differs: fa_floret, 200k rows x 300d, floret mode, trained on the full
|
||||
# Persian Wikipedia dump for 5 epochs (assets/vectors/fa_floret_lg), vs md's 50k rows x 300d
|
||||
# trained on 400k Persian documents. Seed, widths, rows, batcher, patience, eval_frequency all
|
||||
# held constant so the delta measures the vector table and nothing else.
|
||||
#
|
||||
# No `ner` here by design; see configs/fa_ner_lg.cfg and project.yml.
|
||||
|
||||
[paths]
|
||||
train = null
|
||||
dev = null
|
||||
vectors = null
|
||||
init_tok2vec = null
|
||||
|
||||
[system]
|
||||
gpu_allocator = null
|
||||
seed = 0
|
||||
|
||||
[nlp]
|
||||
lang = "fa"
|
||||
pipeline = ["tok2vec", "tagger", "morphologizer", "trainable_lemmatizer", "parser"]
|
||||
batch_size = 1000
|
||||
disabled = []
|
||||
before_creation = null
|
||||
after_creation = null
|
||||
after_pipeline_creation = null
|
||||
|
||||
[corpora]
|
||||
|
||||
[training]
|
||||
dev_corpus = "corpora.dev"
|
||||
train_corpus = "corpora.train"
|
||||
seed = ${system.seed}
|
||||
gpu_allocator = ${system.gpu_allocator}
|
||||
dropout = 0.1
|
||||
accumulate_gradient = 1
|
||||
patience = 1600
|
||||
max_epochs = 0
|
||||
max_steps = 20000
|
||||
eval_frequency = 400
|
||||
frozen_components = []
|
||||
annotating_components = []
|
||||
before_to_disk = null
|
||||
before_update = null
|
||||
|
||||
[initialize]
|
||||
vectors = ${paths.vectors}
|
||||
init_tok2vec = ${paths.init_tok2vec}
|
||||
vocab_data = null
|
||||
lookups = null
|
||||
before_init = null
|
||||
after_init = null
|
||||
|
||||
[components]
|
||||
|
||||
[pretraining]
|
||||
|
||||
[nlp.tokenizer]
|
||||
@tokenizers = "spacy.Tokenizer.v1"
|
||||
|
||||
[nlp.vectors]
|
||||
@vectors = "spacy.Vectors.v1"
|
||||
|
||||
[corpora.train]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.train}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[corpora.dev]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.dev}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[training.optimizer]
|
||||
@optimizers = "Adam.v1"
|
||||
beta1 = 0.9
|
||||
beta2 = 0.999
|
||||
L2_is_weight_decay = true
|
||||
L2 = 0.01
|
||||
grad_clip = 1.0
|
||||
use_averages = false
|
||||
eps = 1e-08
|
||||
learn_rate = 0.001
|
||||
|
||||
[training.batcher]
|
||||
@batchers = "spacy.batch_by_words.v1"
|
||||
discard_oversize = false
|
||||
tolerance = 0.2
|
||||
get_length = null
|
||||
|
||||
[training.logger]
|
||||
@loggers = "spacy.ConsoleLogger.v1"
|
||||
progress_bar = false
|
||||
|
||||
[training.score_weights]
|
||||
tag_acc = 0.25
|
||||
pos_acc = 0.12
|
||||
tag_micro_p = null
|
||||
tag_micro_r = null
|
||||
tag_micro_f = null
|
||||
morph_acc = 0.12
|
||||
morph_per_feat = null
|
||||
lemma_acc = 0.25
|
||||
dep_uas = 0.12
|
||||
dep_las = 0.12
|
||||
dep_las_per_type = null
|
||||
sents_p = null
|
||||
sents_r = null
|
||||
sents_f = 0.0
|
||||
|
||||
[initialize.tokenizer]
|
||||
|
||||
[initialize.components]
|
||||
|
||||
[components.tok2vec]
|
||||
factory = "tok2vec"
|
||||
|
||||
[components.tagger]
|
||||
factory = "tagger"
|
||||
label_smoothing = 0.05
|
||||
overwrite = false
|
||||
neg_prefix = "!"
|
||||
|
||||
[components.morphologizer]
|
||||
factory = "morphologizer"
|
||||
label_smoothing = 0.05
|
||||
overwrite = true
|
||||
extend = false
|
||||
|
||||
[components.trainable_lemmatizer]
|
||||
factory = "trainable_lemmatizer"
|
||||
backoff = "orth"
|
||||
min_tree_freq = 3
|
||||
overwrite = false
|
||||
top_k = 1
|
||||
|
||||
[components.parser]
|
||||
factory = "parser"
|
||||
moves = null
|
||||
update_with_oracle_cut_size = 100
|
||||
learn_tokens = false
|
||||
min_action_freq = 30
|
||||
|
||||
[training.batcher.size]
|
||||
@schedules = "compounding.v1"
|
||||
start = 100
|
||||
stop = 1000
|
||||
compound = 1.001
|
||||
t = 0.0
|
||||
|
||||
[components.tok2vec.model]
|
||||
@architectures = "spacy.Tok2Vec.v2"
|
||||
|
||||
[components.tagger.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.tagger.scorer]
|
||||
@scorers = "spacy.tagger_scorer.v1"
|
||||
|
||||
[components.morphologizer.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.morphologizer.scorer]
|
||||
@scorers = "spacy.morphologizer_scorer.v1"
|
||||
|
||||
[components.trainable_lemmatizer.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.trainable_lemmatizer.scorer]
|
||||
@scorers = "spacy.lemmatizer_scorer.v1"
|
||||
|
||||
[components.parser.model]
|
||||
@architectures = "spacy.TransitionBasedParser.v2"
|
||||
state_type = "parser"
|
||||
extra_state_tokens = false
|
||||
hidden_width = 128
|
||||
maxout_pieces = 3
|
||||
use_upper = true
|
||||
nO = null
|
||||
|
||||
[components.parser.scorer]
|
||||
@scorers = "spacy.parser_scorer.v1"
|
||||
|
||||
[components.tok2vec.model.embed]
|
||||
@architectures = "spacy.MultiHashEmbed.v2"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
attrs = ["NORM", "PREFIX", "SUFFIX", "SHAPE"]
|
||||
rows = [5000, 1000, 2500, 2500]
|
||||
include_static_vectors = true
|
||||
|
||||
[components.tok2vec.model.encode]
|
||||
@architectures = "spacy.MaxoutWindowEncoder.v2"
|
||||
width = 96
|
||||
depth = 4
|
||||
window_size = 1
|
||||
maxout_pieces = 3
|
||||
|
||||
[components.tagger.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
||||
[components.morphologizer.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
||||
[components.trainable_lemmatizer.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
||||
[components.parser.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
|
@ -0,0 +1,235 @@
|
|||
# fa_dep_news_md — tagger, morphologizer, trainable_lemmatizer, parser, WITH static vectors.
|
||||
#
|
||||
# Byte-identical to configs/fa_dep_news_sm.cfg except:
|
||||
# - [components.tok2vec.model.embed] include_static_vectors: false -> true
|
||||
#
|
||||
# Everything else (seed, widths, rows, batcher, patience, eval_frequency) is held constant so
|
||||
# the sm/md delta measures the floret vectors and nothing else.
|
||||
#
|
||||
# Vectors are supplied at train time via --paths.vectors, pointing at the fa_floret table
|
||||
# (50k rows x 300d, floret mode, minn=maxn=5, hash_count=2) trained on 400k Persian documents.
|
||||
# floret has no OOV: every string hashes into the table, which is the point for Persian, where
|
||||
# ZWNJ inconsistency (میرود / میرود / می رود) would shatter a classic word table.
|
||||
#
|
||||
# No `ner` here by design; see configs/fa_ner_md.cfg and project.yml.
|
||||
|
||||
[paths]
|
||||
train = null
|
||||
dev = null
|
||||
vectors = null
|
||||
init_tok2vec = null
|
||||
|
||||
[system]
|
||||
gpu_allocator = null
|
||||
seed = 0
|
||||
|
||||
[nlp]
|
||||
lang = "fa"
|
||||
pipeline = ["tok2vec", "tagger", "morphologizer", "trainable_lemmatizer", "parser"]
|
||||
batch_size = 1000
|
||||
disabled = []
|
||||
before_creation = null
|
||||
after_creation = null
|
||||
after_pipeline_creation = null
|
||||
|
||||
[corpora]
|
||||
|
||||
[training]
|
||||
dev_corpus = "corpora.dev"
|
||||
train_corpus = "corpora.train"
|
||||
seed = ${system.seed}
|
||||
gpu_allocator = ${system.gpu_allocator}
|
||||
dropout = 0.1
|
||||
accumulate_gradient = 1
|
||||
patience = 1600
|
||||
max_epochs = 0
|
||||
max_steps = 20000
|
||||
eval_frequency = 400
|
||||
frozen_components = []
|
||||
annotating_components = []
|
||||
before_to_disk = null
|
||||
before_update = null
|
||||
|
||||
[initialize]
|
||||
vectors = ${paths.vectors}
|
||||
init_tok2vec = ${paths.init_tok2vec}
|
||||
vocab_data = null
|
||||
lookups = null
|
||||
before_init = null
|
||||
after_init = null
|
||||
|
||||
[components]
|
||||
|
||||
[pretraining]
|
||||
|
||||
[nlp.tokenizer]
|
||||
@tokenizers = "spacy.Tokenizer.v1"
|
||||
|
||||
[nlp.vectors]
|
||||
@vectors = "spacy.Vectors.v1"
|
||||
|
||||
[corpora.train]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.train}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[corpora.dev]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.dev}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[training.optimizer]
|
||||
@optimizers = "Adam.v1"
|
||||
beta1 = 0.9
|
||||
beta2 = 0.999
|
||||
L2_is_weight_decay = true
|
||||
L2 = 0.01
|
||||
grad_clip = 1.0
|
||||
use_averages = false
|
||||
eps = 1e-08
|
||||
learn_rate = 0.001
|
||||
|
||||
[training.batcher]
|
||||
@batchers = "spacy.batch_by_words.v1"
|
||||
discard_oversize = false
|
||||
tolerance = 0.2
|
||||
get_length = null
|
||||
|
||||
[training.logger]
|
||||
@loggers = "spacy.ConsoleLogger.v1"
|
||||
progress_bar = false
|
||||
|
||||
[training.score_weights]
|
||||
tag_acc = 0.25
|
||||
pos_acc = 0.12
|
||||
tag_micro_p = null
|
||||
tag_micro_r = null
|
||||
tag_micro_f = null
|
||||
morph_acc = 0.12
|
||||
morph_per_feat = null
|
||||
lemma_acc = 0.25
|
||||
dep_uas = 0.12
|
||||
dep_las = 0.12
|
||||
dep_las_per_type = null
|
||||
sents_p = null
|
||||
sents_r = null
|
||||
sents_f = 0.0
|
||||
|
||||
[initialize.tokenizer]
|
||||
|
||||
[initialize.components]
|
||||
|
||||
[components.tok2vec]
|
||||
factory = "tok2vec"
|
||||
|
||||
[components.tagger]
|
||||
factory = "tagger"
|
||||
label_smoothing = 0.05
|
||||
overwrite = false
|
||||
neg_prefix = "!"
|
||||
|
||||
[components.morphologizer]
|
||||
factory = "morphologizer"
|
||||
label_smoothing = 0.05
|
||||
overwrite = true
|
||||
extend = false
|
||||
|
||||
[components.trainable_lemmatizer]
|
||||
factory = "trainable_lemmatizer"
|
||||
backoff = "orth"
|
||||
min_tree_freq = 3
|
||||
overwrite = false
|
||||
top_k = 1
|
||||
|
||||
[components.parser]
|
||||
factory = "parser"
|
||||
moves = null
|
||||
update_with_oracle_cut_size = 100
|
||||
learn_tokens = false
|
||||
min_action_freq = 30
|
||||
|
||||
[training.batcher.size]
|
||||
@schedules = "compounding.v1"
|
||||
start = 100
|
||||
stop = 1000
|
||||
compound = 1.001
|
||||
t = 0.0
|
||||
|
||||
[components.tok2vec.model]
|
||||
@architectures = "spacy.Tok2Vec.v2"
|
||||
|
||||
[components.tagger.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.tagger.scorer]
|
||||
@scorers = "spacy.tagger_scorer.v1"
|
||||
|
||||
[components.morphologizer.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.morphologizer.scorer]
|
||||
@scorers = "spacy.morphologizer_scorer.v1"
|
||||
|
||||
[components.trainable_lemmatizer.model]
|
||||
@architectures = "spacy.Tagger.v2"
|
||||
nO = null
|
||||
normalize = false
|
||||
|
||||
[components.trainable_lemmatizer.scorer]
|
||||
@scorers = "spacy.lemmatizer_scorer.v1"
|
||||
|
||||
[components.parser.model]
|
||||
@architectures = "spacy.TransitionBasedParser.v2"
|
||||
state_type = "parser"
|
||||
extra_state_tokens = false
|
||||
hidden_width = 128
|
||||
maxout_pieces = 3
|
||||
use_upper = true
|
||||
nO = null
|
||||
|
||||
[components.parser.scorer]
|
||||
@scorers = "spacy.parser_scorer.v1"
|
||||
|
||||
[components.tok2vec.model.embed]
|
||||
@architectures = "spacy.MultiHashEmbed.v2"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
attrs = ["NORM", "PREFIX", "SUFFIX", "SHAPE"]
|
||||
rows = [5000, 1000, 2500, 2500]
|
||||
include_static_vectors = true
|
||||
|
||||
[components.tok2vec.model.encode]
|
||||
@architectures = "spacy.MaxoutWindowEncoder.v2"
|
||||
width = 96
|
||||
depth = 4
|
||||
window_size = 1
|
||||
maxout_pieces = 3
|
||||
|
||||
[components.tagger.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
||||
[components.morphologizer.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
||||
[components.trainable_lemmatizer.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
||||
[components.parser.model.tok2vec]
|
||||
@architectures = "spacy.Tok2VecListener.v1"
|
||||
width = ${components.tok2vec.model.encode.width}
|
||||
upstream = "*"
|
||||
|
|
@ -0,0 +1,154 @@
|
|||
# fa_ent_news_lg: Persian NER with the lg-tier static floret vectors.
|
||||
#
|
||||
# Identical to configs/fa_ner_md.cfg (which is identical to fa_ner_sm.cfg except
|
||||
# include_static_vectors: true). Only the vector table supplied at train time via
|
||||
# --paths.vectors differs: fa_floret, 200k rows x 300d, floret mode, trained on the full
|
||||
# Persian Wikipedia dump for 5 epochs (assets/vectors/fa_floret_lg), vs md's 50k rows x 300d
|
||||
# trained on 400k Persian documents.
|
||||
#
|
||||
# Same embedded-tok2vec design as sm/md (no Tok2VecListener), so the trained component stays
|
||||
# sourceable into a future fa_core_news_lg via `nlp.add_pipe("ner", source=...)`.
|
||||
|
||||
[paths]
|
||||
train = null
|
||||
dev = null
|
||||
vectors = null
|
||||
init_tok2vec = null
|
||||
|
||||
[system]
|
||||
gpu_allocator = null
|
||||
seed = 0
|
||||
|
||||
[nlp]
|
||||
lang = "fa"
|
||||
pipeline = ["ner"]
|
||||
batch_size = 1000
|
||||
disabled = []
|
||||
before_creation = null
|
||||
after_creation = null
|
||||
after_pipeline_creation = null
|
||||
|
||||
[nlp.tokenizer]
|
||||
@tokenizers = "spacy.Tokenizer.v1"
|
||||
|
||||
[nlp.vectors]
|
||||
@vectors = "spacy.Vectors.v1"
|
||||
|
||||
[components]
|
||||
|
||||
[components.ner]
|
||||
factory = "ner"
|
||||
moves = null
|
||||
update_with_oracle_cut_size = 100
|
||||
incorrect_spans_key = null
|
||||
|
||||
[components.ner.model]
|
||||
@architectures = "spacy.TransitionBasedParser.v2"
|
||||
state_type = "ner"
|
||||
extra_state_tokens = false
|
||||
hidden_width = 64
|
||||
maxout_pieces = 2
|
||||
use_upper = true
|
||||
nO = null
|
||||
|
||||
[components.ner.model.tok2vec]
|
||||
@architectures = "spacy.Tok2Vec.v2"
|
||||
|
||||
[components.ner.model.tok2vec.embed]
|
||||
@architectures = "spacy.MultiHashEmbed.v2"
|
||||
width = ${components.ner.model.tok2vec.encode.width}
|
||||
attrs = ["NORM", "PREFIX", "SUFFIX", "SHAPE"]
|
||||
rows = [5000, 1000, 2500, 2500]
|
||||
include_static_vectors = true
|
||||
|
||||
[components.ner.model.tok2vec.encode]
|
||||
@architectures = "spacy.MaxoutWindowEncoder.v2"
|
||||
width = 96
|
||||
depth = 4
|
||||
window_size = 1
|
||||
maxout_pieces = 3
|
||||
|
||||
[components.ner.scorer]
|
||||
@scorers = "spacy.ner_scorer.v1"
|
||||
|
||||
[corpora]
|
||||
|
||||
[corpora.train]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.train}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[corpora.dev]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.dev}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[training]
|
||||
dev_corpus = "corpora.dev"
|
||||
train_corpus = "corpora.train"
|
||||
seed = ${system.seed}
|
||||
gpu_allocator = ${system.gpu_allocator}
|
||||
dropout = 0.1
|
||||
accumulate_gradient = 1
|
||||
patience = 1600
|
||||
max_epochs = 0
|
||||
max_steps = 20000
|
||||
eval_frequency = 400
|
||||
frozen_components = []
|
||||
annotating_components = []
|
||||
before_to_disk = null
|
||||
before_update = null
|
||||
|
||||
[training.optimizer]
|
||||
@optimizers = "Adam.v1"
|
||||
beta1 = 0.9
|
||||
beta2 = 0.999
|
||||
L2_is_weight_decay = true
|
||||
L2 = 0.01
|
||||
grad_clip = 1.0
|
||||
use_averages = false
|
||||
eps = 1e-08
|
||||
learn_rate = 0.001
|
||||
|
||||
[training.batcher]
|
||||
@batchers = "spacy.batch_by_words.v1"
|
||||
discard_oversize = false
|
||||
tolerance = 0.2
|
||||
get_length = null
|
||||
|
||||
[training.batcher.size]
|
||||
@schedules = "compounding.v1"
|
||||
start = 100
|
||||
stop = 1000
|
||||
compound = 1.001
|
||||
t = 0.0
|
||||
|
||||
[training.logger]
|
||||
@loggers = "spacy.ConsoleLogger.v1"
|
||||
progress_bar = false
|
||||
|
||||
[training.score_weights]
|
||||
ents_f = 1.0
|
||||
ents_p = 0.0
|
||||
ents_r = 0.0
|
||||
ents_per_type = null
|
||||
|
||||
[initialize]
|
||||
vectors = ${paths.vectors}
|
||||
init_tok2vec = ${paths.init_tok2vec}
|
||||
vocab_data = null
|
||||
lookups = null
|
||||
before_init = null
|
||||
after_init = null
|
||||
|
||||
[initialize.tokenizer]
|
||||
|
||||
[initialize.components]
|
||||
|
||||
[pretraining]
|
||||
|
|
@ -0,0 +1,154 @@
|
|||
# fa_ent_news_md — Persian NER with static floret vectors.
|
||||
#
|
||||
# Identical to configs/fa_ner_sm.cfg except:
|
||||
# - [components.ner.model.tok2vec.embed] include_static_vectors: false -> true
|
||||
#
|
||||
# Same embedded-tok2vec design as the sm variant (no Tok2VecListener), so the trained
|
||||
# component stays sourceable into fa_core_news_md via `nlp.add_pipe("ner", source=...)`.
|
||||
#
|
||||
# Vectors supplied at train time via --paths.vectors (fa_floret, 50k rows x 300d,
|
||||
# trained on 400k Persian documents).
|
||||
|
||||
[paths]
|
||||
train = null
|
||||
dev = null
|
||||
vectors = null
|
||||
init_tok2vec = null
|
||||
|
||||
[system]
|
||||
gpu_allocator = null
|
||||
seed = 0
|
||||
|
||||
[nlp]
|
||||
lang = "fa"
|
||||
pipeline = ["ner"]
|
||||
batch_size = 1000
|
||||
disabled = []
|
||||
before_creation = null
|
||||
after_creation = null
|
||||
after_pipeline_creation = null
|
||||
|
||||
[nlp.tokenizer]
|
||||
@tokenizers = "spacy.Tokenizer.v1"
|
||||
|
||||
[nlp.vectors]
|
||||
@vectors = "spacy.Vectors.v1"
|
||||
|
||||
[components]
|
||||
|
||||
[components.ner]
|
||||
factory = "ner"
|
||||
moves = null
|
||||
update_with_oracle_cut_size = 100
|
||||
incorrect_spans_key = null
|
||||
|
||||
[components.ner.model]
|
||||
@architectures = "spacy.TransitionBasedParser.v2"
|
||||
state_type = "ner"
|
||||
extra_state_tokens = false
|
||||
hidden_width = 64
|
||||
maxout_pieces = 2
|
||||
use_upper = true
|
||||
nO = null
|
||||
|
||||
[components.ner.model.tok2vec]
|
||||
@architectures = "spacy.Tok2Vec.v2"
|
||||
|
||||
[components.ner.model.tok2vec.embed]
|
||||
@architectures = "spacy.MultiHashEmbed.v2"
|
||||
width = ${components.ner.model.tok2vec.encode.width}
|
||||
attrs = ["NORM", "PREFIX", "SUFFIX", "SHAPE"]
|
||||
rows = [5000, 1000, 2500, 2500]
|
||||
include_static_vectors = true
|
||||
|
||||
[components.ner.model.tok2vec.encode]
|
||||
@architectures = "spacy.MaxoutWindowEncoder.v2"
|
||||
width = 96
|
||||
depth = 4
|
||||
window_size = 1
|
||||
maxout_pieces = 3
|
||||
|
||||
[components.ner.scorer]
|
||||
@scorers = "spacy.ner_scorer.v1"
|
||||
|
||||
[corpora]
|
||||
|
||||
[corpora.train]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.train}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[corpora.dev]
|
||||
@readers = "spacy.Corpus.v1"
|
||||
path = ${paths.dev}
|
||||
max_length = 0
|
||||
gold_preproc = false
|
||||
limit = 0
|
||||
augmenter = null
|
||||
|
||||
[training]
|
||||
dev_corpus = "corpora.dev"
|
||||
train_corpus = "corpora.train"
|
||||
seed = ${system.seed}
|
||||
gpu_allocator = ${system.gpu_allocator}
|
||||
dropout = 0.1
|
||||
accumulate_gradient = 1
|
||||
patience = 1600
|
||||
max_epochs = 0
|
||||
max_steps = 20000
|
||||
eval_frequency = 400
|
||||
frozen_components = []
|
||||
annotating_components = []
|
||||
before_to_disk = null
|
||||
before_update = null
|
||||
|
||||
[training.optimizer]
|
||||
@optimizers = "Adam.v1"
|
||||
beta1 = 0.9
|
||||
beta2 = 0.999
|
||||
L2_is_weight_decay = true
|
||||
L2 = 0.01
|
||||
grad_clip = 1.0
|
||||
use_averages = false
|
||||
eps = 1e-08
|
||||
learn_rate = 0.001
|
||||
|
||||
[training.batcher]
|
||||
@batchers = "spacy.batch_by_words.v1"
|
||||
discard_oversize = false
|
||||
tolerance = 0.2
|
||||
get_length = null
|
||||
|
||||
[training.batcher.size]
|
||||
@schedules = "compounding.v1"
|
||||
start = 100
|
||||
stop = 1000
|
||||
compound = 1.001
|
||||
t = 0.0
|
||||
|
||||
[training.logger]
|
||||
@loggers = "spacy.ConsoleLogger.v1"
|
||||
progress_bar = false
|
||||
|
||||
[training.score_weights]
|
||||
ents_f = 1.0
|
||||
ents_p = 0.0
|
||||
ents_r = 0.0
|
||||
ents_per_type = null
|
||||
|
||||
[initialize]
|
||||
vectors = ${paths.vectors}
|
||||
init_tok2vec = ${paths.init_tok2vec}
|
||||
vocab_data = null
|
||||
lookups = null
|
||||
before_init = null
|
||||
after_init = null
|
||||
|
||||
[initialize.tokenizer]
|
||||
|
||||
[initialize.components]
|
||||
|
||||
[pretraining]
|
||||
|
|
@ -56,7 +56,12 @@ From <https://github.com/explosion/spaCy/blob/master/CONTRIBUTING.md>:
|
|||
python -m spacy package training/fa_dep_news_sm packages --name dep_news_sm --version 3.8.0 --build wheel
|
||||
python -m spacy huggingface-hub push packages/fa_dep_news_sm-3.8.0/dist/fa_dep_news_sm-3.8.0-py3-none-any.whl --org <org>
|
||||
```
|
||||
Users then `pip install https://huggingface.co/<org>/fa_dep_news_sm/resolve/main/fa_dep_news_sm-any-py3-none-any.whl`.
|
||||
Users then `pip install https://huggingface.co/<org>/fa_dep_news_sm/resolve/main/fa_dep_news_sm-3.8.0-py3-none-any.whl`.
|
||||
Note the filename: `spacy huggingface-hub push` uploads the wheel as `<name>-any-py3-none-any.whl`,
|
||||
but `"any"` is not a valid PEP 440 version and current pip rejects it
|
||||
(`Invalid wheel filename (invalid version)`). Upload a second copy under its real versioned
|
||||
filename too (`api.upload_file(path_in_repo=f"{name}-{version}-py3-none-any.whl", ...)`) and
|
||||
link to that one instead.
|
||||
2. PyPI or a self-hosted wheel: `spacy package … --build sdist,wheel` then `twine upload`, or
|
||||
attach the wheel to a GitHub Release. See <https://spacy.io/api/cli#package>.
|
||||
3. spaCy Universe, which lists the package on spacy.io but hosts nothing. Per
|
||||
|
|
|
|||
252
docs/MODELS.md
252
docs/MODELS.md
|
|
@ -21,6 +21,13 @@ requirement. That ordering sets the roadmap below.
|
|||
|
||||
Source: <https://spacy.io/models/en>.
|
||||
|
||||
That last point does **not** transfer to Persian. The `sm` -> `md` step measured on this
|
||||
project buys +1.19 LAS and +2.85 NER F (§6), where English gets ~0.00 LAS. Two reasons: PerDT
|
||||
is roughly a tenth the size of OntoNotes, so hash embeddings have far less signal to learn a
|
||||
lexicon from, and floret's subword hashing gives 0% OOV on a language whose ZWNJ variation
|
||||
(میرود / میرود / می رود) fragments any fixed word-key table. English `md` uses 20k classic
|
||||
word vectors and hits OOV constantly. Do not use the English row as the Persian prior.
|
||||
|
||||
## 2. Target: the Persian pipelines
|
||||
|
||||
Naming follows `[lang]_[type]_[genre]_[size]` (<https://spacy.io/models#conventions>). The
|
||||
|
|
@ -35,9 +42,12 @@ pipelines such as `de_core_news_sm` as `news`.
|
|||
| `fa_core_news_sm` | the above plus ner | hash embeddings | built, shipping |
|
||||
| `fa_ent_news_sm` | ner (own internal tok2vec) | hash embeddings | built, optional |
|
||||
| `fa_core_web_sm` | same as core, mixed-genre training data | hash embeddings | not built; would add ParsTwiNER to cover social media |
|
||||
| `fa_core_news_md` | + static vectors | floret, 50k rows | vectors must be trained first (CPU-days on fa Wikipedia + OSCAR) |
|
||||
| `fa_core_news_lg` | same | floret, 200k rows | same as md, bigger table |
|
||||
| `fa_core_news_trf` | transformer instead of tok2vec | `HooshvareLab/roberta-fa-zwnj-base` (Apache-2.0) | not on this hardware; 2 GB VRAM cannot fine-tune a 125M-param encoder |
|
||||
| `fa_dep_news_md` | same as `fa_dep_news_sm` | floret, 50k rows / 300d | built, shipping |
|
||||
| `fa_core_news_md` | same as `fa_core_news_sm` | floret, 50k rows / 300d | built, shipping |
|
||||
| `fa_dep_news_lg` | same as `fa_dep_news_sm` | floret, 200k rows / 300d, full-wiki 5 epochs | built, shipping |
|
||||
| `fa_core_news_lg` | same as `fa_core_news_sm` | floret, 200k rows / 300d, full-wiki 5 epochs | built, shipping |
|
||||
| `fa_ent_news_lg` | ner (own internal tok2vec) | floret, 200k rows / 300d, full-wiki 5 epochs | built, optional |
|
||||
| `fa_core_news_trf` | transformer instead of tok2vec | `HooshvareLab/bert-base-parsbert-uncased`, fine-tuned | built on a rented Colab T4 (not on this hardware: 2 GB VRAM cannot fine-tune a 125M-param encoder), shipping with a redistribution caveat because that encoder's card states no licence; §3.4 and §8 |
|
||||
|
||||
### Why `core` is honest here
|
||||
|
||||
|
|
@ -276,3 +286,239 @@ Sources are recorded in each `meta.json` with their licences, per
|
|||
crediting Beheshti-NER. CC BY-SA 4.0 on PerDT means every package carries attribution and a
|
||||
share-alike notice, handled in `scripts/finalize_pipeline.py`, which also enforces the shape of
|
||||
each variant: it refuses to publish a `dep` pipeline containing `ner`, or a `core` one without it.
|
||||
|
||||
## 6. The `md` tier: floret static vectors
|
||||
|
||||
Built after the `sm` tier, from `fa_floret`: 50,000 rows x 300d, floret mode, `minn=maxn=5`,
|
||||
`hash_count=2`, trained on 400,000 Persian documents. The wheel is a vectors-only pipeline;
|
||||
`scripts/unpack_vectors.py` unwraps it into a directory `--paths.vectors` can read, so nothing
|
||||
needs pip-installing to train against it.
|
||||
|
||||
`configs/fa_dep_news_md.cfg` and `configs/fa_ner_md.cfg` are their `sm` counterparts with one
|
||||
line changed, `include_static_vectors = false -> true`. Same seed, same corpus, same widths,
|
||||
same batcher, same patience. The deltas below are therefore attributable to the vector table
|
||||
and nothing else. Reproduce with `spacy project run md`, or the table alone with
|
||||
`python scripts/compare_tiers.py`.
|
||||
|
||||
### UD_Persian-PerDT test split
|
||||
|
||||
| Metric | `sm` | `md` | Delta |
|
||||
| --- | --- | --- | --- |
|
||||
| `TAG_ACC` | 95.96 | 96.25 | +0.29 |
|
||||
| `POS_ACC` | 96.24 | 96.64 | +0.40 |
|
||||
| `MORPH_ACC` | 96.29 | 96.64 | +0.35 |
|
||||
| `LEMMA_ACC` | 97.91 | 97.96 | +0.05 |
|
||||
| `SENTS_F` | 99.25 | 99.28 | +0.03 |
|
||||
| `DEP_UAS` | 89.69 | 90.52 | +0.83 |
|
||||
| `DEP_LAS` | 85.15 | 86.34 | +1.19 |
|
||||
| Speed (dep) | 12,505 w/s | 10,493 w/s | -16.1% |
|
||||
|
||||
### PerDT NER test split, `fa_core_news_md`
|
||||
|
||||
| Metric | `sm` | `md` | Delta |
|
||||
| --- | --- | --- | --- |
|
||||
| `ENTS_P` | 77.67 | 76.56 | -1.10 |
|
||||
| `ENTS_R` | 66.87 | 72.95 | +6.08 |
|
||||
| `ENTS_F` | 71.87 | 74.71 | +2.85 |
|
||||
|
||||
Almost all of the NER gain is recall. That is the expected shape of a fix for a coverage
|
||||
problem: hash embeddings had no lexical prior for rare proper nouns, so the `sm` model
|
||||
declined to tag them. Precision slips ~1 point because the model now guesses more.
|
||||
|
||||
| Label | Gold in test | `sm` F | `md` F | Delta |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `PER` | 297 | 65.29 | 68.18 | +2.89 |
|
||||
| `LOC` | 273 | 80.24 | 84.05 | +3.81 |
|
||||
| `ORG` | 144 | 68.77 | 70.25 | +1.48 |
|
||||
| `DAT` | 69 | 74.45 | 76.19 | +1.74 |
|
||||
| `MON` | 10 | 73.68 | 84.21 | +10.53 |
|
||||
| `TIM` | 9 | 66.67 | 66.67 | +0.00 |
|
||||
| `PCT` | 4 | 57.14 | 33.33 | -23.81 |
|
||||
|
||||
Read the bottom three rows as noise, not signal. `PCT` has four gold entities in the whole
|
||||
test split, so its -23.81 F is one entity changing hands; `MON`'s +10.53 is likewise one of
|
||||
ten. The three labels with real support (`PER`, `LOC`, `ORG`, 714 entities between them) all
|
||||
improve, which is the finding.
|
||||
|
||||
### Cost
|
||||
|
||||
The vectors dominate the artifact: `fa_dep_news_md` is a 62 MB wheel against 7.5 MB for `sm`,
|
||||
`fa_core_news_md` 68 MB against 13 MB. Inference is ~16% slower across all three pipelines,
|
||||
a uniform hit consistent with the extra 300d concatenation per token rather than anything
|
||||
component-specific. Training cost was comparable to `sm` (early stop at step 12,400 of 20,000,
|
||||
best checkpoint near 10,800).
|
||||
|
||||
Whether that trade is worth it depends on deployment. For a 1.19 LAS and 2.85 NER F gain, a
|
||||
9x larger download and 16% slower parse is a good deal on a server and a bad one in a browser
|
||||
or a Lambda cold start. Both tiers ship; pick per target.
|
||||
|
||||
## 7. The `lg` tier: bigger floret table, full pipeline
|
||||
|
||||
Built after `md`, from a new `fa_floret` table: 200,000 rows x 300d, floret mode,
|
||||
`minn=maxn=5`, `hash_count=2`, trained on the full Persian Wikipedia dump for 5 epochs (4x
|
||||
the rows of `md`'s 50k-row table trained on 400k documents). Raw `.floret`/`.vec` and the
|
||||
packaged spaCy wheel are at <https://huggingface.co/Phazel/fa-floret-wiki-vectors>. Unpacked
|
||||
the same way as `md` via `scripts/unpack_vectors.py`, into `assets/vectors/fa_floret_lg`.
|
||||
|
||||
`configs/fa_ner_lg.cfg` and `configs/fa_dep_news_lg.cfg` are `fa_ner_md.cfg`/
|
||||
`fa_dep_news_md.cfg` unchanged except `--paths.vectors`. Same seed, same corpus, same
|
||||
architecture as `sm`/`md` throughout, so the deltas below are attributable to the vector
|
||||
table alone. Reproduce with `spacy project run lg`, or the tables alone with
|
||||
`python scripts/compare_tiers.py`.
|
||||
|
||||
### UD test split, `fa_dep_news_lg` / `fa_core_news_lg`
|
||||
|
||||
| Metric | `sm` | `md` | `lg` | Delta (lg vs sm) | Delta (lg vs md) |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `TAG_ACC` | 95.96 | 96.25 | 96.55 | +0.59 | +0.30 |
|
||||
| `POS_ACC` | 96.24 | 96.64 | 96.68 | +0.44 | +0.04 |
|
||||
| `MORPH_ACC` | 96.29 | 96.64 | 96.70 | +0.41 | +0.06 |
|
||||
| `LEMMA_ACC` | 97.91 | 97.96 | 98.08 | +0.17 | +0.12 |
|
||||
| `DEP_UAS` | 89.69 | 90.52 | 90.96 | +1.27 | +0.44 |
|
||||
| `DEP_LAS` | 85.15 | 86.34 | 86.60 | +1.45 | +0.26 |
|
||||
|
||||
`lg` beats `md` on every UD metric, the same monotonic pattern as `md` beating `sm` in §6.
|
||||
The bigger, less collision-prone floret table keeps paying off, though the `md`-to-`lg`
|
||||
gains (4x the vector rows) are smaller than the `sm`-to-`md` gains (going from none to 50k
|
||||
rows): diminishing returns, as expected.
|
||||
|
||||
### PerDT NER test split, `fa_ent_news_lg` (identical `ner` component embedded in `fa_core_news_lg`)
|
||||
|
||||
| Metric | `sm` | `md` | `lg` | Delta (lg vs sm) | Delta (lg vs md) |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `ENTS_P` | 77.67 | 76.56 | 81.51 | +3.84 | +4.95 |
|
||||
| `ENTS_R` | 66.87 | 72.95 | 71.09 | +4.22 | -1.86 |
|
||||
| `ENTS_F` | 71.87 | 74.71 | 75.94 | +4.08 | +1.23 |
|
||||
|
||||
`lg` beats both `sm` and `md` on `ENTS_F`, and unlike `md`'s recall-only gain over `sm`, `lg`
|
||||
improves precision too (+3.84 over `sm`, whereas `md` cost -1.10). Consistent with a bigger,
|
||||
less collision-prone floret table giving both better recall on rare proper nouns and fewer
|
||||
false positives from hash collisions.
|
||||
|
||||
| Label | Gold in test | `sm` F | `md` F | `lg` F | Delta (lg vs sm) |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| `PER` | 297 | 65.29 | 68.18 | 72.63 | +7.33 |
|
||||
| `LOC` | 273 | 80.24 | 84.05 | 83.66 | +3.42 |
|
||||
| `ORG` | 144 | 68.77 | 70.25 | 71.01 | +2.24 |
|
||||
| `DAT` | 69 | 74.45 | 76.19 | 70.83 | -3.62 |
|
||||
| `MON` | 10 | 73.68 | 84.21 | 88.89 | +15.20 |
|
||||
| `TIM` | 9 | 66.67 | 66.67 | 61.54 | -5.13 |
|
||||
| `PCT` | 4 | 57.14 | 33.33 | 57.14 | +0.00 |
|
||||
|
||||
`PER`, `LOC` and `ORG` (714 entities, the labels with real support) all improve over both
|
||||
smaller tiers. `DAT` and `TIM` regress a few points against `md`; `MON`/`TIM`/`PCT` swings are
|
||||
one-or-two-entity noise, same caveat as §6.
|
||||
|
||||
### Cost
|
||||
|
||||
The bigger table dominates the artifact even more than `md`'s did: the 200k x 300d float32
|
||||
vector table is ~240 MB uncompressed, so `fa_dep_news_lg` is a 219 MB wheel (vs 7.5 MB `sm`,
|
||||
60 MB `md`), `fa_core_news_lg` 225 MB (vs 13 MB `sm`, 66 MB `md`), and `fa_ent_news_lg` alone
|
||||
217 MB (vs 5.6 MB `sm`, 58 MB `md`). Training cost roughly doubled `md`'s: `dep_lg` ran to
|
||||
early stop at step 12,000 of 20,000 over ~2h08m CPU wall time (vs `dep_md`'s single-digit
|
||||
minutes territory implied by its architecture-identical config; `lg`'s extra time is
|
||||
entirely the larger embedding table's per-step cost, not more steps). `ner_lg` early-stopped
|
||||
at step 7,200, ~13 min, in line with `sm`/`md`.
|
||||
|
||||
`words/s` from `spacy benchmark accuracy` were noisier at this tier than `sm`-vs-`md`: dep/core
|
||||
throughput dropped as expected (9,387 / 6,655 words/s vs `sm`'s 12,505 / 8,834, `md`'s
|
||||
10,493 / 7,269 words/s; the larger table costs real lookup time), but the standalone `ent_lg` run
|
||||
showed 15,614 words/s, higher than `sm`/`md`'s ent runs despite an identical `ner`
|
||||
architecture and the same larger table. That figure was single-run CPU contention noise on
|
||||
shared hardware, not a real speedup. Those numbers are superseded by §9, which times
|
||||
`nlp.pipe` alone instead of reading a scoring-contaminated figure off the benchmark command.
|
||||
|
||||
For a 4x download over `md` (and up to 39x over `sm`) buying +1.45 DEP_LAS / +1.23 ENTS_F
|
||||
over `md` (+1.45 DEP_LAS / +4.08 ENTS_F over `sm`), `lg` is a server/offline-batch pipeline,
|
||||
not something to ship to a browser or a cold-start function. All three variants (`dep`,
|
||||
`ent`, `core`) are built and evaluated at this tier, same as `md`.
|
||||
|
||||
## 8. The `trf` tier: one fine-tuned ParsBERT
|
||||
|
||||
`configs/fa_core_news_trf.cfg` replaces the static-vector tok2vec with
|
||||
`HooshvareLab/bert-base-parsbert-uncased`, fine-tuned during training. Trained on a rented
|
||||
Colab T4 in 1h58m: 3000 steps, no early stop, the full learning-rate anneal.
|
||||
|
||||
### One corpus, because a transformer cannot be trained twice
|
||||
|
||||
The `sm`/`md`/`lg` tiers train `ner` as its own pipeline with its own embedded tok2vec and
|
||||
then source it into the dep model. That is affordable because a hash-embed tok2vec is cheap.
|
||||
A 162M-parameter encoder is not: fine-tuning it once per component would double GPU cost and
|
||||
put two encoders in one wheel, and sourcing the second would collide on the `transformer`
|
||||
component name.
|
||||
|
||||
So every component listens to a single shared transformer through a `TransformerListener`,
|
||||
which requires one corpus carrying both the UD and NER annotation layers on the same `Doc`.
|
||||
`scripts/merge_joint_corpus.py` builds it. The fusion is exact rather than approximate:
|
||||
`corpus/perdt-ner/` was converted from the same `--merge-subtokens` CoNLL-U as
|
||||
`corpus/merged/` with the same `--n-sents`, so the two DocBins are token-for-token identical.
|
||||
The script asserts that per document and copies only `doc.ents` across. Char offsets are not
|
||||
usable for the copy, because the two converters differ in trailing whitespace, which shifts
|
||||
`char_span` off the token grid and returns None; the transfer goes by token index.
|
||||
|
||||
### Results against `lg`
|
||||
|
||||
| Metric | `lg` | `trf` | Delta |
|
||||
| --- | ---: | ---: | ---: |
|
||||
| `TAG_ACC` | 96.55 | 97.62 | +1.07 |
|
||||
| `POS_ACC` | 96.68 | 97.63 | +0.95 |
|
||||
| `MORPH_ACC` | 96.70 | 97.82 | +1.12 |
|
||||
| `LEMMA_ACC` | 98.08 | 97.31 | -0.77 |
|
||||
| `DEP_UAS` | 90.96 | 93.87 | +2.91 |
|
||||
| `DEP_LAS` | 86.60 | 90.79 | +4.19 |
|
||||
| `SENTS_F` | 99.18 | 97.35 | -1.83 |
|
||||
| `ENTS_F` | 75.94 | 82.89 | +6.95 |
|
||||
|
||||
The parser gain is the headline: `DEP_LAS` 90.79 passes the hazm+ParsBERT reference of 89.34,
|
||||
which no CPU tier reached. NER gains 6.95 F, almost all of it recall (71.09 to 81.76) at
|
||||
higher precision, which is what a pretrained encoder buys on the difflib-transferred layer.
|
||||
|
||||
Two metrics regress. `SENTS_F` drops 1.83, most likely because `strided_spans` at
|
||||
`window = 128, stride = 96` leaves 32 tokens of overlap, so tokens near a span edge see
|
||||
truncated right context where the CPU tiers' tok2vec sees the whole doc. `LEMMA_ACC` drops
|
||||
0.77 and is the one metric where a static-vector tier wins: `trainable_lemmatizer` reads a
|
||||
single `reduce_mean`-pooled vector per token, while `lg` runs an edit-tree lemmatizer over
|
||||
floret subwords that model Persian orthography directly. Neither is a training-length
|
||||
problem; see TODO.md for the evidence that more steps do not help.
|
||||
|
||||
### Cost, and the licence problem
|
||||
|
||||
608 MB wheel, 2.6x `lg` and 45x `sm`. 187 words/s on the laptop CPU against `sm`'s 5,484
|
||||
(§9), so this tier needs a GPU in production rather than merely benefiting from one.
|
||||
|
||||
ParsBERT's model card states no licence. §3.4 picked `HooshvareLab/roberta-fa-zwnj-base`
|
||||
(Apache-2.0) for exactly this reason, and the published wheel therefore embeds weights whose
|
||||
redistribution terms are unknown. `scripts/finalize_pipeline.py` reads the encoder name out
|
||||
of the trained config and writes a redistribution warning into `meta.json` when the encoder
|
||||
has no licence, so the artifact carries the caveat. Retraining on the Apache-2.0 encoder is a
|
||||
one-line change to `name` in the config.
|
||||
|
||||
## 9. Throughput
|
||||
|
||||
Measured with `scripts/benchmark_throughput.py`, which times `nlp.pipe` and nothing else.
|
||||
The `words/s` printed by `spacy benchmark accuracy` runs the Scorer's per-token alignment
|
||||
inside the timed region, which is why the §7 numbers disagree with these and why one of them
|
||||
was impossible.
|
||||
|
||||
Median of repeated passes over the 146-document PerDT test split (23,825 tokens), batch 32,
|
||||
warmup discarded. Raw records in `metrics/throughput-*.json`.
|
||||
|
||||
| Tier | CPU, i5-7200U | GPU, GeForce 940MX | CPU, Xeon @ 2.00GHz | GPU, Tesla T4 |
|
||||
| --- | ---: | ---: | ---: | ---: |
|
||||
| `sm` | 5,484 | 10,235 | | |
|
||||
| `md` | 5,408 | 9,058 | | |
|
||||
| `lg` | 4,715 | 9,215 | | |
|
||||
| `trf` | 187 | 1,106 | 336 | 8,320 |
|
||||
|
||||
The CPU tiers sit within about 15% of each other, less than their vector-table sizes suggest,
|
||||
so the tok2vec lookup is not the bottleneck; the parser and lemmatizer are. Run-to-run spread
|
||||
on the laptop is roughly 10% either way with thermal state, and a background rsync halved
|
||||
every number, so treat small differences as noise.
|
||||
|
||||
`trf` is 29x slower than `sm` on the same CPU. The T4 and Xeon columns come from the same Colab
|
||||
VM, giving a clean 25x GPU speedup for the transformer.
|
||||
|
||||
`trf` on the 940MX needs a `cu126` torch build. sm_50 kernels were dropped from the `cu128` and
|
||||
`cu129` wheels at torch 2.8, which is what `pip install torch` resolves to. `.venv-trf-gpu` pins
|
||||
`torch==2.7.1+cu126`, separate from `.venv` because torch's pinned `nvidia-*` wheels downgrade
|
||||
the CUDA libraries cupy uses there from 12.9 to 12.6. Batch 32 fits in 2 GB.
|
||||
|
|
|
|||
409
project.yml
409
project.yml
|
|
@ -29,6 +29,22 @@ vars:
|
|||
# -1 = CPU. A GTX 940MX (2 GB) is not worth the transfer overhead for an sm pipeline.
|
||||
gpu: -1
|
||||
n_sents: 10
|
||||
# md tier. Same architecture as sm plus the fa_floret static vector table.
|
||||
dep_md_package_name: "dep_news_md"
|
||||
core_md_package_name: "core_news_md"
|
||||
floret_wheel: "fa_floret-0.1.0-py3-none-any-400k-documents.whl"
|
||||
vectors_dir: "assets/vectors/fa_floret_400k"
|
||||
# lg tier: same architecture as sm/md, larger floret table (200k rows x 300d, trained on
|
||||
# the full Persian Wikipedia dump for 5 epochs, vs md's 50k rows / 400k documents).
|
||||
ent_lg_package_name: "ent_news_lg"
|
||||
dep_lg_package_name: "dep_news_lg"
|
||||
core_lg_package_name: "core_news_lg"
|
||||
floret_lg_wheel: "fa_floret-0.1.0-py3-none-any-full-wiki-200k-5epoch.whl"
|
||||
vectors_lg_dir: "assets/vectors/fa_floret_lg"
|
||||
# trf tier: one fine-tuned ParsBERT shared by every component. Needs a real GPU; the
|
||||
# 940MX cannot fine-tune a 162M-parameter encoder, so `gpu_trf` is set for a rented card.
|
||||
core_trf_package_name: "core_news_trf"
|
||||
gpu_trf: 0
|
||||
|
||||
directories:
|
||||
- "assets"
|
||||
|
|
@ -88,6 +104,42 @@ workflows:
|
|||
- finalize-ent
|
||||
- evaluate-ent
|
||||
- package-ent
|
||||
# The lg tier: same corpus and architecture as sm/md, with a bigger floret table (200k
|
||||
# rows, full Persian Wikipedia, 5 epochs) than md's (50k rows, 400k documents).
|
||||
lg:
|
||||
- vectors-lg
|
||||
- train-dep-lg
|
||||
- train-ner-lg
|
||||
- finalize-dep-lg
|
||||
- assemble-core-lg
|
||||
- evaluate-lg
|
||||
- finalize-meta-lg
|
||||
- compare-lg
|
||||
- package-lg
|
||||
- smoke-lg
|
||||
# The md tier: same corpus and architecture, plus the fa_floret static vectors.
|
||||
md:
|
||||
- vectors-md
|
||||
- train-dep-md
|
||||
- train-ner-md
|
||||
- finalize-dep-md
|
||||
- assemble-core-md
|
||||
- evaluate-md
|
||||
- finalize-meta-md
|
||||
- compare-md
|
||||
- package-md
|
||||
- smoke-md
|
||||
# The trf tier: one fine-tuned ParsBERT shared by every component, including ner, so it
|
||||
# trains against a single joint corpus instead of the sm/md/lg dep+ner split. GPU only.
|
||||
trf:
|
||||
- merge-joint
|
||||
- debug-data-trf
|
||||
- train-trf
|
||||
- finalize-trf
|
||||
- evaluate-trf
|
||||
- finalize-meta-trf
|
||||
- package-trf
|
||||
- smoke-trf
|
||||
|
||||
commands:
|
||||
- name: "inspect"
|
||||
|
|
@ -297,6 +349,363 @@ commands:
|
|||
outputs:
|
||||
- "packages/${vars.lang}_${vars.ent_package_name}-${vars.package_version}"
|
||||
|
||||
# ---------------------------------------------------------------- lg tier (ner only)
|
||||
|
||||
- name: "vectors-lg"
|
||||
help: >
|
||||
Unpack the lg-tier fa_floret wheel into a plain spaCy model directory. 200k rows x
|
||||
300d in floret mode, trained on the full Persian Wikipedia dump for 5 epochs, vs
|
||||
vectors-md's 50k rows / 400k documents.
|
||||
script:
|
||||
- "python scripts/unpack_vectors.py ${vars.floret_lg_wheel} ${vars.vectors_lg_dir}"
|
||||
deps:
|
||||
- "${vars.floret_lg_wheel}"
|
||||
- "scripts/unpack_vectors.py"
|
||||
outputs:
|
||||
- "${vars.vectors_lg_dir}"
|
||||
|
||||
- name: "train-dep-lg"
|
||||
help: "Train the dep pipeline with the lg-tier static floret vectors"
|
||||
script:
|
||||
- "python -m spacy train configs/fa_dep_news_lg.cfg --output training/dep-lg --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy --paths.vectors ${vars.vectors_lg_dir} --gpu-id ${vars.gpu}"
|
||||
deps:
|
||||
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
||||
- "corpus/merged/${vars.treebank}-ud-dev.spacy"
|
||||
- "configs/fa_dep_news_lg.cfg"
|
||||
- "${vars.vectors_lg_dir}"
|
||||
outputs:
|
||||
- "training/dep-lg/model-best"
|
||||
|
||||
- name: "train-ner-lg"
|
||||
help: "Train the NER component with the lg-tier static floret vectors"
|
||||
script:
|
||||
- "python -m spacy train configs/fa_ner_lg.cfg --output training/perdt-ner-lg --paths.train corpus/perdt-ner/train.spacy --paths.dev corpus/perdt-ner/dev.spacy --paths.vectors ${vars.vectors_lg_dir} --gpu-id ${vars.gpu}"
|
||||
deps:
|
||||
- "corpus/perdt-ner/train.spacy"
|
||||
- "corpus/perdt-ner/dev.spacy"
|
||||
- "configs/fa_ner_lg.cfg"
|
||||
- "${vars.vectors_lg_dir}"
|
||||
outputs:
|
||||
- "training/perdt-ner-lg/model-best"
|
||||
|
||||
- name: "finalize-ent-lg"
|
||||
help: "Write fa_ent_news_lg metadata onto the trained lg model"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/perdt-ner-lg/model-best training/fa_ent_news_lg --variant ent --size lg --version ${vars.package_version}"
|
||||
deps:
|
||||
- "training/perdt-ner-lg/model-best"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
outputs:
|
||||
- "training/fa_ent_news_lg"
|
||||
|
||||
- name: "finalize-dep-lg"
|
||||
help: "Write fa_dep_news_lg metadata onto the trained lg model"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/dep-lg/model-best training/fa_dep_news_lg --variant dep --size lg --version ${vars.package_version}"
|
||||
deps:
|
||||
- "training/dep-lg/model-best"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
outputs:
|
||||
- "training/fa_dep_news_lg"
|
||||
|
||||
- name: "assemble-core-lg"
|
||||
help: "Source the lg ner into the lg dep pipeline to produce fa_core_news_lg"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/dep-lg/model-best training/fa_core_news_lg --variant core --size lg --version ${vars.package_version} --add-ner training/perdt-ner-lg/model-best"
|
||||
deps:
|
||||
- "training/dep-lg/model-best"
|
||||
- "training/perdt-ner-lg/model-best"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
outputs:
|
||||
- "training/fa_core_news_lg"
|
||||
|
||||
- name: "evaluate-ent-lg"
|
||||
help: "Score fa_ent_news_lg on the held-out PerDT NER test split"
|
||||
script:
|
||||
- "python -m spacy benchmark accuracy training/fa_ent_news_lg corpus/perdt-ner/test.spacy --output metrics/lg-perdt-ner-test.json --gpu-id ${vars.gpu}"
|
||||
- "python scripts/finalize_pipeline.py training/perdt-ner-lg/model-best training/fa_ent_news_lg --variant ent --size lg --version ${vars.package_version} --ner-metrics metrics/lg-perdt-ner-test.json"
|
||||
deps:
|
||||
- "training/fa_ent_news_lg"
|
||||
- "corpus/perdt-ner/test.spacy"
|
||||
outputs:
|
||||
- "metrics/lg-perdt-ner-test.json"
|
||||
|
||||
- name: "evaluate-lg"
|
||||
help: "Score both lg packages (dep, core) on the held-out test splits"
|
||||
script:
|
||||
- "python -m spacy benchmark accuracy training/fa_dep_news_lg corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/lg-ud-test.json --gpu-id ${vars.gpu}"
|
||||
- "python -m spacy benchmark accuracy training/fa_core_news_lg corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/lg-core-ud-test.json --gpu-id ${vars.gpu}"
|
||||
- "python -m spacy benchmark accuracy training/fa_core_news_lg corpus/perdt-ner/test.spacy --output metrics/lg-core-perdt-ner-test.json --gpu-id ${vars.gpu}"
|
||||
deps:
|
||||
- "training/fa_dep_news_lg"
|
||||
- "training/fa_core_news_lg"
|
||||
outputs:
|
||||
- "metrics/lg-ud-test.json"
|
||||
- "metrics/lg-core-ud-test.json"
|
||||
- "metrics/lg-core-perdt-ner-test.json"
|
||||
|
||||
- name: "finalize-meta-lg"
|
||||
help: "Fold the lg test scores into both lg meta.json files"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/dep-lg/model-best training/fa_dep_news_lg --variant dep --size lg --version ${vars.package_version} --ud-metrics metrics/lg-ud-test.json"
|
||||
- "python scripts/finalize_pipeline.py training/dep-lg/model-best training/fa_core_news_lg --variant core --size lg --version ${vars.package_version} --add-ner training/perdt-ner-lg/model-best --ud-metrics metrics/lg-core-ud-test.json --ner-metrics metrics/lg-core-perdt-ner-test.json"
|
||||
deps:
|
||||
- "metrics/lg-ud-test.json"
|
||||
- "metrics/lg-core-perdt-ner-test.json"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
|
||||
- name: "compare-lg"
|
||||
help: "Table the sm vs md vs lg deltas from the metrics/ JSON reports"
|
||||
script:
|
||||
- "python scripts/compare_tiers.py"
|
||||
deps:
|
||||
- "metrics/ud-test.json"
|
||||
- "metrics/md-ud-test.json"
|
||||
- "metrics/lg-ud-test.json"
|
||||
- "metrics/perdt-ner-test.json"
|
||||
- "metrics/md-perdt-ner-test.json"
|
||||
- "metrics/lg-perdt-ner-test.json"
|
||||
- "scripts/compare_tiers.py"
|
||||
|
||||
- name: "package-ent-lg"
|
||||
help: "Build the installable fa_ent_news_lg wheel + sdist"
|
||||
script:
|
||||
- "python -m spacy package training/fa_ent_news_lg packages --name ${vars.ent_lg_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||
deps:
|
||||
- "training/fa_ent_news_lg"
|
||||
outputs:
|
||||
- "packages/${vars.lang}_${vars.ent_lg_package_name}-${vars.package_version}"
|
||||
|
||||
- name: "package-lg"
|
||||
help: "Build installable wheels + sdists for both lg packages"
|
||||
script:
|
||||
- "python -m spacy package training/fa_dep_news_lg packages --name ${vars.dep_lg_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||
- "python -m spacy package training/fa_core_news_lg packages --name ${vars.core_lg_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||
deps:
|
||||
- "training/fa_dep_news_lg"
|
||||
- "training/fa_core_news_lg"
|
||||
outputs:
|
||||
- "packages/${vars.lang}_${vars.dep_lg_package_name}-${vars.package_version}"
|
||||
- "packages/${vars.lang}_${vars.core_lg_package_name}-${vars.package_version}"
|
||||
|
||||
- name: "smoke-ent-lg"
|
||||
help: "Load fa_ent_news_lg and run it over real Persian text"
|
||||
script:
|
||||
- "python scripts/smoke_test.py training/fa_ent_news_lg"
|
||||
deps:
|
||||
- "training/fa_ent_news_lg"
|
||||
|
||||
- name: "smoke-lg"
|
||||
help: "Load both lg pipelines and run them over real Persian text"
|
||||
script:
|
||||
- "python scripts/smoke_test.py training/fa_dep_news_lg"
|
||||
- "python scripts/smoke_test.py training/fa_core_news_lg"
|
||||
deps:
|
||||
- "training/fa_dep_news_lg"
|
||||
- "training/fa_core_news_lg"
|
||||
|
||||
# ---------------------------------------------------------------- md tier
|
||||
|
||||
- name: "vectors-md"
|
||||
help: >
|
||||
Unpack the fa_floret wheel into a plain spaCy model directory that
|
||||
`--paths.vectors` can point at. The wheel is a vectors-only pipeline
|
||||
(empty `pipeline: []`), 50k rows x 300d in floret mode, trained on 400k
|
||||
Persian documents, so no `spacy init vectors` step is needed.
|
||||
script:
|
||||
- "python scripts/unpack_vectors.py ${vars.floret_wheel} ${vars.vectors_dir}"
|
||||
deps:
|
||||
- "${vars.floret_wheel}"
|
||||
- "scripts/unpack_vectors.py"
|
||||
outputs:
|
||||
- "${vars.vectors_dir}"
|
||||
|
||||
- name: "train-dep-md"
|
||||
help: "Train the dep pipeline with static floret vectors"
|
||||
script:
|
||||
- "python -m spacy train configs/fa_dep_news_md.cfg --output training/dep-md --paths.train corpus/merged/${vars.treebank}-ud-train.spacy --paths.dev corpus/merged/${vars.treebank}-ud-dev.spacy --paths.vectors ${vars.vectors_dir} --gpu-id ${vars.gpu}"
|
||||
deps:
|
||||
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
||||
- "corpus/merged/${vars.treebank}-ud-dev.spacy"
|
||||
- "configs/fa_dep_news_md.cfg"
|
||||
- "${vars.vectors_dir}"
|
||||
outputs:
|
||||
- "training/dep-md/model-best"
|
||||
|
||||
- name: "train-ner-md"
|
||||
help: "Train the NER component with static floret vectors"
|
||||
script:
|
||||
- "python -m spacy train configs/fa_ner_md.cfg --output training/perdt-ner-md --paths.train corpus/perdt-ner/train.spacy --paths.dev corpus/perdt-ner/dev.spacy --paths.vectors ${vars.vectors_dir} --gpu-id ${vars.gpu}"
|
||||
deps:
|
||||
- "corpus/perdt-ner/train.spacy"
|
||||
- "corpus/perdt-ner/dev.spacy"
|
||||
- "configs/fa_ner_md.cfg"
|
||||
- "${vars.vectors_dir}"
|
||||
outputs:
|
||||
- "training/perdt-ner-md/model-best"
|
||||
|
||||
- name: "finalize-dep-md"
|
||||
help: "Write fa_dep_news_md metadata onto the trained md model"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_dep_news_md --variant dep --size md --version ${vars.package_version}"
|
||||
deps:
|
||||
- "training/dep-md/model-best"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
outputs:
|
||||
- "training/fa_dep_news_md"
|
||||
|
||||
- name: "assemble-core-md"
|
||||
help: "Source the md ner into the md dep pipeline to produce fa_core_news_md"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_core_news_md --variant core --size md --version ${vars.package_version} --add-ner training/perdt-ner-md/model-best"
|
||||
deps:
|
||||
- "training/dep-md/model-best"
|
||||
- "training/perdt-ner-md/model-best"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
outputs:
|
||||
- "training/fa_core_news_md"
|
||||
|
||||
- name: "evaluate-md"
|
||||
help: "Score both md packages on the held-out test splits"
|
||||
script:
|
||||
- "python -m spacy benchmark accuracy training/fa_dep_news_md corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/md-ud-test.json --gpu-id ${vars.gpu}"
|
||||
- "python -m spacy benchmark accuracy training/fa_core_news_md corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/md-core-ud-test.json --gpu-id ${vars.gpu}"
|
||||
- "python -m spacy benchmark accuracy training/fa_core_news_md corpus/perdt-ner/test.spacy --output metrics/md-perdt-ner-test.json --gpu-id ${vars.gpu}"
|
||||
deps:
|
||||
- "training/fa_dep_news_md"
|
||||
- "training/fa_core_news_md"
|
||||
outputs:
|
||||
- "metrics/md-ud-test.json"
|
||||
- "metrics/md-core-ud-test.json"
|
||||
- "metrics/md-perdt-ner-test.json"
|
||||
|
||||
- name: "finalize-meta-md"
|
||||
help: "Fold the md test scores into both md meta.json files"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_dep_news_md --variant dep --size md --version ${vars.package_version} --ud-metrics metrics/md-ud-test.json"
|
||||
- "python scripts/finalize_pipeline.py training/dep-md/model-best training/fa_core_news_md --variant core --size md --version ${vars.package_version} --add-ner training/perdt-ner-md/model-best --ud-metrics metrics/md-core-ud-test.json --ner-metrics metrics/md-perdt-ner-test.json"
|
||||
deps:
|
||||
- "metrics/md-ud-test.json"
|
||||
- "metrics/md-perdt-ner-test.json"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
|
||||
- name: "compare-md"
|
||||
help: "Table the sm vs md deltas from the metrics/ JSON reports"
|
||||
script:
|
||||
- "python scripts/compare_tiers.py"
|
||||
deps:
|
||||
- "metrics/md-ud-test.json"
|
||||
- "metrics/md-perdt-ner-test.json"
|
||||
- "scripts/compare_tiers.py"
|
||||
|
||||
- name: "package-md"
|
||||
help: "Build installable wheels + sdists for both md packages"
|
||||
script:
|
||||
- "python -m spacy package training/fa_dep_news_md packages --name ${vars.dep_md_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||
- "python -m spacy package training/fa_core_news_md packages --name ${vars.core_md_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||
deps:
|
||||
- "training/fa_dep_news_md"
|
||||
- "training/fa_core_news_md"
|
||||
outputs:
|
||||
- "packages/${vars.lang}_${vars.dep_md_package_name}-${vars.package_version}"
|
||||
- "packages/${vars.lang}_${vars.core_md_package_name}-${vars.package_version}"
|
||||
|
||||
- name: "smoke-md"
|
||||
help: "Load both md pipelines and run them over real Persian text"
|
||||
script:
|
||||
- "python scripts/smoke_test.py training/fa_dep_news_md"
|
||||
- "python scripts/smoke_test.py training/fa_core_news_md"
|
||||
deps:
|
||||
- "training/fa_dep_news_md"
|
||||
- "training/fa_core_news_md"
|
||||
|
||||
# ---------------------------------------------------------------- trf tier
|
||||
|
||||
- name: "merge-joint"
|
||||
help: >
|
||||
Fuse the UD layer and the transferred NER layer onto one set of Docs. The trf tier
|
||||
shares a single transformer across every component, so it needs one corpus carrying
|
||||
both annotation layers; the two DocBins are token-for-token identical by construction
|
||||
and the script asserts it.
|
||||
script:
|
||||
- "python scripts/merge_joint_corpus.py --ud-dir corpus/merged --ner-dir corpus/perdt-ner --out corpus/joint"
|
||||
deps:
|
||||
- "corpus/merged/${vars.treebank}-ud-train.spacy"
|
||||
- "corpus/perdt-ner/train.spacy"
|
||||
- "scripts/merge_joint_corpus.py"
|
||||
outputs:
|
||||
- "corpus/joint/train.spacy"
|
||||
- "corpus/joint/dev.spacy"
|
||||
- "corpus/joint/test.spacy"
|
||||
|
||||
- name: "debug-data-trf"
|
||||
help: "Validate the joint corpus against the trf config before renting GPU time"
|
||||
script:
|
||||
- "python -m spacy debug data configs/fa_core_news_trf.cfg --paths.train corpus/joint/train.spacy --paths.dev corpus/joint/dev.spacy"
|
||||
deps:
|
||||
- "corpus/joint/train.spacy"
|
||||
- "configs/fa_core_news_trf.cfg"
|
||||
|
||||
- name: "train-trf"
|
||||
help: "Fine-tune ParsBERT with tagger + morphologizer + lemmatizer + parser + ner listening"
|
||||
script:
|
||||
- "python -m spacy train configs/fa_core_news_trf.cfg --output training/core-trf --paths.train corpus/joint/train.spacy --paths.dev corpus/joint/dev.spacy --gpu-id ${vars.gpu_trf}"
|
||||
deps:
|
||||
- "corpus/joint/train.spacy"
|
||||
- "corpus/joint/dev.spacy"
|
||||
- "configs/fa_core_news_trf.cfg"
|
||||
outputs:
|
||||
- "training/core-trf/model-best"
|
||||
|
||||
- name: "finalize-trf"
|
||||
help: "Write fa_core_news_trf metadata onto the trained model"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/core-trf/model-best training/fa_core_news_trf --variant core --size trf --version ${vars.package_version}"
|
||||
deps:
|
||||
- "training/core-trf/model-best"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
outputs:
|
||||
- "training/fa_core_news_trf"
|
||||
|
||||
- name: "evaluate-trf"
|
||||
help: "Score fa_core_news_trf on the held-out UD and NER test splits"
|
||||
script:
|
||||
- "python -m spacy benchmark accuracy training/fa_core_news_trf corpus/merged/${vars.treebank}-ud-test.spacy --output metrics/trf-core-ud-test.json --gpu-id ${vars.gpu_trf}"
|
||||
- "python -m spacy benchmark accuracy training/fa_core_news_trf corpus/perdt-ner/test.spacy --output metrics/trf-perdt-ner-test.json --gpu-id ${vars.gpu_trf}"
|
||||
deps:
|
||||
- "training/fa_core_news_trf"
|
||||
- "corpus/merged/${vars.treebank}-ud-test.spacy"
|
||||
- "corpus/perdt-ner/test.spacy"
|
||||
outputs:
|
||||
- "metrics/trf-core-ud-test.json"
|
||||
- "metrics/trf-perdt-ner-test.json"
|
||||
|
||||
- name: "finalize-meta-trf"
|
||||
help: "Fold the trf test scores into meta.json"
|
||||
script:
|
||||
- "python scripts/finalize_pipeline.py training/core-trf/model-best training/fa_core_news_trf --variant core --size trf --version ${vars.package_version} --ud-metrics metrics/trf-core-ud-test.json --ner-metrics metrics/trf-perdt-ner-test.json"
|
||||
deps:
|
||||
- "metrics/trf-core-ud-test.json"
|
||||
- "metrics/trf-perdt-ner-test.json"
|
||||
- "scripts/finalize_pipeline.py"
|
||||
|
||||
- name: "package-trf"
|
||||
help: "Build the installable fa_core_news_trf wheel + sdist"
|
||||
script:
|
||||
- "python -m spacy package training/fa_core_news_trf packages --name ${vars.core_trf_package_name} --version ${vars.package_version} --build sdist,wheel --force"
|
||||
deps:
|
||||
- "training/fa_core_news_trf"
|
||||
outputs:
|
||||
- "packages/${vars.lang}_${vars.core_trf_package_name}-${vars.package_version}"
|
||||
|
||||
- name: "smoke-trf"
|
||||
help: "Load fa_core_news_trf and run it over real Persian text"
|
||||
script:
|
||||
- "python scripts/smoke_test.py training/fa_core_news_trf"
|
||||
deps:
|
||||
- "training/fa_core_news_trf"
|
||||
|
||||
|
||||
- name: "clean"
|
||||
help: "Drop corpora, training runs and metrics (keeps downloaded assets)"
|
||||
script:
|
||||
|
|
|
|||
|
|
@ -0,0 +1,112 @@
|
|||
"""Measure inference throughput (words/second) for a pipeline, on CPU or GPU.
|
||||
|
||||
`spacy benchmark accuracy` prints a speed number, but it is scoring-contaminated: the
|
||||
Scorer's per-token alignment and per-type bookkeeping run inside the timed region, which
|
||||
matters a lot for the cheap CPU tiers and understates them. This times `nlp.pipe` only.
|
||||
|
||||
Reported figure is the median of `--runs` passes over the same texts, after a discarded
|
||||
warmup pass. Median rather than mean because the first CUDA kernel launches, cuBLAS
|
||||
autotuning and any page-cache miss produce outliers that a mean would smear into the result.
|
||||
|
||||
Batch size matters far more for the trf tier than the CPU tiers (a transformer amortizes a
|
||||
GEMM over the batch; a hash-embed tok2vec barely cares), so it is a parameter and gets
|
||||
recorded in the output rather than being left implicit.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import platform
|
||||
import statistics
|
||||
import subprocess
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
import spacy
|
||||
from spacy.tokens import DocBin
|
||||
|
||||
|
||||
def cpu_model():
|
||||
try:
|
||||
for line in Path("/proc/cpuinfo").read_text().splitlines():
|
||||
if line.startswith("model name"):
|
||||
return line.split(":", 1)[1].strip()
|
||||
except OSError:
|
||||
pass
|
||||
return platform.processor() or "unknown"
|
||||
|
||||
|
||||
def gpu_model():
|
||||
try:
|
||||
out = subprocess.run(
|
||||
["nvidia-smi", "--query-gpu=name,memory.total", "--format=csv,noheader"],
|
||||
capture_output=True, text=True, timeout=30,
|
||||
)
|
||||
if out.returncode == 0:
|
||||
return out.stdout.strip().splitlines()[0].strip()
|
||||
except (OSError, subprocess.SubprocessError):
|
||||
pass
|
||||
return "unknown"
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("model", help="installed package name or path to a pipeline")
|
||||
ap.add_argument("--corpus", default="corpus/merged/fa_perdt-ud-test.spacy",
|
||||
help="DocBin whose raw texts are used as input")
|
||||
ap.add_argument("--gpu-id", type=int, default=-1, help="-1 for CPU")
|
||||
ap.add_argument("--batch-size", type=int, default=32)
|
||||
ap.add_argument("--runs", type=int, default=3)
|
||||
ap.add_argument("--limit", type=int, default=0, help="cap number of docs (0 = all)")
|
||||
ap.add_argument("--output", default=None, help="write a JSON record here")
|
||||
args = ap.parse_args()
|
||||
|
||||
if args.gpu_id >= 0:
|
||||
# require_gpu, not prefer_gpu: a silent fall back to CPU would be reported as a GPU
|
||||
# number, which is exactly the measurement error this script exists to avoid.
|
||||
spacy.require_gpu(args.gpu_id)
|
||||
device = f"gpu:{args.gpu_id} ({gpu_model()})"
|
||||
else:
|
||||
device = f"cpu ({cpu_model()})"
|
||||
|
||||
nlp = spacy.load(args.model)
|
||||
vocab_docs = list(DocBin().from_disk(args.corpus).get_docs(spacy.blank("fa").vocab))
|
||||
if args.limit:
|
||||
vocab_docs = vocab_docs[:args.limit]
|
||||
texts = [d.text for d in vocab_docs]
|
||||
n_words = sum(len(d) for d in vocab_docs)
|
||||
|
||||
# Warmup: first pass pays for lazy CUDA context creation, cuBLAS handles and any
|
||||
# transformer weight transfer. Timing it would misattribute setup cost to throughput.
|
||||
for _ in nlp.pipe(texts[:args.batch_size], batch_size=args.batch_size):
|
||||
pass
|
||||
|
||||
wps = []
|
||||
for _ in range(args.runs):
|
||||
t0 = time.perf_counter()
|
||||
for _ in nlp.pipe(texts, batch_size=args.batch_size):
|
||||
pass
|
||||
elapsed = time.perf_counter() - t0
|
||||
wps.append(n_words / elapsed)
|
||||
|
||||
median = statistics.median(wps)
|
||||
record = {
|
||||
"model": args.model,
|
||||
"pipeline": list(nlp.pipe_names),
|
||||
"device": device,
|
||||
"batch_size": args.batch_size,
|
||||
"docs": len(texts),
|
||||
"words": n_words,
|
||||
"runs": [round(w, 1) for w in wps],
|
||||
"wps_median": round(median, 1),
|
||||
"spacy_version": spacy.__version__,
|
||||
}
|
||||
print(json.dumps(record, indent=2, ensure_ascii=False))
|
||||
if args.output:
|
||||
p = Path(args.output)
|
||||
p.parent.mkdir(parents=True, exist_ok=True)
|
||||
p.write_text(json.dumps(record, indent=2, ensure_ascii=False) + "\n")
|
||||
print(f"wrote {p}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
|
@ -0,0 +1,166 @@
|
|||
"""Table the sm vs md vs lg test-set deltas.
|
||||
|
||||
All tiers are trained from the same corpus, the same seed and the same architecture; the
|
||||
only difference is the static vector table (none for sm, fa_floret 50k rows for md, fa_floret
|
||||
200k rows for lg) via `include_static_vectors`. So the delta printed here is attributable to
|
||||
the vector table and nothing else.
|
||||
|
||||
Reads the `spacy benchmark accuracy` reports written by the `evaluate-*` targets. Missing
|
||||
files are reported rather than fatal, so this is runnable mid-build.
|
||||
|
||||
Usage:
|
||||
python scripts/compare_tiers.py [--metrics-dir metrics]
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
# (label, {tier_label: report_filename})
|
||||
GROUPS = [
|
||||
(
|
||||
"dep pipeline, UD test",
|
||||
{"sm": "ud-test.json", "md": "md-ud-test.json", "lg": "lg-ud-test.json"},
|
||||
),
|
||||
(
|
||||
"core pipeline, UD test",
|
||||
{
|
||||
"sm": "core-ud-test.json",
|
||||
"md": "md-core-ud-test.json",
|
||||
"lg": "lg-core-ud-test.json",
|
||||
},
|
||||
),
|
||||
(
|
||||
"ent NER test",
|
||||
{
|
||||
"sm": "perdt-ner-test.json",
|
||||
"md": "md-perdt-ner-test.json",
|
||||
"lg": "lg-perdt-ner-test.json",
|
||||
},
|
||||
),
|
||||
]
|
||||
|
||||
SCALARS = [
|
||||
("tag_acc", "TAG_ACC"),
|
||||
("pos_acc", "POS_ACC"),
|
||||
("morph_acc", "MORPH_ACC"),
|
||||
("lemma_acc", "LEMMA_ACC"),
|
||||
("dep_uas", "DEP_UAS"),
|
||||
("dep_las", "DEP_LAS"),
|
||||
("sents_f", "SENTS_F"),
|
||||
("ents_p", "ENTS_P"),
|
||||
("ents_r", "ENTS_R"),
|
||||
("ents_f", "ENTS_F"),
|
||||
]
|
||||
|
||||
|
||||
def load(path):
|
||||
return json.loads(path.read_text(encoding="utf8")) if path.exists() else None
|
||||
|
||||
|
||||
def table(title, tiers, rows):
|
||||
"""tiers: list of (label, data-dict-or-None), first tier is the baseline for deltas."""
|
||||
labels = [label for label, _ in tiers]
|
||||
base_label, base = tiers[0]
|
||||
print(f"\n## {title}\n")
|
||||
header = " | ".join(f"{label:>7}" for label in labels)
|
||||
delta_header = " | ".join(f"{'d(' + label + ')':>9}" for label, _ in tiers[1:])
|
||||
print(f"| {'metric':<12} | {header} | {delta_header} |")
|
||||
sep = " | ".join("-" * 7 for _ in labels)
|
||||
delta_sep = " | ".join("-" * 9 for _ in tiers[1:])
|
||||
print(f"| {'-' * 12} | {sep} | {delta_sep} |")
|
||||
for key, label in rows:
|
||||
values = [d.get(key) if d is not None else None for _, d in tiers]
|
||||
if all(v is None for v in values):
|
||||
continue
|
||||
# The NER report scores tag_acc 0.0 because its corpus has no gold tags.
|
||||
if all(v == 0.0 for v in values):
|
||||
continue
|
||||
cells = [f"{v * 100:.2f}" if isinstance(v, float) else "-" for v in values]
|
||||
deltas = []
|
||||
for v in values[1:]:
|
||||
a, b = values[0], v
|
||||
deltas.append(
|
||||
f"{(b - a) * 100:+.2f}" if isinstance(a, float) and isinstance(b, float) else "-"
|
||||
)
|
||||
row = " | ".join(f"{c:>7}" for c in cells)
|
||||
drow = " | ".join(f"{d:>9}" for d in deltas)
|
||||
print(f"| {label:<12} | {row} | {drow} |")
|
||||
speeds = [d.get("speed") if d is not None else None for _, d in tiers]
|
||||
if isinstance(speeds[0], float):
|
||||
cells = [f"{s:.0f}" if isinstance(s, float) else "-" for s in speeds]
|
||||
deltas = [
|
||||
f"{s / speeds[0] - 1:+.1%}" if isinstance(s, float) else "-" for s in speeds[1:]
|
||||
]
|
||||
row = " | ".join(f"{c:>7}" for c in cells)
|
||||
drow = " | ".join(f"{d:>9}" for d in deltas)
|
||||
print(f"| {'words/s':<12} | {row} | {drow} |")
|
||||
|
||||
|
||||
def per_type(title, tiers):
|
||||
per_types = [(label, (d or {}).get("ents_per_type")) for label, d in tiers]
|
||||
if not any(pt for _, pt in per_types):
|
||||
return
|
||||
labels = [label for label, _ in tiers]
|
||||
|
||||
def pct(v):
|
||||
return f"{v * 100:.2f}" if v is not None else "-"
|
||||
|
||||
all_labels = set()
|
||||
for _, pt in per_types:
|
||||
if pt:
|
||||
all_labels |= set(pt)
|
||||
|
||||
print(f"\n### {title}, per label\n")
|
||||
header = " | ".join(f"{label + ' F':>7}" for label in labels)
|
||||
delta_header = " | ".join(f"{'d(' + label + ')':>9}" for label in labels[1:])
|
||||
print(f"| {'label':<6} | {header} | {delta_header} |")
|
||||
sep = " | ".join("-" * 7 for _ in labels)
|
||||
delta_sep = " | ".join("-" * 9 for _ in labels[1:])
|
||||
print(f"| {'-' * 6} | {sep} | {delta_sep} |")
|
||||
|
||||
def sort_key(entity_label):
|
||||
last_pt = per_types[-1][1] or {}
|
||||
return -last_pt.get(entity_label, {}).get("f", 0)
|
||||
|
||||
for entity_label in sorted(all_labels, key=sort_key):
|
||||
fs = [(pt or {}).get(entity_label, {}).get("f") for _, pt in per_types]
|
||||
cells = [pct(f) for f in fs]
|
||||
deltas = []
|
||||
for f in fs[1:]:
|
||||
a = fs[0]
|
||||
deltas.append(f"{(f - a) * 100:+.2f}" if a is not None and f is not None else "-")
|
||||
row = " | ".join(f"{c:>7}" for c in cells)
|
||||
drow = " | ".join(f"{d:>9}" for d in deltas)
|
||||
print(f"| {entity_label:<6} | {row} | {drow} |")
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--metrics-dir", type=Path, default=Path("metrics"))
|
||||
args = ap.parse_args()
|
||||
|
||||
print("# sm vs md vs lg (fa_floret static vectors)")
|
||||
print("\nSame corpus, same seed, same architecture per group. Only difference:")
|
||||
print("`include_static_vectors = false -> true`, and which floret table (md: 50k rows,")
|
||||
print("400k documents; lg: 200k rows, full Persian Wikipedia, 5 epochs).")
|
||||
|
||||
for title, reports in GROUPS:
|
||||
tiers = []
|
||||
missing = []
|
||||
for label, fname in reports.items():
|
||||
data = load(args.metrics_dir / fname)
|
||||
if data is None:
|
||||
missing.append(fname)
|
||||
tiers.append((label, data))
|
||||
if tiers[0][1] is None:
|
||||
print(f"\n## {title}\n\n (skipped, missing baseline {reports[list(reports)[0]]})")
|
||||
continue
|
||||
table(title, tiers, SCALARS)
|
||||
per_type(title, tiers)
|
||||
if missing:
|
||||
print(f"\n (missing: {', '.join(missing)})")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
|
@ -3,9 +3,13 @@
|
|||
Three variants, following spaCy's `[lang]_[type]_[genre]_[size]` naming
|
||||
(https://spacy.io/models#conventions):
|
||||
|
||||
dep -> fa_dep_news_sm tagger + morphologizer + trainable_lemmatizer + parser
|
||||
core -> fa_core_news_sm the above plus ner
|
||||
ent -> fa_ent_news_sm ner only
|
||||
dep -> fa_dep_news_<size> tagger + morphologizer + trainable_lemmatizer + parser
|
||||
core -> fa_core_news_<size> the above plus ner
|
||||
ent -> fa_ent_news_<size> ner only
|
||||
|
||||
`--size` fills the size slot: `sm` (hash embeddings only, the default) or `md` (the same
|
||||
architecture plus the fa_floret static vector table). It is metadata only; which vectors a
|
||||
model actually carries is decided at train time by `--paths.vectors`.
|
||||
|
||||
All three are built from UD_Persian-PerDT alone, including the NER, which comes from that
|
||||
treebank's own `not-to-release/Dadegan with NER tag/` layer. That is what makes `core`
|
||||
|
|
@ -56,6 +60,29 @@ LANG_DATA = {
|
|||
"author": "Explosion and spaCy contributors",
|
||||
"license": "MIT",
|
||||
}
|
||||
FLORET = {
|
||||
"name": "fa_floret static vectors (50k rows x 300d, floret mode, 400k Persian documents)",
|
||||
"url": PROJECT_URL,
|
||||
"author": "Kiyarash Fazeli",
|
||||
"license": "CC BY-SA 4.0",
|
||||
}
|
||||
FLORET_LG = {
|
||||
"name": "fa_floret static vectors (lg tier: 200k rows x 300d floret table trained on "
|
||||
"the full Persian Wikipedia dump, 5 epochs, via spacy-vectors-builder)",
|
||||
"url": "https://huggingface.co/Phazel/fa-floret-wiki-vectors",
|
||||
"author": "Kiyarash Fazeli",
|
||||
"license": "CC BY-SA 4.0",
|
||||
}
|
||||
# Whatever encoder the config actually names wins; hardcoding one would silently mislabel a
|
||||
# wheel the moment configs/fa_core_news_trf.cfg's `name` changes. Licences are recorded per
|
||||
# encoder because they differ sharply, and two of the Persian ones have none at all.
|
||||
ENCODER_LICENSES = {
|
||||
"HooshvareLab/roberta-fa-zwnj-base": ("Hooshvare Team", "Apache-2.0"),
|
||||
"HooshvareLab/bert-fa-zwnj-base": ("Hooshvare Team", "Apache-2.0"),
|
||||
"m3hrdadfi/albert-fa-base-v2": ("Mehrdad Farahani", "Apache-2.0"),
|
||||
"HooshvareLab/bert-base-parsbert-uncased": ("Hooshvare Team", "no licence stated on the model card"),
|
||||
"sbunlp/fabert": ("SBU NLP Lab", "no licence stated on the model card"),
|
||||
}
|
||||
|
||||
NER_NOTE = (
|
||||
"The ner component is trained on the NER layer shipped in UD_Persian-PerDT's "
|
||||
|
|
@ -78,6 +105,73 @@ CHUNK_NOTE = (
|
|||
"ClearNLP labels that do not exist in Universal Dependencies, see "
|
||||
"docs/upstream/fa-noun-chunks.md."
|
||||
)
|
||||
VECTORS_NOTE = (
|
||||
"This is the `md` tier: identical architecture to the `sm` pipeline plus static floret "
|
||||
"vectors (50,000 rows x 300 dimensions, minn=maxn=5, hash_count=2) trained on 400,000 "
|
||||
"Persian documents. floret hashes subwords into a fixed table, so there are no "
|
||||
"out-of-vocabulary tokens and `token.has_vector` is always True. That matters for "
|
||||
"Persian, where inconsistent ZWNJ (U+200C) usage splits one word across several surface "
|
||||
"forms (mi-ravad written joined, with ZWNJ, or with a space) that a classic word-vector "
|
||||
"table would miss."
|
||||
)
|
||||
|
||||
|
||||
def vectors_note_lg(nlp):
|
||||
"""Row/dim counts come from the trained model, not a hardcoded description, because the
|
||||
lg-tier floret table is still being iterated on (unlike md's fixed, shipped table)."""
|
||||
rows, dim = nlp.vocab.vectors.shape
|
||||
return (
|
||||
f"This is the `lg` tier: identical architecture to `sm`/`md` but a larger static "
|
||||
f"floret vector table ({rows:,} rows x {dim} dimensions, minn=maxn=5, hash_count=2) "
|
||||
f"trained on the full Persian Wikipedia dump for 5 epochs via spacy-vectors-builder. "
|
||||
f"Same zero-OOV rationale as `md` (see docs/MODELS.md): floret hashes subwords into "
|
||||
f"a fixed table, so `token.has_vector` is always True despite Persian's ZWNJ "
|
||||
f"(U+200C) inconsistency."
|
||||
)
|
||||
|
||||
|
||||
def encoder_name(nlp):
|
||||
"""Read the encoder out of the trained pipeline's own config."""
|
||||
try:
|
||||
return nlp.config["components"]["transformer"]["model"]["name"]
|
||||
except KeyError:
|
||||
raise SystemExit(
|
||||
"--size trf expects a pipeline with a `transformer` component whose model names "
|
||||
f"an encoder; got pipeline {list(nlp.pipe_names)}"
|
||||
)
|
||||
|
||||
|
||||
def transformer_source(nlp):
|
||||
name = encoder_name(nlp)
|
||||
author, license_ = ENCODER_LICENSES.get(name, ("unknown", "unknown, check the model card"))
|
||||
return {
|
||||
"name": name,
|
||||
"url": f"https://huggingface.co/{name}",
|
||||
"author": author,
|
||||
"license": license_,
|
||||
}
|
||||
|
||||
|
||||
def transformer_note(nlp):
|
||||
name = encoder_name(nlp)
|
||||
_, license_ = ENCODER_LICENSES.get(name, ("unknown", "unknown, check the model card"))
|
||||
note = (
|
||||
f"This is the `trf` tier: no static vectors. Contextual embeddings come from a "
|
||||
f"fine-tuned {name} ({license_}) via spacy-transformers, shared by every component "
|
||||
f"through a TransformerListener, so one encoder forward pass serves the tagger, "
|
||||
f"morphologizer, lemmatizer, parser and ner. Unlike the sm/md/lg tiers the ner is "
|
||||
f"trained jointly rather than sourced, because a shared encoder cannot be fine-tuned "
|
||||
f"twice and then merged. GPU is strongly recommended for both training and inference."
|
||||
)
|
||||
if "no licence" in license_ or license_.startswith("unknown"):
|
||||
note += (
|
||||
f" REDISTRIBUTION WARNING: {name} states no licence, so this wheel embeds weights "
|
||||
f"whose terms are unknown and must not be republished. Retrain against "
|
||||
f"HooshvareLab/roberta-fa-zwnj-base (Apache-2.0) for a publishable artifact."
|
||||
)
|
||||
return note
|
||||
|
||||
|
||||
# CC BY-SA 4.0 on the treebank propagates to anything derived from it.
|
||||
PERDT_LICENSE = "CC BY-SA 4.0"
|
||||
ATTRIBUTION = (
|
||||
|
|
@ -92,10 +186,10 @@ NER_KEYS = ("ents_p", "ents_r", "ents_f", "ents_per_type")
|
|||
|
||||
VARIANTS = {
|
||||
"dep": {
|
||||
"name": "dep_news_sm",
|
||||
"name": "dep_news_{size}",
|
||||
"description": (
|
||||
"Persian dependency pipeline optimized for CPU. Components: tok2vec, tagger, "
|
||||
"morphologizer, trainable_lemmatizer, parser. No NER, see fa_core_news_sm."
|
||||
"morphologizer, trainable_lemmatizer, parser. No NER, see fa_core_news_{size}."
|
||||
),
|
||||
"license": PERDT_LICENSE,
|
||||
"sources": [PERDT, LANG_DATA],
|
||||
|
|
@ -105,7 +199,7 @@ VARIANTS = {
|
|||
"require_msg": "a 'dep' pipeline must not contain an ner component",
|
||||
},
|
||||
"core": {
|
||||
"name": "core_news_sm",
|
||||
"name": "core_news_{size}",
|
||||
"description": (
|
||||
"Persian pipeline optimized for CPU. Components: tok2vec, tagger, morphologizer, "
|
||||
"trainable_lemmatizer, parser, ner. Entity labels: PER, LOC, ORG, DAT, MON, TIM, "
|
||||
|
|
@ -119,7 +213,7 @@ VARIANTS = {
|
|||
"require_msg": "a 'core' pipeline must contain both parser and ner",
|
||||
},
|
||||
"ent": {
|
||||
"name": "ent_news_sm",
|
||||
"name": "ent_news_{size}",
|
||||
"description": (
|
||||
"Persian named entity recognizer optimized for CPU, with its own internal "
|
||||
"tok2vec. Labels: PER, LOC, ORG, DAT, MON, TIM, PCT."
|
||||
|
|
@ -140,6 +234,10 @@ def main():
|
|||
ap.add_argument("output", help="destination directory")
|
||||
ap.add_argument("--variant", choices=sorted(VARIANTS), required=True)
|
||||
ap.add_argument("--version", default="3.8.0")
|
||||
ap.add_argument("--size", choices=("sm", "md", "lg", "trf"), default="sm",
|
||||
help="size slot in the package name. 'md'/'lg' additionally record the "
|
||||
"floret vector table as a source and append a vectors note; 'trf' "
|
||||
"records the transformer source and appends a transformer note.")
|
||||
ap.add_argument("--ud-metrics", default=None,
|
||||
help="benchmark accuracy JSON scored on the UD test split; supplies the "
|
||||
"tagger/morph/lemma/parser keys only")
|
||||
|
|
@ -152,7 +250,27 @@ def main():
|
|||
args = ap.parse_args()
|
||||
|
||||
spec = VARIANTS[args.variant]
|
||||
name = spec["name"].format(size=args.size)
|
||||
description = spec["description"].format(size=args.size)
|
||||
sources = list(spec["sources"])
|
||||
notes = spec["notes"]
|
||||
nlp = spacy.load(args.model)
|
||||
if args.size == "md":
|
||||
sources.append(FLORET)
|
||||
notes = " ".join([notes, VECTORS_NOTE])
|
||||
elif args.size == "lg":
|
||||
sources.append(FLORET_LG)
|
||||
notes = " ".join([notes, vectors_note_lg(nlp)])
|
||||
elif args.size == "trf":
|
||||
sources.append(transformer_source(nlp))
|
||||
notes = " ".join([notes, transformer_note(nlp)])
|
||||
# The stock description advertises a CPU tok2vec pipeline, which is wrong here.
|
||||
description = (
|
||||
"Persian pipeline built on a fine-tuned "
|
||||
f"{encoder_name(nlp)} transformer. Components: transformer, tagger, "
|
||||
"morphologizer, trainable_lemmatizer, parser, ner. Entity labels: PER, LOC, ORG, "
|
||||
"DAT, MON, TIM, PCT. GPU recommended."
|
||||
)
|
||||
if args.add_ner:
|
||||
ner_nlp = spacy.load(args.add_ner)
|
||||
if ner_nlp.pipe_names != ["ner"]:
|
||||
|
|
@ -202,22 +320,22 @@ def main():
|
|||
nlp.meta.update(
|
||||
{
|
||||
"lang": "fa",
|
||||
"name": spec["name"],
|
||||
"name": name,
|
||||
"version": args.version,
|
||||
"description": spec["description"],
|
||||
"description": description,
|
||||
"author": AUTHOR,
|
||||
"email": EMAIL,
|
||||
"url": PROJECT_URL,
|
||||
"license": spec["license"],
|
||||
"sources": spec["sources"],
|
||||
"notes": spec["notes"],
|
||||
"sources": sources,
|
||||
"notes": notes,
|
||||
"performance": performance,
|
||||
}
|
||||
)
|
||||
|
||||
out = Path(args.output)
|
||||
nlp.to_disk(out)
|
||||
print(f"wrote {out} as fa_{spec['name']} {args.version} ({spec['license']})")
|
||||
print(f"wrote {out} as fa_{name} {args.version} ({spec['license']})")
|
||||
scalars = {k: round(v * 100, 2) for k, v in performance.items() if isinstance(v, float)}
|
||||
print(json.dumps(scalars, indent=2))
|
||||
|
||||
|
|
|
|||
|
|
@ -0,0 +1,136 @@
|
|||
"""Build the Hugging Face model card for a packaged pipeline.
|
||||
|
||||
`spacy package` already writes a README into the wheel, and `spacy huggingface-hub push`
|
||||
uploads it as the card. That card is a metadata dump: no install line, no usage, no
|
||||
throughput, and no YAML frontmatter, so the Hub cannot index the model by language or task.
|
||||
|
||||
This composes a card from the same sources of truth (`meta.json` and the JSON written by
|
||||
scripts/benchmark_throughput.py) rather than from hand-copied numbers, so the card cannot
|
||||
drift from the artifact it describes.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
# meta.json key -> (row label, reference note). Only keys the pipeline actually evidences
|
||||
# are emitted; a missing key means the corpus could not score it.
|
||||
METRICS = [
|
||||
("token_acc", "Tokenization accuracy", ""),
|
||||
("tag_acc", "XPOS tag accuracy", ""),
|
||||
("pos_acc", "UPOS tag accuracy", ""),
|
||||
("morph_acc", "Morphological features", ""),
|
||||
("lemma_acc", "Lemma accuracy", ""),
|
||||
("dep_uas", "Unlabelled attachment (UAS)", ""),
|
||||
("dep_las", "Labelled attachment (LAS)", ""),
|
||||
("sents_f", "Sentence segmentation F", ""),
|
||||
("ents_p", "NER precision", ""),
|
||||
("ents_r", "NER recall", ""),
|
||||
("ents_f", "NER F-score", ""),
|
||||
]
|
||||
|
||||
|
||||
def load(path):
|
||||
return json.loads(Path(path).read_text())
|
||||
|
||||
|
||||
def throughput_rows(paths):
|
||||
rows = []
|
||||
for p in paths:
|
||||
if not Path(p).exists():
|
||||
continue
|
||||
d = load(p)
|
||||
rows.append((d["device"], d["batch_size"], d["wps_median"]))
|
||||
return rows
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--meta", required=True, help="meta.json of the finalized pipeline")
|
||||
ap.add_argument("--throughput", nargs="*", default=[], help="benchmark_throughput JSONs")
|
||||
ap.add_argument("--repo-id", required=True, help="e.g. Phazel/fa_core_news_trf")
|
||||
ap.add_argument("--wheel-name", required=True)
|
||||
ap.add_argument("--out", required=True)
|
||||
args = ap.parse_args()
|
||||
|
||||
meta = load(args.meta)
|
||||
name = f"{meta['lang']}_{meta['name']}"
|
||||
perf = meta.get("performance", {})
|
||||
|
||||
lines = []
|
||||
# Frontmatter: without this the Hub cannot filter the model by language or library.
|
||||
lines += [
|
||||
"---",
|
||||
"language:",
|
||||
"- fa",
|
||||
f"license: {meta.get('license', 'cc-by-sa-4.0').lower().replace(' ', '-')}",
|
||||
"library_name: spacy",
|
||||
"pipeline_tag: token-classification",
|
||||
"tags:",
|
||||
"- spacy",
|
||||
"- token-classification",
|
||||
"- persian",
|
||||
"- farsi",
|
||||
"---",
|
||||
"",
|
||||
f"# {name}",
|
||||
"",
|
||||
meta.get("description", "").strip(),
|
||||
"",
|
||||
]
|
||||
|
||||
lines += [
|
||||
"## Install",
|
||||
"",
|
||||
"```bash",
|
||||
f"pip install https://huggingface.co/{args.repo_id}/resolve/main/{args.wheel_name}",
|
||||
"```",
|
||||
"",
|
||||
"```python",
|
||||
"import spacy",
|
||||
f'nlp = spacy.load("{name}")',
|
||||
'doc = nlp("شرکت ایران خودرو اعلام کرد که تولید خود را افزایش می\u200cدهد.")',
|
||||
"print([(t.text, t.pos_, t.lemma_, t.dep_) for t in doc])",
|
||||
"print([(e.text, e.label_) for e in doc.ents])",
|
||||
"```",
|
||||
"",
|
||||
]
|
||||
|
||||
lines += ["## Accuracy", "",
|
||||
"Scored with `spacy benchmark accuracy` on the held-out PerDT test split.",
|
||||
"", "| Metric | Score |", "| --- | ---: |"]
|
||||
for key, label, _ in METRICS:
|
||||
v = perf.get(key)
|
||||
if isinstance(v, (int, float)):
|
||||
lines.append(f"| {label} | {v * 100:.2f} |")
|
||||
lines.append("")
|
||||
|
||||
rows = throughput_rows(args.throughput)
|
||||
if rows:
|
||||
lines += ["## Throughput", "",
|
||||
"Median of repeated `nlp.pipe` passes over the 146-document PerDT test",
|
||||
"split (23,825 tokens), timing the pipe only. Warmup pass discarded.",
|
||||
"", "| Device | Batch | Words/s |", "| --- | ---: | ---: |"]
|
||||
for device, batch, wps in rows:
|
||||
lines.append(f"| {device} | {batch} | {wps:,.0f} |")
|
||||
lines.append("")
|
||||
|
||||
lines += ["## Sources", "", "| Source | Author | Licence |", "| --- | --- | --- |"]
|
||||
for s in meta.get("sources", []):
|
||||
url, nm = s.get("url"), s.get("name", "")
|
||||
label = f"[{nm}]({url})" if url else nm
|
||||
lines.append(f"| {label} | {s.get('author', '')} | {s.get('license', '')} |")
|
||||
lines.append("")
|
||||
|
||||
notes = (meta.get("notes") or "").strip()
|
||||
if notes:
|
||||
lines += ["## Notes", "", notes, ""]
|
||||
|
||||
out = Path(args.out)
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_text("\n".join(lines), encoding="utf-8")
|
||||
print(f"wrote {out} ({out.stat().st_size} bytes)")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
|
@ -0,0 +1,68 @@
|
|||
"""Fuse the UD annotation layer and the transferred NER layer into one DocBin.
|
||||
|
||||
The sm/md/lg tiers train `ner` as a separate pipeline with its own embedded tok2vec, then
|
||||
source it into the dep model (project.yml `assemble-core`). That works because a hash-embed
|
||||
tok2vec is cheap enough to train twice.
|
||||
|
||||
A transformer is not. Fine-tuning ParsBERT once per component would double GPU cost and
|
||||
produce a package carrying two independent 162M-parameter encoders, and sourcing the second
|
||||
one would collide on the `transformer` component name. So the trf tier trains every component
|
||||
against a single shared transformer via TransformerListener, which requires a single corpus
|
||||
carrying both annotation layers on the same Doc.
|
||||
|
||||
That fusion is exact, not approximate: `corpus/perdt-ner/` was produced by
|
||||
scripts/transfer_perdt_ner.py from the same `--merge-subtokens` CoNLL-U as `corpus/merged/`,
|
||||
then converted with the same `--n-sents`, so the two DocBins are token-for-token identical
|
||||
(verified below and asserted at runtime). Only `doc.ents` is copied across; every other
|
||||
annotation stays on the UD doc.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
from pathlib import Path
|
||||
|
||||
import spacy
|
||||
from spacy.tokens import DocBin, Span
|
||||
|
||||
SPLITS = (("train", "fa_perdt-ud-train"), ("dev", "fa_perdt-ud-dev"), ("test", "fa_perdt-ud-test"))
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--ud-dir", default="corpus/merged")
|
||||
ap.add_argument("--ner-dir", default="corpus/perdt-ner")
|
||||
ap.add_argument("--out", default="corpus/joint")
|
||||
ap.add_argument("--lang", default="fa")
|
||||
args = ap.parse_args()
|
||||
|
||||
nlp = spacy.blank(args.lang)
|
||||
out = Path(args.out)
|
||||
out.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
for split, ud_stem in SPLITS:
|
||||
ud_docs = list(DocBin().from_disk(Path(args.ud_dir) / f"{ud_stem}.spacy").get_docs(nlp.vocab))
|
||||
ner_docs = list(DocBin().from_disk(Path(args.ner_dir) / f"{split}.spacy").get_docs(nlp.vocab))
|
||||
if len(ud_docs) != len(ner_docs):
|
||||
raise SystemExit(
|
||||
f"{split}: {len(ud_docs)} UD docs vs {len(ner_docs)} NER docs; the two corpora "
|
||||
"were not converted from the same source with the same --n-sents"
|
||||
)
|
||||
|
||||
db = DocBin(store_user_data=True)
|
||||
n_ents = 0
|
||||
for i, (ud, ner) in enumerate(zip(ud_docs, ner_docs)):
|
||||
if [t.text for t in ud] != [t.text for t in ner]:
|
||||
raise SystemExit(f"{split} doc {i}: tokenization differs between UD and NER layers")
|
||||
# Tokens are index-aligned, so rebuild by token index. Char offsets are NOT
|
||||
# safe here: the two converters can differ in trailing whitespace, which shifts
|
||||
# `char_span` off the token grid and silently yields None.
|
||||
ud.ents = [Span(ud, e.start, e.end, label=e.label_) for e in ner.ents]
|
||||
n_ents += len(ud.ents)
|
||||
db.add(ud)
|
||||
|
||||
dest = out / f"{split}.spacy"
|
||||
db.to_disk(dest)
|
||||
print(f"{dest}: {len(ud_docs)} docs, {n_ents} entities")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
|
@ -0,0 +1,53 @@
|
|||
"""Unpack a spaCy vectors-only wheel into a plain model directory.
|
||||
|
||||
`spacy train --paths.vectors` wants a directory it can `spacy.load()`. The fa_floret wheel
|
||||
already contains exactly that (an empty pipeline carrying only vocab/vectors), it is just
|
||||
buried under the wheel's package layout, so this unwraps it rather than pip-installing a
|
||||
package whose only job is to hold a 57 MB array.
|
||||
|
||||
Usage:
|
||||
python scripts/unpack_vectors.py fa_floret-0.1.0-py3-none-any-400k-documents.whl \\
|
||||
assets/vectors/fa_floret_400k
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import shutil
|
||||
import sys
|
||||
import tempfile
|
||||
import zipfile
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("wheel", type=Path)
|
||||
ap.add_argument("output", type=Path)
|
||||
args = ap.parse_args()
|
||||
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
tmp = Path(tmp)
|
||||
with zipfile.ZipFile(args.wheel) as z:
|
||||
z.extractall(tmp)
|
||||
# The model directory is the one holding config.cfg, e.g. fa_floret/fa_floret-0.1.0/.
|
||||
models = sorted(p.parent for p in tmp.rglob("config.cfg"))
|
||||
if len(models) != 1:
|
||||
sys.exit(f"expected exactly one config.cfg in {args.wheel}, found {len(models)}")
|
||||
if args.output.exists():
|
||||
shutil.rmtree(args.output)
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
shutil.move(str(models[0]), str(args.output))
|
||||
|
||||
import spacy
|
||||
|
||||
nlp = spacy.load(args.output)
|
||||
vectors = nlp.vocab.vectors
|
||||
if vectors.shape[0] == 0:
|
||||
sys.exit(f"{args.output} has no vectors")
|
||||
print(
|
||||
f"{args.output}: mode={vectors.mode} shape={vectors.shape} "
|
||||
f"n_keys={vectors.n_keys} pipeline={nlp.pipe_names}"
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Loading…
Reference in New Issue