Lead the README tables with trf instead of sm
The Hazm and en_core_web_sm comparison now cites fa_core_news_trf, which is the strongest tier and the only one to pass the hazm+ParsBERT DEP_LAS reference. Drop the package-size row from that table: it compared a 608 MB transformer against a 7 MB rule-based toolkit, which says nothing about accuracy. Add trf to the package list and to the per-label NER table, alongside lg, which had never been added. Its largest per-label gains are PER +9.25 and DAT +11.69 over lg. Drop the gold-count column there and trim the surrounding notes.
This commit is contained in:
parent
5da9dd1524
commit
6fa8aa70b8
19
README.fa.md
19
README.fa.md
|
|
@ -42,10 +42,9 @@ print(doc.ents) # (محمدرضا شجریان, مشهد)
|
||||||
| سرعت (940MX، دستهٔ ۳۲) | ۱۰٬۲۳۵ | ۹٬۰۵۸ | ۹٬۲۱۵ | بخش توان عملیاتی | |
|
| سرعت (940MX، دستهٔ ۳۲) | ۱۰٬۲۳۵ | ۹٬۰۵۸ | ۹٬۲۱۵ | بخش توان عملیاتی | |
|
||||||
| حجم بستهٔ نصب | ۱۳٫۵ مگابایت | ۶۸٫۵ مگابایت | ۲۳۵ مگابایت | ۶۰۸ مگابایت | |
|
| حجم بستهٔ نصب | ۱۳٫۵ مگابایت | ۶۸٫۵ مگابایت | ۲۳۵ مگابایت | ۶۰۸ مگابایت | |
|
||||||
|
|
||||||
ردهٔ `trf` که ParsBERT را ریزتنظیم میکند در همهجا جلو است مگر در واژهیابی و مرزبندی جمله، که
|
ردهٔ `trf` در همهجا جلو است مگر در واژهیابی و مرزبندی جمله، و تنها ردهٔای است که از مرجع
|
||||||
`lg` با واژهیاب درختویرایش روی زیرواژههای floret همچنان بهتر عمل میکند. تنها ردهٔای است که از
|
`DEP_LAS` برابر ۸۹٫۳۴ عبور میکند. به کارت گرافیک نیاز دارد و مدل پایهٔ آن پروانهٔ مشخصی ندارد،
|
||||||
مرجع `DEP_LAS` برابر ۸۹٫۳۴ عبور میکند. این رده به کارت گرافیک نیاز دارد و مدل پایهٔ آن پروانهٔ
|
پس قابل بازانتشار نیست (`docs/MODELS.md` بخش ۸).
|
||||||
مشخصی ندارد، پس قابل بازانتشار نیست؛ هر دو نکته در `docs/MODELS.md` بخش ۸ آمده است.
|
|
||||||
|
|
||||||
برچسبهای موجودیت «نقرهای» هستند: از لایهای در خود پیکره میآیند که با برچسبزن Beheshti-NER
|
برچسبهای موجودیت «نقرهای» هستند: از لایهای در خود پیکره میآیند که با برچسبزن Beheshti-NER
|
||||||
تولید و سپس دستی اصلاح شده است. بنابراین `ENTS_F` تا اندازهای همخوانی با آن برچسبزن را
|
تولید و سپس دستی اصلاح شده است. بنابراین `ENTS_F` تا اندازهای همخوانی با آن برچسبزن را
|
||||||
|
|
@ -68,14 +67,10 @@ print(doc.ents) # (محمدرضا شجریان, مشهد)
|
||||||
| `lg` | ۴٬۷۱۵ | ۹٬۲۱۵ | |
|
| `lg` | ۴٬۷۱۵ | ۹٬۲۱۵ | |
|
||||||
| `trf` | ۱۸۷ | | ۸٬۳۲۰ |
|
| `trf` | ۱۸۷ | | ۸٬۳۲۰ |
|
||||||
|
|
||||||
ردهٔ `trf` جنس دیگری دارد: روی همان پردازندهٔ لپتاپ ۱۸۷ واژه بر ثانیه است، یعنی حدود ۲۹ برابر
|
ردهٔ `trf` روی یک پردازنده حدود ۲۹ برابر کندتر از `sm` است، و روی T4 نسبت به پردازندهٔ همان
|
||||||
کندتر از `sm` با ۵٬۴۸۴. روی T4 به ۸٬۳۲۰ میرسد و روی پردازندهٔ همان ماشین ۳۳۶، یعنی شتاب ۲۵
|
ماشین (۳۳۶ واژه بر ثانیه) ۲۵ برابر سریعتر، پس کارت گرافیک برای آن یک نیاز است نه بهینهسازی.
|
||||||
برابری. پس کارت گرافیک برای این رده یک نیاز است نه یک بهینهسازی. کارت 940MX اصلاً `trf` را
|
فاصلهٔ ردههای پردازندهای کمتر از ۱۵ درصد است، یعنی گلوگاه جستوجوی tok2vec نیست بلکه تجزیهگر
|
||||||
اجرا نمیکند، چون نسخههای امروزی PyTorch پشتیبانی از معماری sm_50 را کنار گذاشتهاند.
|
و واژهیاب است. پراکندگی اجراها روی لپتاپ بسته به دمای دستگاه حدود ۱۰± درصد است.
|
||||||
|
|
||||||
فاصلهٔ `sm` و `md` و `lg` روی پردازنده کمتر از ۱۵ درصد است، یعنی کمتر از آنچه تفاوت اندازهٔ
|
|
||||||
جدول بردارها نشان میدهد: گلوگاه tok2vec نیست، تجزیهگر و واژهیاب است. پراکندگی اجراها روی
|
|
||||||
لپتاپ بسته به دمای دستگاه حدود ۱۰± درصد است، پس تفاوتهای کمتر از آن نویز شمرده میشوند.
|
|
||||||
|
|
||||||
گامهای تبدیل پیکره، آموزش، ارزیابی و بستهبندی در [`project.yml`](project.yml) تعریف شدهاند.
|
گامهای تبدیل پیکره، آموزش، ارزیابی و بستهبندی در [`project.yml`](project.yml) تعریف شدهاند.
|
||||||
توضیح بیشتر دربارهٔ گزینش پیکره و پروانهها در [`docs/MODELS.md`](docs/MODELS.md) و شرح انگلیسی
|
توضیح بیشتر دربارهٔ گزینش پیکره و پروانهها در [`docs/MODELS.md`](docs/MODELS.md) و شرح انگلیسی
|
||||||
|
|
|
||||||
66
README.md
66
README.md
|
|
@ -25,13 +25,12 @@ pip install https://huggingface.co/Phazel/fa_core_news_sm/resolve/main/fa_core_n
|
||||||
|
|
||||||
Compared against Hazm (the most-used Persian toolkit) and `en_core_web_sm` (English reference).
|
Compared against Hazm (the most-used Persian toolkit) and `en_core_web_sm` (English reference).
|
||||||
|
|
||||||
| Metric | **`spacy-persian`**<br>`fa_core_news_sm` | **Hazm**<br>(Persian toolkit) | `en_core_web_sm`<br>(English reference) |
|
| Metric | **`spacy-persian`**<br>`fa_core_news_trf` | **Hazm**<br>(Persian toolkit) | `en_core_web_sm`<br>(English reference) |
|
||||||
|--------|:---:|:---:|:---:|
|
|--------|:---:|:---:|:---:|
|
||||||
| **POS Accuracy (UPOS)** | **96.24%** | ~95.69%¹ | 97.21%² |
|
| **POS Accuracy (UPOS)** | **97.63%** | ~95.69%¹ | 97.21%² |
|
||||||
| **Lemma Accuracy** | **97.91%** | 89.9%¹ | — |
|
| **Lemma Accuracy** | **97.31%** | 89.9%¹ | — |
|
||||||
| **Dependency LAS** | 85.15% | 85.6%¹ | 91.85%² |
|
| **Dependency LAS** | **90.79%** | 85.6%¹ | 91.85%² |
|
||||||
| **NER F-score** | 71.87% | — | 83.80%² |
|
| **NER F-score** | **82.89%** | — | 83.80%² |
|
||||||
| **Package Size** | **13 MB** (syntax+NER)<br>**7.5 MB** (syntax-only) | ~7 MB | 12 MB |
|
|
||||||
|
|
||||||
> **¹** Hazm scores from its official README
|
> **¹** Hazm scores from its official README
|
||||||
> **²** `en_core_web_sm` scores from spaCy's official model card
|
> **²** `en_core_web_sm` scores from spaCy's official model card
|
||||||
|
|
@ -48,6 +47,7 @@ From `spacy benchmark accuracy`, stored in `metrics/`.
|
||||||
| `fa_dep_news_md` | same as `fa_dep_news_sm`, plus floret vectors | CC BY-SA 4.0 | LEMMA 97.96 | 62 MB |
|
| `fa_dep_news_md` | same as `fa_dep_news_sm`, plus floret vectors | CC BY-SA 4.0 | LEMMA 97.96 | 62 MB |
|
||||||
| `fa_core_news_md` | same as `fa_core_news_sm`, plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 68 MB |
|
| `fa_core_news_md` | same as `fa_core_news_sm`, plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 68 MB |
|
||||||
| `fa_ent_news_md` | `ner` alone (own embedded tok2vec), plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 58 MB |
|
| `fa_ent_news_md` | `ner` alone (own embedded tok2vec), plus floret vectors | CC BY-SA 4.0 | ENTS_F 74.71 | 58 MB |
|
||||||
|
| `fa_core_news_trf` | transformer, tagger, morphologizer, trainable_lemmatizer, parser, ner | see §8, encoder unlicensed | ENTS_F 82.89, LAS 90.79 | 608 MB |
|
||||||
|
|
||||||
The `md` tier adds a 50k x 300d floret vector table trained on 400k Persian documents. Its
|
The `md` tier adds a 50k x 300d floret vector table trained on 400k Persian documents. Its
|
||||||
config differs from `sm` by exactly one line (`include_static_vectors`), so the columns below
|
config differs from `sm` by exactly one line (`include_static_vectors`), so the columns below
|
||||||
|
|
@ -69,11 +69,9 @@ isolate what the vectors buy. Full breakdown in `docs/MODELS.md` §6.
|
||||||
| Speed (940MX, batch 32) | 10,235 words/s | 9,058 words/s | 9,215 words/s | see §Throughput | |
|
| Speed (940MX, batch 32) | 10,235 words/s | 9,058 words/s | 9,215 words/s | see §Throughput | |
|
||||||
| Wheel size | 13.5 MB | 68.5 MB | 235 MB | 608 MB | |
|
| Wheel size | 13.5 MB | 68.5 MB | 235 MB | 608 MB | |
|
||||||
|
|
||||||
`trf` fine-tunes ParsBERT and wins everywhere except lemmatization and sentence
|
`trf` leads everywhere except lemmatization and sentence segmentation, and is the only tier to
|
||||||
segmentation, where `lg`'s edit-tree lemmatizer over floret subwords still leads. It is the
|
pass the hazm+ParsBERT `DEP_LAS` reference of 89.34. It needs a GPU, and its encoder states no
|
||||||
only tier to pass the hazm+ParsBERT `DEP_LAS` reference of 89.34. It needs a GPU and its
|
licence so it is not redistributable (`docs/MODELS.md` §8).
|
||||||
encoder has no stated licence, so it is not redistributable; `docs/MODELS.md` §8 has both
|
|
||||||
caveats.
|
|
||||||
|
|
||||||
Entity scores are `fa_core_news_*` on the PerDT NER test split; per-label breakdown and
|
Entity scores are `fa_core_news_*` on the PerDT NER test split; per-label breakdown and
|
||||||
caveats are in [Named entity recognition](#named-entity-recognition).
|
caveats are in [Named entity recognition](#named-entity-recognition).
|
||||||
|
|
@ -97,16 +95,10 @@ timing the pipe only, warmup discarded. Reproduce with
|
||||||
| `lg` | 4,715 | 9,215 | |
|
| `lg` | 4,715 | 9,215 | |
|
||||||
| `trf` | 187 | | 8,320 |
|
| `trf` | 187 | | 8,320 |
|
||||||
|
|
||||||
The `trf` tier is a different kind of thing: 187 words/s on the same laptop CPU that runs
|
`trf` runs 29x slower than `sm` on the same CPU, and 25x faster on a T4 than on that VM's own
|
||||||
`sm` at 5,484, so about 29x slower. On a T4 it reaches 8,320, and on that VM's own Xeon it
|
Xeon (336 words/s), so a GPU is a requirement rather than an optimization. The CPU tiers sit
|
||||||
manages 336, a 25x GPU speedup. Treat GPU as a requirement rather than an optimization.
|
within 15% of each other, so the tok2vec lookup is not the bottleneck; the parser and
|
||||||
The 940MX cannot run `trf` at all, since current PyTorch wheels have dropped its sm_50
|
lemmatizer are. Laptop spread is about 10% with thermal state.
|
||||||
compute capability.
|
|
||||||
|
|
||||||
`sm`, `md` and `lg` are within about 15% of each other on CPU, which is smaller than the
|
|
||||||
gap in vector-table size suggests: the tok2vec is not the bottleneck, the parser and
|
|
||||||
lemmatizer are. Run-to-run spread on the laptop is roughly +/-10% depending on thermal
|
|
||||||
state, so treat differences under that as noise.
|
|
||||||
|
|
||||||
## Named entity recognition
|
## Named entity recognition
|
||||||
|
|
||||||
|
|
@ -119,23 +111,25 @@ recall, so the `ENTS_F` numbers below partly reflect agreement with that tagger,
|
||||||
human annotation.
|
human annotation.
|
||||||
|
|
||||||
`ner` runs standalone with its own embedded tok2vec (`fa_ent_news_sm`, `fa_ent_news_md`), or
|
`ner` runs standalone with its own embedded tok2vec (`fa_ent_news_sm`, `fa_ent_news_md`), or
|
||||||
bundled into `fa_core_news_sm`/`fa_core_news_md` alongside the syntax pipeline.
|
bundled into `fa_core_news_sm`/`fa_core_news_md` alongside the syntax pipeline. In `trf` it is
|
||||||
|
trained jointly against the shared transformer instead, so there is no standalone trf variant.
|
||||||
|
|
||||||
| Label | Gold in test | `sm` F | `md` F | Train examples |
|
| Label | `sm` F | `md` F | `lg` F | `trf` F | Train examples |
|
||||||
| --- | --- | --- | --- | --- |
|
| --- | --- | --- | --- | --- | --- |
|
||||||
| `LOC` | 273 | 80.24 | 84.05 | 4,954 |
|
| `LOC` | 80.24 | 84.05 | 83.66 | **87.78** | 4,954 |
|
||||||
| `PER` | 297 | 65.29 | 68.18 | 4,847 |
|
| `PER` | 65.29 | 68.18 | 72.63 | **81.88** | 4,847 |
|
||||||
| `ORG` | 144 | 68.77 | 70.25 | 2,643 |
|
| `ORG` | 68.77 | 70.25 | 71.01 | **78.50** | 2,643 |
|
||||||
| `DAT` | 69 | 74.45 | 76.19 | 1,323 |
|
| `DAT` | 74.45 | 76.19 | 70.83 | **82.52** | 1,323 |
|
||||||
| `MON` | 10 | 73.68 | 84.21 | 205 |
|
| `MON` | 73.68 | 84.21 | 88.89 | 88.89 | 205 |
|
||||||
| `TIM` | 9 | 66.67 | 66.67 | 135 |
|
| `TIM` | 66.67 | 66.67 | 61.54 | 50.00 | 135 |
|
||||||
| `PCT` | 4 | 57.14 | 33.33 | 121 |
|
| `PCT` | 57.14 | 33.33 | 57.14 | 33.33 | 121 |
|
||||||
|
|
||||||
`MON`, `TIM` and `PCT` have single-digit support in the test split, so their deltas are one
|
`MON`, `TIM` and `PCT` have single-digit support in the test split, so their deltas are one or
|
||||||
or two entities changing hands, not signal. `PER`, `LOC` and `ORG` carry the split and all
|
two entities changing hands, not signal. `PER`, `LOC` and `ORG` carry the split. The `md` gain
|
||||||
improve with floret vectors; the `md` gain over `sm` (`ENTS_F` 71.87 to 74.71) is almost
|
over `sm` (`ENTS_F` 71.87 to 74.71) is almost entirely recall (+6.08), the lexical prior static
|
||||||
entirely recall (+6.08), the lexical prior static vectors give rare proper nouns that hash
|
vectors give rare proper nouns that hash embeddings never had. `trf` adds another +6.95 F over
|
||||||
embeddings never had.
|
`lg`, again mostly recall (71.09 to 81.76), and its largest per-label gains are `PER` (+9.25)
|
||||||
|
and `DAT` (+11.69).
|
||||||
|
|
||||||
|
|
||||||
## Install
|
## Install
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue