224 lines
8.8 KiB
Markdown
224 lines
8.8 KiB
Markdown
# Upstream PR candidate: `spacy/lang/fa/syntax_iterators.py` uses non-UD dependency labels
|
|
|
|
## The bug
|
|
|
|
`spacy/lang/fa/syntax_iterators.py` matches these dependency labels as noun-phrase heads:
|
|
|
|
```python
|
|
labels = ["nsubj", "dobj", "nsubjpass", "pcomp", "pobj", "dative", "appos", "attr", "ROOT"]
|
|
```
|
|
|
|
`dobj`, `nsubjpass`, `pobj`, `dative` and `attr` are ClearNLP/English labels. They are not
|
|
Universal Dependencies relations, and every Persian treebank is UD (`UD_Persian-PerDT`,
|
|
`UD_Persian-Seraji`, `UD_Persian-PUD`, `UD_Persian-IPerUDT`). Any trained `fa` pipeline emits
|
|
UD labels, so five of the nine labels are dead code and `doc.noun_chunks` returns bare head
|
|
nouns.
|
|
|
|
`spacy/lang/fr/syntax_iterators.py` and `spacy/lang/es/syntax_iterators.py` use UD labels
|
|
(`nsubj`, `nsubj:pass`, `obj`, `obl`, `nmod`, `appos`, `ROOT`) because their treebanks are UD.
|
|
`spacy/lang/de` uses TIGER labels (`sb`, `oa`, `nk`, …) because the German pipelines are
|
|
trained on TIGER. Persian has no equivalent reason.
|
|
|
|
A second, smaller problem: the iterator yields `word.left_edge.i` to `word.i + 1`, so it only
|
|
expands left. Persian noun phrases expand right through ezafe. `رئیس انجمن جراحان قلب ایران`
|
|
("the head of the Iranian society of heart surgeons") is one NP whose head is the leftmost
|
|
token, and left-only expansion truncates it to `رئیس`.
|
|
|
|
## Evidence
|
|
|
|
Measured with `scripts/check_noun_chunks.py` on the `UD_Persian-PerDT` dev split (146 docs)
|
|
parsed by this project's trained pipeline:
|
|
|
|
```
|
|
shipped noun_chunks: 1503 chunks, 1.31 tokens/chunk
|
|
proposed noun_chunks: 5300 chunks, 2.77 tokens/chunk
|
|
|
|
deprels on NOUN/PROPN/PRON tokens in dev (top 15):
|
|
nmod 2894 -
|
|
obl 1401 -
|
|
nsubj 1326 shipped
|
|
compound:lvc 1294 -
|
|
obl:arg 1095 -
|
|
obj 973 -
|
|
conj 583 shipped
|
|
flat:name 389 -
|
|
ROOT 101 shipped
|
|
xcomp 89 -
|
|
appos 51 shipped
|
|
nsubj:pass 38 -
|
|
```
|
|
|
|
At 1.31 tokens per chunk the shipped iterator is returning single head nouns. The four most
|
|
common noun-bearing relations after `nsubj`, namely `nmod` (2894), `obl` (1401), `obl:arg`
|
|
(1095) and `obj` (973), are all unreachable.
|
|
|
|
Live text:
|
|
|
|
```
|
|
رئیس انجمن جراحان قلب ایران تأکید کرد که امکانات پیشرفته در ایران وجود دارد.
|
|
shipped : ['رئیس', 'امکانات']
|
|
proposed: ['رئیس انجمن جراحان قلب ایران', 'امکانات پیشرفته', 'ایران']
|
|
|
|
دانشگاه تهران بزرگترین دانشگاه ایران است.
|
|
shipped : ['دانشگاه']
|
|
proposed: ['دانشگاه تهران', 'ایران']
|
|
```
|
|
|
|
## Proposed patch
|
|
|
|
Swap in UD labels, expand right through modifier chains, and drop a leading `ADP` or `CCONJ`
|
|
as `fr` and `es` already do:
|
|
|
|
```diff
|
|
--- a/spacy/lang/fa/syntax_iterators.py
|
|
+++ b/spacy/lang/fa/syntax_iterators.py
|
|
@@
|
|
def noun_chunks(doclike: Union[Doc, Span]) -> Iterator[Tuple[int, int, int]]:
|
|
"""
|
|
Detect base noun phrases from a dependency parse. Works on both Doc and Span.
|
|
"""
|
|
- labels = [
|
|
- "nsubj",
|
|
- "dobj",
|
|
- "nsubjpass",
|
|
- "pcomp",
|
|
- "pobj",
|
|
- "dative",
|
|
- "appos",
|
|
- "attr",
|
|
- "ROOT",
|
|
- ]
|
|
+ # Persian pipelines are trained on Universal Dependencies treebanks
|
|
+ # (UD_Persian-PerDT, UD_Persian-Seraji), so these are UD relations.
|
|
+ labels = [
|
|
+ "nsubj",
|
|
+ "nsubj:pass",
|
|
+ "obj",
|
|
+ "iobj",
|
|
+ "obl",
|
|
+ "obl:arg",
|
|
+ "nmod",
|
|
+ "appos",
|
|
+ "vocative",
|
|
+ "ROOT",
|
|
+ ]
|
|
+ # Persian noun phrases grow to the right through ezafe constructions.
|
|
+ post_modifiers = [
|
|
+ "nmod",
|
|
+ "nmod:poss",
|
|
+ "amod",
|
|
+ "det",
|
|
+ "nummod",
|
|
+ "flat",
|
|
+ "flat:name",
|
|
+ "flat:num",
|
|
+ "fixed",
|
|
+ "compound",
|
|
+ ]
|
|
doc = doclike.doc # Ensure works on both Doc and Span.
|
|
|
|
if not doc.has_annotation("DEP"):
|
|
raise ValueError(Errors.E029)
|
|
|
|
np_deps = [doc.vocab.strings.add(label) for label in labels]
|
|
+ np_modifs = {doc.vocab.strings.add(label) for label in post_modifiers}
|
|
conj = doc.vocab.strings.add("conj")
|
|
np_label = doc.vocab.strings.add("NP")
|
|
+ adp_pos = doc.vocab.strings.add("ADP")
|
|
+ cconj_pos = doc.vocab.strings.add("CCONJ")
|
|
prev_end = -1
|
|
for i, word in enumerate(doclike):
|
|
if word.pos not in (NOUN, PROPN, PRON):
|
|
continue
|
|
# Prevent nested chunks from being produced
|
|
if word.left_edge.i <= prev_end:
|
|
continue
|
|
- if word.dep in np_deps:
|
|
- prev_end = word.i
|
|
- yield word.left_edge.i, word.i + 1, np_label
|
|
- elif word.dep == conj:
|
|
+ if word.dep == conj:
|
|
head = word.head
|
|
while head.dep == conj and head.head.i < head.i:
|
|
head = head.head
|
|
- # If the head is an NP, and we're coordinated to it, we're an NP
|
|
- if head.dep in np_deps:
|
|
- prev_end = word.i
|
|
- yield word.left_edge.i, word.i + 1, np_label
|
|
+ # If the head is an NP, and we're coordinated to it, we're an NP
|
|
+ if head.dep not in np_deps:
|
|
+ continue
|
|
+ elif word.dep not in np_deps:
|
|
+ continue
|
|
+ # Expand right through the ezafe / modifier chain.
|
|
+ right = word
|
|
+ for child in word.rights:
|
|
+ if child.dep in np_modifs:
|
|
+ right = child.right_edge
|
|
+ else:
|
|
+ break
|
|
+ start, end = word.left_edge.i, max(word.i, right.i) + 1
|
|
+ # A leading preposition or coordinator is not part of the NP.
|
|
+ while start < word.i and doc[start].pos in (adp_pos, cconj_pos):
|
|
+ start += 1
|
|
+ if end <= prev_end:
|
|
+ continue
|
|
+ prev_end = end - 1
|
|
+ yield start, end, np_label
|
|
```
|
|
|
|
The working implementation lives in `scripts/check_noun_chunks.py::proposed_noun_chunks`.
|
|
|
|
## Test to add
|
|
|
|
`spacy/tests/lang/fa/test_noun_chunks.py` currently has one test (a hand-built `Doc`, 296 B).
|
|
Add UD-label coverage:
|
|
|
|
```python
|
|
def test_fa_noun_chunks_ezafe(fa_vocab):
|
|
# رئیس انجمن جراحان, "head of the surgeons' society"
|
|
words = ["رئیس", "انجمن", "جراحان", "آمد"]
|
|
heads = [3, 0, 1, 3]
|
|
deps = ["nsubj", "nmod", "nmod", "ROOT"]
|
|
pos = ["NOUN", "NOUN", "NOUN", "VERB"]
|
|
doc = Doc(fa_vocab, words=words, heads=heads, deps=deps, pos=pos)
|
|
assert [c.text for c in doc.noun_chunks] == ["رئیس انجمن جراحان"]
|
|
|
|
|
|
def test_fa_noun_chunks_drops_leading_adp(fa_vocab):
|
|
words = ["در", "ایران", "بود"]
|
|
heads = [1, 2, 2]
|
|
deps = ["case", "obl", "ROOT"]
|
|
pos = ["ADP", "NOUN", "VERB"]
|
|
doc = Doc(fa_vocab, words=words, heads=heads, deps=deps, pos=pos)
|
|
assert [c.text for c in doc.noun_chunks] == ["ایران"]
|
|
```
|
|
|
|
## Other `spacy/lang/fa` gaps found while building this pipeline
|
|
|
|
Ordered by how much they cost a Persian pipeline.
|
|
|
|
1. Clitic splitting. The `fa` tokenizer cannot split pronominal enclitics or the enclitic
|
|
copula (`پدرم` = `پدر` + `م`, `ساکتند` = `ساکت` + `ند`), which UD treebanks annotate as
|
|
multiword tokens. Measured on PerDT dev, gold-against-tokenizer token F is 0.9823 when
|
|
clitics are kept split, affecting 1.49% of tokens. Persian-specific `TOKENIZER_SUFFIXES`
|
|
entries for the enclitic set would remove the need for `--merge-subtokens` and the
|
|
composite XPOS tags it produces.
|
|
2. `punctuation.py` defines only `TOKENIZER_SUFFIXES`, with no `TOKENIZER_PREFIXES` and no
|
|
`TOKENIZER_INFIXES`. ZWNJ (U+200C) is handled only implicitly through the 65 KB generated
|
|
verb-exception table. The missing infix rules break numerics: `spacy/lang/fa/examples.py`
|
|
ships the sentence `دیروز علی به من ۲۰۰۰.۱﷼ پول نقد داد.`, and the trained pipeline splits
|
|
`۲۰۰۰.۱﷼` into `۲۰۰۰` + `.` + `۱﷼`, promoting the stray `.` to a sentence boundary, so one
|
|
input sentence comes out as three. `LIKE_NUM` in `lex_attrs.py` recognises Persian digits,
|
|
but no tokenizer rule keeps a Persian decimal or a currency sign attached.
|
|
3. No tokenizer tests for `fa`. `spacy/tests/lang/fa/` contains only `test_noun_chunks.py`,
|
|
which guards none of the 65 KB exception table.
|
|
4. `spacy-lookups-data` has no `fa_license.txt`, although `fa_source.txt` records that the
|
|
lemma tables were "extracted from Mojgan Seraji's Persian Universal Dependencies Corpus",
|
|
which is CC BY-SA 4.0. Catalan ships a `ca_license.txt`.
|
|
5. No `fa_lemma_lookup.json`. Only rule-mode lemmatizer assets exist, so `mode="lookup"` is
|
|
unavailable for Persian.
|
|
|
|
Items 1 to 3 are self-contained code PRs. Item 4 is a licence-hygiene PR against
|
|
`spacy-lookups-data`, and affects anyone redistributing a Persian pipeline.
|