spacy-fa-pipeline/docs/upstream/fa-noun-chunks.md

224 lines
8.8 KiB
Markdown

# Upstream PR candidate: `spacy/lang/fa/syntax_iterators.py` uses non-UD dependency labels
## The bug
`spacy/lang/fa/syntax_iterators.py` matches these dependency labels as noun-phrase heads:
```python
labels = ["nsubj", "dobj", "nsubjpass", "pcomp", "pobj", "dative", "appos", "attr", "ROOT"]
```
`dobj`, `nsubjpass`, `pobj`, `dative` and `attr` are ClearNLP/English labels. They are not
Universal Dependencies relations, and every Persian treebank is UD (`UD_Persian-PerDT`,
`UD_Persian-Seraji`, `UD_Persian-PUD`, `UD_Persian-IPerUDT`). Any trained `fa` pipeline emits
UD labels, so five of the nine labels are dead code and `doc.noun_chunks` returns bare head
nouns.
`spacy/lang/fr/syntax_iterators.py` and `spacy/lang/es/syntax_iterators.py` use UD labels
(`nsubj`, `nsubj:pass`, `obj`, `obl`, `nmod`, `appos`, `ROOT`) because their treebanks are UD.
`spacy/lang/de` uses TIGER labels (`sb`, `oa`, `nk`, …) because the German pipelines are
trained on TIGER. Persian has no equivalent reason.
A second, smaller problem: the iterator yields `word.left_edge.i` to `word.i + 1`, so it only
expands left. Persian noun phrases expand right through ezafe. `رئیس انجمن جراحان قلب ایران`
("the head of the Iranian society of heart surgeons") is one NP whose head is the leftmost
token, and left-only expansion truncates it to `رئیس`.
## Evidence
Measured with `scripts/check_noun_chunks.py` on the `UD_Persian-PerDT` dev split (146 docs)
parsed by this project's trained pipeline:
```
shipped noun_chunks: 1503 chunks, 1.31 tokens/chunk
proposed noun_chunks: 5300 chunks, 2.77 tokens/chunk
deprels on NOUN/PROPN/PRON tokens in dev (top 15):
nmod 2894 -
obl 1401 -
nsubj 1326 shipped
compound:lvc 1294 -
obl:arg 1095 -
obj 973 -
conj 583 shipped
flat:name 389 -
ROOT 101 shipped
xcomp 89 -
appos 51 shipped
nsubj:pass 38 -
```
At 1.31 tokens per chunk the shipped iterator is returning single head nouns. The four most
common noun-bearing relations after `nsubj`, namely `nmod` (2894), `obl` (1401), `obl:arg`
(1095) and `obj` (973), are all unreachable.
Live text:
```
رئیس انجمن جراحان قلب ایران تأکید کرد که امکانات پیشرفته در ایران وجود دارد.
shipped : ['رئیس', 'امکانات']
proposed: ['رئیس انجمن جراحان قلب ایران', 'امکانات پیشرفته', 'ایران']
دانشگاه تهران بزرگ‌ترین دانشگاه ایران است.
shipped : ['دانشگاه']
proposed: ['دانشگاه تهران', 'ایران']
```
## Proposed patch
Swap in UD labels, expand right through modifier chains, and drop a leading `ADP` or `CCONJ`
as `fr` and `es` already do:
```diff
--- a/spacy/lang/fa/syntax_iterators.py
+++ b/spacy/lang/fa/syntax_iterators.py
@@
def noun_chunks(doclike: Union[Doc, Span]) -> Iterator[Tuple[int, int, int]]:
"""
Detect base noun phrases from a dependency parse. Works on both Doc and Span.
"""
- labels = [
- "nsubj",
- "dobj",
- "nsubjpass",
- "pcomp",
- "pobj",
- "dative",
- "appos",
- "attr",
- "ROOT",
- ]
+ # Persian pipelines are trained on Universal Dependencies treebanks
+ # (UD_Persian-PerDT, UD_Persian-Seraji), so these are UD relations.
+ labels = [
+ "nsubj",
+ "nsubj:pass",
+ "obj",
+ "iobj",
+ "obl",
+ "obl:arg",
+ "nmod",
+ "appos",
+ "vocative",
+ "ROOT",
+ ]
+ # Persian noun phrases grow to the right through ezafe constructions.
+ post_modifiers = [
+ "nmod",
+ "nmod:poss",
+ "amod",
+ "det",
+ "nummod",
+ "flat",
+ "flat:name",
+ "flat:num",
+ "fixed",
+ "compound",
+ ]
doc = doclike.doc # Ensure works on both Doc and Span.
if not doc.has_annotation("DEP"):
raise ValueError(Errors.E029)
np_deps = [doc.vocab.strings.add(label) for label in labels]
+ np_modifs = {doc.vocab.strings.add(label) for label in post_modifiers}
conj = doc.vocab.strings.add("conj")
np_label = doc.vocab.strings.add("NP")
+ adp_pos = doc.vocab.strings.add("ADP")
+ cconj_pos = doc.vocab.strings.add("CCONJ")
prev_end = -1
for i, word in enumerate(doclike):
if word.pos not in (NOUN, PROPN, PRON):
continue
# Prevent nested chunks from being produced
if word.left_edge.i <= prev_end:
continue
- if word.dep in np_deps:
- prev_end = word.i
- yield word.left_edge.i, word.i + 1, np_label
- elif word.dep == conj:
+ if word.dep == conj:
head = word.head
while head.dep == conj and head.head.i < head.i:
head = head.head
- # If the head is an NP, and we're coordinated to it, we're an NP
- if head.dep in np_deps:
- prev_end = word.i
- yield word.left_edge.i, word.i + 1, np_label
+ # If the head is an NP, and we're coordinated to it, we're an NP
+ if head.dep not in np_deps:
+ continue
+ elif word.dep not in np_deps:
+ continue
+ # Expand right through the ezafe / modifier chain.
+ right = word
+ for child in word.rights:
+ if child.dep in np_modifs:
+ right = child.right_edge
+ else:
+ break
+ start, end = word.left_edge.i, max(word.i, right.i) + 1
+ # A leading preposition or coordinator is not part of the NP.
+ while start < word.i and doc[start].pos in (adp_pos, cconj_pos):
+ start += 1
+ if end <= prev_end:
+ continue
+ prev_end = end - 1
+ yield start, end, np_label
```
The working implementation lives in `scripts/check_noun_chunks.py::proposed_noun_chunks`.
## Test to add
`spacy/tests/lang/fa/test_noun_chunks.py` currently has one test (a hand-built `Doc`, 296 B).
Add UD-label coverage:
```python
def test_fa_noun_chunks_ezafe(fa_vocab):
# رئیس انجمن جراحان, "head of the surgeons' society"
words = ["رئیس", "انجمن", "جراحان", "آمد"]
heads = [3, 0, 1, 3]
deps = ["nsubj", "nmod", "nmod", "ROOT"]
pos = ["NOUN", "NOUN", "NOUN", "VERB"]
doc = Doc(fa_vocab, words=words, heads=heads, deps=deps, pos=pos)
assert [c.text for c in doc.noun_chunks] == ["رئیس انجمن جراحان"]
def test_fa_noun_chunks_drops_leading_adp(fa_vocab):
words = ["در", "ایران", "بود"]
heads = [1, 2, 2]
deps = ["case", "obl", "ROOT"]
pos = ["ADP", "NOUN", "VERB"]
doc = Doc(fa_vocab, words=words, heads=heads, deps=deps, pos=pos)
assert [c.text for c in doc.noun_chunks] == ["ایران"]
```
## Other `spacy/lang/fa` gaps found while building this pipeline
Ordered by how much they cost a Persian pipeline.
1. Clitic splitting. The `fa` tokenizer cannot split pronominal enclitics or the enclitic
copula (`پدرم` = `پدر` + `م`, `ساکتند` = `ساکت` + `ند`), which UD treebanks annotate as
multiword tokens. Measured on PerDT dev, gold-against-tokenizer token F is 0.9823 when
clitics are kept split, affecting 1.49% of tokens. Persian-specific `TOKENIZER_SUFFIXES`
entries for the enclitic set would remove the need for `--merge-subtokens` and the
composite XPOS tags it produces.
2. `punctuation.py` defines only `TOKENIZER_SUFFIXES`, with no `TOKENIZER_PREFIXES` and no
`TOKENIZER_INFIXES`. ZWNJ (U+200C) is handled only implicitly through the 65 KB generated
verb-exception table. The missing infix rules break numerics: `spacy/lang/fa/examples.py`
ships the sentence `دیروز علی به من ۲۰۰۰.۱﷼ پول نقد داد.`, and the trained pipeline splits
`۲۰۰۰.۱﷼` into `۲۰۰۰` + `.` + `۱﷼`, promoting the stray `.` to a sentence boundary, so one
input sentence comes out as three. `LIKE_NUM` in `lex_attrs.py` recognises Persian digits,
but no tokenizer rule keeps a Persian decimal or a currency sign attached.
3. No tokenizer tests for `fa`. `spacy/tests/lang/fa/` contains only `test_noun_chunks.py`,
which guards none of the 65 KB exception table.
4. `spacy-lookups-data` has no `fa_license.txt`, although `fa_source.txt` records that the
lemma tables were "extracted from Mojgan Seraji's Persian Universal Dependencies Corpus",
which is CC BY-SA 4.0. Catalan ships a `ca_license.txt`.
5. No `fa_lemma_lookup.json`. Only rule-mode lemmatizer assets exist, so `mode="lookup"` is
unavailable for Persian.
Items 1 to 3 are self-contained code PRs. Item 4 is a licence-hygiene PR against
`spacy-lookups-data`, and affects anyone redistributing a Persian pipeline.