15 lines
1.8 KiB
Markdown
15 lines
1.8 KiB
Markdown
An implementation of stochastic context free grammar induction, following Stolcke and Omohundro's
|
|
[Inducing Probabilistic Grammars by Bayesian Model Merging](https://arxiv.org/abs/cmp-lg/9409010) (ICGI 1994).
|
|
|
|
<p align="center">
|
|
<img src="meta/paper-and-induced.svg" alt="the paper's Figure 2 and Table 1, and the eleven induced grammars" width="900">
|
|
</p>
|
|
|
|
We start from the most specific grammar the data permits: every sample contributes its own production, and every terminal that occurs gets a corresponding nonterminal. At this stage there's effectively no sharing between samples, so the grammar is just memorising the corpus rather than generalising beyond it.
|
|
|
|
From there, we generalise with two operators, merging and chunking. Merging takes a pair of nonterminals and folds them into a single nonterminal containing the union of their productions; chunking replaces a contiguous sequence of symbols with a fresh nonterminal. Chunking doesn't itself change the language the grammar generates, but it changes the internal structure in a way that can expose useful merges which weren't previously available.
|
|
|
|
We rank candidate grammars by the posterior `P(M | X) ∝ P(M) P(X | M)`, where the prior is a description length and the likelihood integrates over the production probabilities under symmetric Dirichlet priors. The scoring therefore accounts for uncertainty in the production probabilities, rather than relying on a single fitted parameterisation.
|
|
|
|
We explore the resulting grammar space using either beam or best first search, then fit the parameters by expectation maximisation once the grammar structure is fixed. Parsing is a generalised CYK or inside computation over spans, which lets the grammar be scored directly without an intermediate conversion into Chomsky normal form.
|