<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>nlp &middot; Arguable Intelligence</title><link>https://ojmason.net/tags/nlp/</link><description>Personal musings on computational linguistics, AI, sailing, and generated fiction.</description><language>en-gb</language><managingEditor>Oliver Mason</managingEditor><webMaster>Oliver Mason</webMaster><lastBuildDate>Mon, 10 Apr 2017 11:27:24 +0100</lastBuildDate><atom:link href="https://ojmason.net/tags/nlp/index.xml" rel="self" type="application/rss+xml"/><item><title>Toki Pona</title><link>https://ojmason.net/toki-pona/</link><guid isPermaLink="true">https://ojmason.net/toki-pona/</guid><pubDate>Mon, 10 Apr 2017 11:27:24 +0100</pubDate><category>toki pona</category><category>conlang</category><category>nlp</category><description>Since working in computational linguistics I have been interested in constructed languages (&lsquo;conlangs&rsquo;). When trying to process any natural language with computer programs, you constantly run into inconvenient exceptions, in morphology, syntax, etc. A conlang such as Esperanto promises a great simplification: as it is completely regular, NLP software is likely to be much simpler and less complex. However, given that Esperanto is a full-scale language, it&rsquo;s still not trivial to work with. It&rsquo;s got a large vocabulary, and the syntax is not that easy to parse.</description><content:encoded><![CDATA[<p>Since working in computational linguistics I have been interested in constructed languages (&lsquo;conlangs&rsquo;).
When trying to process any natural language with computer programs, you constantly run into inconvenient
exceptions, in morphology, syntax, etc. A conlang such as <a href="https://en.wikipedia.org/wiki/Esperanto">Esperanto</a>
promises a great simplification: as it is completely regular, NLP software is likely to be much simpler and
less complex. However, given that Esperanto is a full-scale language, it&rsquo;s still not trivial to work with. It&rsquo;s
got a large vocabulary, and the syntax is not that easy to parse.</p>
<p>In a post on the <a href="http://esperanto.stackexchange.com/">Esperanto Language Stack Exchange</a> I then heard of another
conlang, <a href="https://en.wikipedia.org/wiki/Toki_Pona">Toki Pona</a>. This is billed as a <em>minimalist</em> language: it only
has about 120 words, and a very fixed and simplistic sentence structure. Unlike Esperanto, it is not really a
proper language for everyday use, but more of a philosophical experiment. How does your language influence the
way you think about the world? This is kind of related to the <a href="https://en.wikipedia.org/wiki/Linguistic_relativity">Sapir-Whorf Hypothesis</a>.
When you have to limit yourself to describing/paraphrasing everything with a limited vocabulary, you need to
reflect more about what it is you&rsquo;re talking about. For example, a friend is a <em>jan pona</em> (&lsquo;good person&rsquo;), and a
bad person is a <em>jan ike</em>. So how do you say &lsquo;bad friend&rsquo;? You are not able to say someone is both good and bad
at the same time. So a friend cannot be a bad person, or vice versa.</p>
<p>Toki Pona arguably is a toy language only. There is just one word for <em>fruit</em>, <em>vegetable</em>, etc.: <em>kili</em>. If you want
to say <em>banana</em>, you say <em>kili jelo</em> (&ldquo;yellow fruit&rdquo;). If you want to talk about lemons as well, you&rsquo;re out of luck.
It&rsquo;s very context dependent, and definitely not useful for a scientific treatise, or even recipes. But it&rsquo;s fine
for basic stories, myths and legends, and so on. And, most importantly, it&rsquo;s easy to learn. No morphology. Very
limited syntax. Small vocabulary. The main difficulty is to express yourself given those limited means. But
we grow only when challenged!</p>
<p>I&rsquo;m interested in doing NLP with Toki Pona, as it is so limited. It should be possible to quickly get to the semantic
or pragmatic levels, as morphology and syntax will be dealt with easily. Analysing and generating sentences should
be extremely easy. More on that as it materialises.</p>
<p>If you&rsquo;re interested in languages, and what to explore the way you express meaning, give Toki Pona a try. There
are various on-line resources available, plus a text book <a href="http://amzn.to/2oN4Zo7"><em>Toki Pona &ndash; The Language of Good</em></a>.</p>
<p>I&rsquo;ll post more on this at a later time&hellip;</p>
]]></content:encoded></item><item><title>Replacing a stack with concurrency</title><link>https://ojmason.net/replacing-a-stack-with-concurrency/</link><guid isPermaLink="true">https://ojmason.net/replacing-a-stack-with-concurrency/</guid><pubDate>Wed, 23 Apr 2008 00:00:00 +0000</pubDate><category>nlp</category><category>erlang</category><description>For some language processing task I needed a reasonably powerful parser (a program to identify the syntactic structure of a sentence). So I dug out my copy of Winograd (1983) (Language as a Cognitive Process) and set about implementing an Augmented Transition Network parser in Erlang.
Now, the first thing you learn about natural language is that it is full of ambiguities, and so there will always be several alternatives available, several possible paths through the network which defines the grammar. The traditional solution is to dump all the alternatives on a stack, and look at them when the current path has been finished with. You can either go depth-first, where you complete the current path before you get the next one off the stack, or breadth-first, where you advance all paths by one step at a time, kind of pseudo-parallel.</description><content:encoded><![CDATA[<p>For some language processing task I needed a reasonably powerful
parser (a program to identify the syntactic structure of a sentence).
So I dug out my copy of Winograd (1983) (<em>Language as a Cognitive
Process</em>) and set about implementing an <a href="https://en.wikipedia.org/wiki/Augmented_transition_network">Augmented Transition Network</a>
parser in Erlang.</p>
<p>Now, the first thing you learn about natural language is that it
is full of ambiguities, and so there will always be several
alternatives available, several possible paths through the network
which defines the grammar. The traditional solution is to dump all
the alternatives on a stack, and look at them when the current path
has been finished with. You can either go depth-first, where you
complete the current path before you get the next one off the stack,
or breadth-first, where you advance all paths by one step at a time,
kind of pseudo-parallel.</p>
<p>Having to deal with a stack is tedious, as you need to keep track
of the current configuration: which network are you at, what node,
what position in the sentence, etc. But then, it occurred to me,
there’s an easier way to do it (at least it’s easier in Erlang!):
every time you come to a point where you have multiple alternatives,
you spawn a new process and pursue all of them in parallel.</p>
<p>The only overhead you need is a loop which keeps track of all the
processes currently running. This loop receives the results of
successful paths, and gets notified of unsuccessful ones (where the
process terminates without having found a valid structure). No need
for a stack, and hopefully very efficient processing on multi-core
machines as a free side-effect.</p>
<p>I’m still amazed how easy it was to implement. I wouldn’t have
fancied doing that in Java or even C. For my test sentences I had
about 8 to 10 processes running in parallel most of the time, but
it depends on the size of the grammar and the length of the sentence
really. What I liked about this was that it seemed the natural way
to do in Erlang, where working with processes is just so easy.</p>
<p>And also, another nail in the coffin for the claim that you can’t
use Erlang for handling texts easily!</p>
]]></content:encoded></item></channel></rss>