Huge News!Announcing our $40M Series B led by Abstract Ventures.Learn More →

parse-latin

Package Overview

Dependencies

Advanced tools

Install Socket

Detect and block malicious and high-risk dependencies

Install

parse-latin

Latin-script (natural language) parser

6.0.0
Source
npm

Version published: 2 years ago

Weekly downloads: 468K; increased by2.2%

Maintainers: 1

Weekly downloads

Created: 10 years ago

What is parse-latin?

The parse-latin npm package is a JavaScript library used to parse Latin-script natural language into a syntax tree. It is particularly useful for text processing tasks such as tokenization, sentence splitting, and word segmentation.

What are parse-latin's main functionalities?

Tokenization

This feature allows you to tokenize a given text into individual tokens (words, punctuation, etc.). The code sample demonstrates how to tokenize a simple sentence.

const ParseLatin = require('parse-latin');
const parser = new ParseLatin();
const tokens = parser.tokenize('This is a sentence.');
console.log(tokens);

Sentence Splitting

This feature enables you to split a paragraph into individual sentences. The code sample shows how to split a paragraph into separate sentences.

const ParseLatin = require('parse-latin');
const parser = new ParseLatin();
const sentences = parser.tokenizeParagraph('This is a sentence. This is another sentence.');
console.log(sentences);

Word Segmentation

This feature allows you to segment a sentence into individual words. The code sample demonstrates how to segment a sentence into words.

const ParseLatin = require('parse-latin');
const parser = new ParseLatin();
const words = parser.tokenizeWords('This is a sentence.');
console.log(words);

Other packages similar to parse-latin

parse-latin

A natural language parser, for Latin-script languages, that produces nlcst.

What is this?

This package exposes a parser that takes Latin-script natural language and produces a syntax tree.

When should I use this?

If you want to handle natural language as syntax trees manually, use this.

Alternatively, you can use the retext plugin retext-latin, which wraps this project to also parse natural language at a higher-level (easier) abstraction.

Whether Old-English (“þā gewearþ þǣm hlāforde and þǣm hȳrigmannum wiþ ānum penninge”), Icelandic (“Hvað er að frétta”), French (“Où sont les toilettes?”), this project does a good job at tokenizing it.

For English and Dutch, you can instead use parse-english and parse-dutch.

You can somewhat use this for Latin-like scripts, such as Cyrillic (“Добро пожаловать!”), Georgian (“როგორა ხარ?”), Armenian (“Շատ հաճելի է”), and such.

Install

This package is ESM only. In Node.js (version 14.14+, 16.0+), install with npm:

npm install parse-latin

In Deno with esm.sh:

import {ParseLatin} from 'https://esm.sh/parse-latin@6'

In browsers with esm.sh:

<script type="module">
  import {ParseLatin} from 'https://esm.sh/parse-latin@6?bundle'
</script>

Use

import {inspect} from 'unist-util-inspect'
import {ParseLatin} from 'parse-latin'

const tree = new ParseLatin().parse('A simple sentence.')

console.log(inspect(tree))

Yields:

RootNode[1] (1:1-1:19, 0-18)
└─0 ParagraphNode[1] (1:1-1:19, 0-18)
    └─0 SentenceNode[6] (1:1-1:19, 0-18)
        ├─0 WordNode[1] (1:1-1:2, 0-1)
        │   └─0 TextNode "A" (1:1-1:2, 0-1)
        ├─1 WhiteSpaceNode " " (1:2-1:3, 1-2)
        ├─2 WordNode[1] (1:3-1:9, 2-8)
        │   └─0 TextNode "simple" (1:3-1:9, 2-8)
        ├─3 WhiteSpaceNode " " (1:9-1:10, 8-9)
        ├─4 WordNode[1] (1:10-1:18, 9-17)
        │   └─0 TextNode "sentence" (1:10-1:18, 9-17)
        └─5 PunctuationNode "." (1:18-1:19, 17-18)

API

This package exports the identifier ParseLatin. There is no default export.

`ParseLatin()`

Create a new parser.

`ParseLatin#parse(value)`

Turn natural language into a syntax tree.

Parameters

`value`

Value to parse (string).

Returns

RootNode.

Algorithm

👉 Note: The easiest way to see how parse-latin parses, is by using the online parser demo, which shows the syntax tree corresponding to the typed text.

parse-latin splits text into white space, punctuation, symbol, and word tokens:

“word” is one or more unicode letters or numbers
“white space” is one or more unicode white space characters
“punctuation” is one or more unicode punctuation characters
“symbol” is one or more of anything else

Then, it manipulates and merges those tokens into a syntax tree, adding sentences and paragraphs where needed.

some punctuation marks are part of the word they occur in, such as non-profit, she’s, G.I., 11:00, N/A, &c, nineteenth- and…
some periods do not mark a sentence end, such as 1., e.g., id.
although periods, question marks, and exclamation marks (sometimes) end a sentence, that end might not occur directly after the mark, such as .), ."
…and many more exceptions

Types

This package is fully typed with TypeScript. It exports no additional types.

Compatibility

This package is at least compatible with all maintained versions of Node.js. As of now, that is Node.js 14.14+ and 16.0+. It also works in Deno and modern browsers.

parse-english — English (natural language) parser
parse-dutch — Dutch (natural language) parser

Contribute

Yes please! See How to Contribute to Open Source.

Security

This package is safe.

License

Keywords

FAQs

What is parse-latin?

Is parse-latin popular?

Is parse-latin well maintained?

Package last updated on 11 Nov 2022

Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

parse-latin

What is parse-latin?

What are parse-latin's main functionalities?

Other packages similar to parse-latin

parse-latin

Contents

What is this?

When should I use this?

Install

Use

API

`ParseLatin()`

`ParseLatin#parse(value)`

Parameters

`value`

Returns

Algorithm

Types

Compatibility

Contribute

Security

License

Keywords

Related posts

parse-latin

What is parse-latin?

What are parse-latin's main functionalities?

Other packages similar to parse-latin

compromise

natural

parse-latin

Contents

What is this?

When should I use this?

Install

Use

API

ParseLatin()

ParseLatin#parse(value)

Parameters

value

Returns

Algorithm

Types

Compatibility

Related

Contribute

Security

License

Keywords

Related posts

Massive npm Malware Campaign Leverages Ethereum Smart Contracts To Evade Detection and Maintain Control

Author Typosquatting on npm: Attackers Impersonate Sindre Sorhus with Malicious ‘chalk-node’ Package

Supply Chain Attack on LottieFiles Player Caused by Compromised npmjs Credentials

`ParseLatin()`

`ParseLatin#parse(value)`

`value`