t.on('token', function(token, type) {
    // do something useful
    // type is the type of the token (specified with addRule)
    // token is the actual matching string
})
// alternatively you can listen on the 'data' event

look out for the end

t.on('end', callback);

the optional callback argument for the constructor is a function that will be called for each token in order to specify a different type by returning a string. The parameters passed to the function are token(the token that we found) and match, an object like this

{
    regex: /whatever/ // the regex that matched the token
    type: 'type' // the type of the token
}

Have a look in the example folder

Rules

rules are regular expressions associated with a type name. The tokenizer tries to find the longest string matching one or more rules. When several rules match the same string, priority is given to the rule which was added first. (this may change)

Please note that your regular expressions should use ^ and $ in order to test the whole string. If these are not used, you rule will match every string that contains what you specified, this could be the whole file!

To do

a lot of optimisation
being able to share rules across several tokenizers (although this can be achieved through inheritance)
probably more hooks
more checking

FAQs

What is tokenizer?

Is tokenizer well maintained?

Package last updated on 14 Feb 2013

Did you know?

Socket for GitHub automatically highlights issues in each pull request and monitors the health of all your open source dependencies. Discover the contents of your packages and block harmful activity before you install or update your dependencies.

Install

tokenizer

Synopsis

How to

Rules

To do

Related posts

require(esm) Backported to Node.js 20, Paving the Way for ESM-Only Packages

PyPI Now Supports iOS and Android Wheels for Mobile Python Development