# doken

> A minimalistic, general purpose tokenizer generator

Latest version **1.0.2** (published 2020-06-19) · MIT license · 0 weekly downloads

## Install

```sh
npm install doken
pnpm add doken
yarn add doken
bun add doken
```

## Health

**Score 25/100 (F)** — status: abandoned.

Positive: has types; no vulnerabilities; high quality score.

Warnings: low downloads; no esm support.

Negative: abandoned; low maintenance score.

## Facts

| | |
|---|---|
| Version | 1.0.2 |
| Published | 2020-06-19 |
| First published | 2020-01-28 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | bundled |
| Module format | CommonJS |
| Dependencies | 0 |
| Unpacked size | 18.9 KB |
| Known vulnerabilities | 0 |
| Install scripts | no |
| GitHub stars | 3 |
| Author | Yichuan Shen |
| Maintainers | yishn |
| Keywords | lexer, tokenizer, parser, generator |

## Links

- npm: https://www.npmjs.com/package/doken
- Repository: https://github.com/yishn/doken
- Homepage: https://github.com/yishn/doken#readme
- Issues: https://github.com/yishn/doken/issues
- npm.io page: https://npm.io/package/doken

## Alternatives

- [babylon](https://npm.io/package/babylon.md) — 5.1M weekly downloads
- [csscolorparser](https://npm.io/package/csscolorparser.md) — 3.7M weekly downloads
- [expr-eval-fork](https://npm.io/package/expr-eval-fork.md) — 1.5M weekly downloads
- [@leeoniya/ufuzzy](https://npm.io/package/@leeoniya/ufuzzy.md) — 247.7K weekly downloads
- [xml-parser](https://npm.io/package/xml-parser.md) — 78.4K weekly downloads

## Recent versions

- 1.0.2 (latest) — 2020-06-19
- 1.0.1 — 2020-04-08
- 1.0.0 — 2020-02-02
- 0.0.8-alpha — 2020-01-30
- 0.0.7-alpha — 2020-01-29
- 0.0.6-alpha — 2020-01-29
- 0.0.5-alpha — 2020-01-29
- 0.0.4-alpha — 2020-01-29
- 0.0.3-alpha — 2020-01-29
- 0.0.2-alpha — 2020-01-29
- 0.0.1-alpha — 2020-01-28

## README

# doken [![CI Status](https://github.com/yishn/doken/workflows/CI/badge.svg?branch=master)](https://github.com/yishn/doken/actions)

A minimalistic, general purpose tokenizer generator.

## Usage

Use npm to install:

```
$ npm install doken
```

Import doken and create a tokenizer with regular expression rules:

```js
const {createTokenizer, regexRule} = require('doken')

const tokenizeJSON = createTokenizer({
  rules: [
    regexRule('_whitespace', /\s+/y, {lineBreaks: true}),
    regexRule('brace', /[{}]/y),
    regexRule('bracket', /[\[\]]/y),
    regexRule('colon', /:/y),
    regexRule('comma', /,/y),
    regexRule('string', /"([^"\n\\]|\\[^\n])*"/y),
    regexRule('number', /(-|\+)?\d+(.\d+)?/y),
    regexRule('boolean', /(true|false)\b/y),
    regexRule('null', /null\b/y)
  ]
})

let tokens = tokenizeJSON(`{"a": "Hello World!"}`)

console.log([...tokens])
```

## API

### Rule object

A rule object contains the following fields:

- `type` `<string>` - The type of the token this rule generates. If `type`
  starts with an underscore `_`, the token will not be emitted by the tokenizer.
- `lineBreaks` `<boolean>` _(Optional)_ - Set this property to `true` if this
  rule might match line breaks and you want to track it correctly.
- `match` `<Function>` - A function with the following signature:

  <!-- prettier-ignore -->
  ```ts
  (input: string, position: number) =>
    null |
    {
      length: number,
      value: any
    }
  ```

  This function will try to get the token of given `type` at given `position` in
  `input` if applicable. Return `null` if `input` at `position` is not a token
  with the given `type`, otherwise return an object.

  The first `length` characters of `input` after `position` will be the matched
  token. You can optionally return a `value` which can contain any data that
  will be attached to the token. If `value` is not given, it will default to the
  first `length` characters of `input` after `position`.

### Token object

A token will be represented by an object with the following fields:

- `type` `<string>` | `null` - The type of the token or `null` if no given rules
  match the input.
- `length` `<number>` - The length of the token.
- `value` `<any>` - The `value` generated by the rule.
- `pos` `<number>` - The zero-based position of the first character of the
  token.
- `row` `<number>` - The zero-based row of the first character of the token.
- `col` `<number>` - The zero-based column of the first character of the token.

### `doken.createTokenizer(options)`

- `options` `<object>`
  - `rules` [`<Array<Rule>>`](#rule-object)
  - `strategy` `'first' | 'longest'` _(Optional)_ - Default: `'first'`
- Returns: `<Function>`

Generates a tokenize function with the following signature:

```ts
(input: string) => IterableIterator<Token>
```

This function will attempt to tokenize given `input`, yielding
[tokens](#token-object) matched by given `rules` one by one.

Set `strategy` to `'longest'` to match the token with the rule that matches the
most characters instead of using the rule that matches first.

### `doken.regexRule(type, regex[, options])`

- `type` `<string>`
- `regex`
  [`<RegExp>`](https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/RegExp)
- `options` `<object>`
  - `lineBreaks` `<boolean>` _(Optional)_ - Set this property to `true` if this
    rule might match line breaks and you want to track it correctly.
  - `value` `<Function>` _(Optional)_ - A function for calculating the token
    value out of the match.
  - `condition` `<Function>` _(Optional)_ - A function for indicating whether to
    discard match or not.
- Returns: [`<Rule>`](#rule-object)

Returns a [rule](#rule-object) that attempts to match input string with the
given `regex`.

`value` can be set to a function `(match: RegExpExecArray) => any`. The
generated token will have the returned value as `value`.

`condition` can be set to a function `(match: RegExpExecArray) => boolean`.
Return `false` to indicate to discard matched result and go on with the next
rule.

---
_Source: https://npm.io/package/doken · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
