# @jscpd/tokenizer

> tokenizer of source code for jscpd

Latest version **4.2.6** (published 2026-08-13) · MIT license · 0 weekly downloads

## Install

```sh
npm install @jscpd/tokenizer
pnpm add @jscpd/tokenizer
yarn add @jscpd/tokenizer
bun add @jscpd/tokenizer
```

## Health

**Score 70/100 (B)** — status: active.

Positive: has types; esm support; no vulnerabilities; recently updated; high maintenance score; high quality score.

Warnings: low downloads.

## Facts

| | |
|---|---|
| Version | 4.2.6 |
| Published | 2026-08-13 |
| First published | 2020-04-29 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | bundled |
| Module format | ESM + CommonJS |
| Dependencies | 2 |
| Unpacked size | 2.4 MB |
| Known vulnerabilities | 0 |
| Install scripts | no |
| GitHub stars | 6334 |
| Author | Andrey Kucherenko |
| Maintainers | apk |

## Links

- npm: https://www.npmjs.com/package/@jscpd/tokenizer
- Repository: https://github.com/kucherenko/jscpd
- Homepage: https://jscpd.dev
- Issues: https://github.com/kucherenko/jscpd/issues
- npm.io page: https://npm.io/package/@jscpd/tokenizer

## Dependencies (2)

- [spark-md5](https://npm.io/package/spark-md5.md) ^3.0.2
- [@jscpd/core](https://npm.io/package/@jscpd/core.md) 4.2.5

## Recent versions

- 4.2.6 (latest) — 2026-08-13
- 4.0.0 (rc) — 2024-05-26
- 3.3.0-alpha.8 (canary) — 2020-04-29
- 4.2.5 — 2026-06-07
- 4.2.4 — 2026-05-25
- 4.2.3 — 2026-05-17
- 4.2.2 — 2026-05-15
- 4.2.1 — 2026-05-15
- 4.2.0 — 2026-05-14
- 4.1.1 — 2026-05-12
- 4.1.0 — 2026-05-09
- 4.0.5 — 2026-04-10
- 4.0.4 — 2026-01-30
- 4.0.3 — 2026-01-11
- 4.0.2 — 2026-01-11
- … 23 more at https://npm.io/package/@jscpd/tokenizer/versions

## README

# `@jscpd/tokenizer`

> Tokenizer package for [@jscpd](https://github.com/kucherenko/jscpd) — converts source code into a list of tokens for duplicate detection.

Supports **223 programming languages and formats** via a self-contained [reprism](https://github.com/tannerlinsley/reprism)-based grammar engine. Grammars are loaded lazily for fast startup, with O(n) hot paths for high-throughput scanning.

Special tokenization modes handle multi-language files:

- **Vue SFC** (`.vue`) — `<template>`, `<script>`, and `<style>` blocks each tokenized by their own language
- **Svelte** (`.svelte`) — per-block tokenization for HTML, JS, and CSS sections
- **Astro** (`.astro`) — frontmatter and template blocks tokenized independently
- **Markdown** (`.md`) — fenced code blocks tokenized by the declared language

This enables cross-format clone detection: a `<script lang="ts">` block in a `.vue` file can match a plain `.ts` file.

## Installation

```bash
npm install @jscpd/tokenizer --save
```

## Usage

```typescript
import { IOptions, ITokensMap } from '@jscpd/core';
import { Tokenizer } from '@jscpd/tokenizer';

const tokenizer = new Tokenizer();
const options: IOptions = {};

const maps: ITokensMap[] = tokenizer.generateMaps('source_id', 'let a = "11"', 'javascript', options);
```

## Supported formats

The full list of 223 supported formats is available in [FORMATS.md](../../FORMATS.md) at the repository root, or at runtime:

```bash
jscpd --list
```


## License

[MIT](LICENSE) © Andrey Kucherenko

---
_Source: https://npm.io/package/@jscpd/tokenizer · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
