# chardet

> Character encoding detector

Latest version **2.2.0** (published 2026-06-20) · MIT license · 0 weekly downloads

## Install

```sh
npm install chardet
pnpm add chardet
yarn add chardet
bun add chardet
```

## Health

**Score 65/100 (B)** — status: active.

Positive: has types; no vulnerabilities; has provenance; high maintenance score; high quality score.

Warnings: low downloads; no esm support.

## Facts

| | |
|---|---|
| Version | 2.2.0 |
| Published | 2026-06-20 |
| First published | 2013-04-30 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | bundled |
| Module format | CommonJS |
| Dependencies | 0 |
| Unpacked size | 174.8 KB |
| Known vulnerabilities | 0 |
| Install scripts | no |
| Provenance | attested (GitHub Actions) |
| GitHub stars | 311 |
| Author | Dmitry Shirokov |
| Maintainers | runk |
| Keywords | encoding, character, utf8, detector, chardet, icu, character detection, character encoding, language, iconv, iconv-light, UTF-8, UTF-16, UTF-32, ISO-2022-JP, ISO-2022-KR, ISO-2022-CN, Shift_JIS, Big5, EUC-JP, EUC-KR, GB18030, ISO-8859-1, ISO-8859-2, ISO-8859-5, ISO-8859-6, ISO-8859-7, ISO-8859-8, ISO-8859-9, windows-1250, windows-1251, windows-1252, windows-1253, windows-1254, windows-1255, windows-1256, windows-1257, windows-1258, windows-874, TIS-620, KOI8-R |

## Links

- npm: https://www.npmjs.com/package/chardet
- Repository: https://github.com/runk/node-chardet
- Issues: http://github.com/runk/node-chardet/issues
- npm.io page: https://npm.io/package/chardet

## Alternatives

- [@opentelemetry/exporter-zipkin](https://npm.io/package/@opentelemetry/exporter-zipkin.md) — 14.8M weekly downloads
- [pusher-js](https://npm.io/package/pusher-js.md) — 2.0M weekly downloads
- [browserify](https://npm.io/package/browserify.md) — 1.7M weekly downloads
- [sqs-consumer](https://npm.io/package/sqs-consumer.md) — 1.7M weekly downloads
- [@sanity/eventsource](https://npm.io/package/@sanity/eventsource.md) — 930.8K weekly downloads

## Recent versions

- 2.2.0 (latest) — 2026-06-20
- 2.1.1 — 2025-10-29
- 2.1.0 — 2025-02-24
- 2.0.0 — 2023-09-28
- 1.6.1 — 2023-09-28
- 1.6.0 — 2023-06-16
- 1.5.1 — 2023-01-05
- 1.5.0 — 2022-10-09
- 1.4.0 — 2021-10-19
- 1.3.0 — 2020-09-25
- 1.2.2 — 2020-09-23
- 1.2.1 — 2020-07-06
- 1.2.0 — 2020-07-02
- 1.1.0 — 2020-05-07
- 1.0.0 — 2020-03-31
- … 15 more at https://npm.io/package/chardet/versions

## README

# chardet

_Chardet_ is a character detection module written in pure JavaScript (TypeScript). Module uses occurrence analysis to determine the most probable encoding.

- Packed size is only **22 KB**
- Works in all environments: Node / Browser / Native
- Works on all platforms: Linux / Mac / Windows
- No dependencies
- No native code / bindings
- 100% written in TypeScript
- Extensive code coverage

## Installation

```
npm i chardet
```

## Usage

To return the encoding with the highest confidence:

```javascript
import chardet from 'chardet';

const encoding = chardet.detect(Buffer.from('hello there!'));
// or
const encoding = await chardet.detectFile('/path/to/file');
// or
const encoding = chardet.detectFileSync('/path/to/file');
```

To return the full list of possible encodings use `analyse` method.

```javascript
import chardet from 'chardet';
chardet.analyse(Buffer.from('hello there!'));
```

Returned value is an array of objects sorted by confidence value in descending order

```javascript
[
  { confidence: 90, name: 'UTF-8' },
  { confidence: 20, name: 'windows-1252', lang: 'fr' },
];
```

In browser, you can use [Uint8Array](https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/Uint8Array) instead of the `Buffer`:

```javascript
import chardet from 'chardet';
chardet.analyse(new Uint8Array([0x68, 0x65, 0x6c, 0x6c, 0x6f]));
```

## Working with large data sets

Sometimes, when data set is huge and you want to optimize performance (with a trade off of less accuracy),
you can sample only the first N bytes of the buffer:

```javascript
const encoding = await chardet.detectFile('/path/to/file', { sampleSize: 32 });
```

You can also specify where to begin reading from in the buffer:

```javascript
const encoding = await chardet.detectFile('/path/to/file', {
  sampleSize: 32,
  offset: 128,
});
```

## Working with strings

In both Node.js and browsers, all strings in memory are represented in UTF-16 encoding. This is a fundamental aspect of the JavaScript language specification. Therefore, you cannot use plain strings directly as input for `chardet.analyse()` or `chardet.detect()`. Instead, you need the original string data in the form of a Buffer or Uint8Array.

In other words, if you receive a piece of data over the network and want to detect its encoding, use the original data payload, not its string representation. By the time you convert data to a string, it will be in UTF-16 encoding.

Note on [TextEncoder](https://developer.mozilla.org/en-US/docs/Web/API/TextEncoder/TextEncoder): By default, it returns a UTF-8 encoded buffer, which means the buffer will not be in the original encoding of the string.

## Supported Encodings:

- UTF-8
- UTF-16 LE
- UTF-16 BE
- UTF-32 LE
- UTF-32 BE
- ISO-2022-JP
- ISO-2022-KR
- ISO-2022-CN
- Shift_JIS
- Big5
- EUC-JP
- EUC-KR
- GB18030
- ISO-8859-1
- ISO-8859-2
- ISO-8859-5
- ISO-8859-6
- ISO-8859-7
- ISO-8859-8
- ISO-8859-9
- windows-1250
- windows-1251
- windows-1252
- windows-1253
- windows-1254
- windows-1255
- windows-1256
- windows-1257
- windows-1258
- windows-874 (TIS-620 / ISO-8859-11)
- KOI8-R

Currently only these encodings are supported.

## TypeScript?

Yes. Type definitions are included.

### References

- http://site.icu-project.org/
- https://github.com/TypesettingTools/uchardet

### TODO
- [ ] KOI8-U for Ukrainian
- [ ] IBM866 for DOS Cyrillic
- [ ] macintosh and x-mac-cyrillic
- [ ] CP949 / UHC support within the Korean recognizer
- [ ] ISO-8859-3, ISO-8859-4
- [ ] ISO-8859-10, ISO-8859-13
- [ ] ISO-8859-14, ISO-8859-15, ISO-8859-16
- [ ] DOS CP850, CP852, CP855

---
_Source: https://npm.io/package/chardet · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
