# scrpr

> a scraper multitool

Latest version **0.1.29** (published 2026-01-27) · UNLICENSE license · 0 weekly downloads

## Install

```sh
npm install scrpr
pnpm add scrpr
yarn add scrpr
bun add scrpr
```

## Health

**Score 45/100 (D)** — status: stable.

Positive: no vulnerabilities.

Warnings: low downloads; no types; no esm support; pre 1.0.

## Facts

| | |
|---|---|
| Version | 0.1.29 |
| Published | 2026-01-27 |
| First published | 2021-02-17 |
| Weekly downloads | 0 |
| License | UNLICENSE |
| TypeScript types | none |
| Module format | CommonJS |
| Dependencies | 3 |
| Unpacked size | 24.4 KB |
| Known vulnerabilities | 0 |
| Install scripts | no |
| GitHub stars | 2 |
| Author | yetzt |
| Maintainers | yetzt |
| Keywords | scraper, scrape, ddj, crawler, scraping, data, tool |

## Links

- npm: https://www.npmjs.com/package/scrpr
- Repository: https://github.com/yetzt/node-scrpr
- Homepage: https://github.com/yetzt/node-scrpr#readme
- Issues: https://github.com/yetzt/node-scrpr/issues
- npm.io page: https://npm.io/package/scrpr

## Dependencies (3)

- [quu](https://npm.io/package/quu.md) ^0.5.0
- [needle](https://npm.io/package/needle.md) ^3.3.1
- [mime-types](https://npm.io/package/mime-types.md) ^3.0.2

## Alternatives

- [@tsparticles/shape-image](https://npm.io/package/@tsparticles/shape-image.md) — 303.7K weekly downloads
- [@tsparticles/shape-line](https://npm.io/package/@tsparticles/shape-line.md) — 233.7K weekly downloads
- [stringify-attributes](https://npm.io/package/stringify-attributes.md) — 58.6K weekly downloads
- [mobile-drag-drop](https://npm.io/package/mobile-drag-drop.md) — 46.3K weekly downloads
- [@comunica/actor-rdf-parse-html](https://npm.io/package/@comunica/actor-rdf-parse-html.md) — 29.2K weekly downloads

## Recent versions

- 0.1.29 (latest) — 2026-01-27
- 0.1.28 — 2024-09-30
- 0.1.27 — 2024-07-29
- 0.1.26 — 2024-07-29
- 0.1.25 — 2024-04-08
- 0.1.24 — 2023-09-13
- 0.1.23 — 2023-05-16
- 0.1.21 — 2022-12-04
- 0.1.20 — 2022-12-04
- 0.1.19 — 2022-11-21
- 0.1.18 — 2022-11-15
- 0.1.17 — 2022-11-07
- 0.1.16 — 2022-04-07
- 0.1.15 — 2022-03-29
- 0.1.14 — 2022-01-31
- … 31 more at https://npm.io/package/scrpr/versions

## README

# scrpr

scrpr is a lightweight scraper multitool. it can fetch data via https, detect changes and parse the most common formats.

## Usage Example

```javascript
const scrpr = require("scrpr");

const scraper = scrpr({
	concurrency: 5,
	cachedir: '/tmp/scraper-cache',
});


scraper("https://example.org/data.csv", { 
	parse: "csv", 
}, function(err, change, data){

	if (err) console.error(err);
	if (change) console.log(data);
	
});
```

### `scrpr(opts)` → *function scraper*

Constructor, returns scraper function

Opts:
* `concurrency` — number of parallel requests; default: `1`
* `cachedir` — directory to save cache data in; default: `<root module>/.scrpr-cache`

### `scraper([url], [opts], [callback(err, change, data)])`

Scraper, delivers data

Opts:
* `method` — http method; default: `get`
* `url` — URL, alternative to `url` parameter
* `headers` — additional http request headers, default: `{}`
* `data` — http data to be sent, default: `null`
* `cache` — use cache, default: `true`
* `cacheid` — override cache id, default: `hash(url, opts)`
* `parse` — format to parse, default: `null` (raw data)
* `successCodes` — array of http status codes considered successful, default: `[ 200 ]`
* `needle` — options passed on to `needle`, default `{}`
* `xlsx` — options passed on to `xlsx`, default `{}`
* `xsv` — options passed on to `xsv`, default `{}`
* `pdf` — options passed on to `pdf.js-extract`, default `{}`
* `preprocess(data, callback(err, data))` — modify data before parsing
* `postprocess(data, callback(err, data))` — modify data after parsing
* `stream` — deliver data as `ReadableStream` — no parsing or processing, default: `false`
* `metaredirects` — follow `<meta http-equiv="refresh">` style redirects, default: `false`
* `iconv` — decode stream or data as this charset with iconv-lite before parsing, default: `false`
* `cooldown` — microseconds since last fetch before a resource is fetched again, default: `false`
* `sizechange` — treat unchanged content-length as same file, default: `false`

Callback:
* `err` — contains Error or `null`
* `change` — `true` if data changed
* `data` — raw or parsed data when changed, otherwise status string

## Parsers

* `csv` — Comma Seperated Values; `data` is an Object, parsed with [xsv](https://npmjs.com/package/xsv)
* `tsv` — Tab Separated Values; `data` is an Object, parsed with [xsv](https://npmjs.com/package/xsv)
* `ssv` — Semicolon Separated Values (data has been exported "as csv" with some localizations of Microsoft Excel): `data` is an Object, parsed with [xsv](https://npmjs.com/package/xsv)
* `xml` — eXtensible Markup Language; `data` is an Object, parsed with [xml2js](https://npmjs.com/package/xml2js)
* `json` — JavaScript object Notation; `data` is an Object, parsed natively
* `html` — HyperText Markup Language; `data` is an instance of [cheerio](https://npmjs.com/package/cheerio)
* `yaml` — YAML Ain't Markup Language; `data` is an Object, parsed with [yaml](https://npmjs.com/package/yaml)
* `xlsx` — Office Open XML Workbook; `data` is an Object, parsed with [xlsx](https://npmjs.com/package/xlsx); `{ "<sheetname>": [ [ cell, cell, cell, ... ], ... ] }`
* `pdf` — Portable Document Format; `data` is an Object, parsed with [pdf.js-extract](https://npmjs.com/package/pdf.js-extract);
* `kdl` — KDL Document Language; `data` is an Object, parsed with [kdljs](https://npmjs.com/package/kdljs);
* `dw` — Datawrapper Visualisation; `data` is an Object, extracted with [dataunwrapper](https://npmjs.com/package/dataunwrapper);

## FTP

Rudimentary handling for `ftp` URLs is available if the optional `get-uri` dependency is installed.

## Local Files

Rudimentary handling for local files is available with the `file:/` pseude-protocol.

## Optional dependencies

`xsv`, `xlsx`, `xml2js`, `yaml`, `cheerio`, `dataunwrapper`, `iconv-lite`, `kdljs`, `pdf.js-extract` and `get-uri` are optional dependencies. They should only be installed if their use is required.

## License

[UNLICENSE](UNLICENSE)

---
_Source: https://npm.io/package/scrpr · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
