# scrapeteer

> A web scraper based on puppeteer.

Latest version **1.0.4** (published 2022-10-07) · MIT license · 0 weekly downloads

## Install

```sh
npm install scrapeteer
pnpm add scrapeteer
yarn add scrapeteer
bun add scrapeteer
```

## Health

**Score 35/100 (D)** — status: abandoned.

Positive: has types; no vulnerabilities; high quality score.

Warnings: low downloads; no esm support.

Negative: abandoned.

## Facts

| | |
|---|---|
| Version | 1.0.4 |
| Published | 2022-10-07 |
| First published | 2020-09-17 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | bundled |
| Module format | CommonJS |
| Dependencies | 1 |
| Unpacked size | 14.4 KB |
| Known vulnerabilities | 0 |
| Install scripts | no |
| GitHub stars | 4 |
| Author | Adalberto Aucar Kutuxidis |
| Maintainers | frypizzabox |
| Keywords | scraper, puppeteer, typescript, nodejs, web, page, html, parser |

## Links

- npm: https://www.npmjs.com/package/scrapeteer
- Repository: https://github.com/frypizzabox/Scrapeteer
- Homepage: https://github.com/frypizzabox/Scrapeteer#readme
- Issues: https://github.com/frypizzabox/Scrapeteer/issues
- npm.io page: https://npm.io/package/scrapeteer

## Dependencies (1)

- [puppeteer](https://npm.io/package/puppeteer.md) ^18.2.1

## Alternatives

- [@tsparticles/shape-image](https://npm.io/package/@tsparticles/shape-image.md) — 303.7K weekly downloads
- [@tsparticles/shape-line](https://npm.io/package/@tsparticles/shape-line.md) — 233.7K weekly downloads
- [stringify-attributes](https://npm.io/package/stringify-attributes.md) — 58.6K weekly downloads
- [mobile-drag-drop](https://npm.io/package/mobile-drag-drop.md) — 46.3K weekly downloads
- [@comunica/actor-rdf-parse-html](https://npm.io/package/@comunica/actor-rdf-parse-html.md) — 29.2K weekly downloads

## Recent versions

- 1.0.4 (latest) — 2022-10-07
- 1.0.3 — 2020-09-21
- 1.0.2 — 2020-09-18
- 1.0.1 — 2020-09-18
- 0.0.6 — 2020-09-17
- 0.0.5 — 2020-09-17

## README

# Scrapeteer

Scrapeteer is a NodeJS library for web-scraping using [Puppeteer](https://www.npmjs.com/package/puppeteer).

## Installation

Use the package manager [npm](https://www.npmjs.com/) to install Scrapeteer.

```bash
npm install scrapeteer
```

## Usage

Scrapeteer will extract data from an web-page using an array of objects where which one will represent rules of how to fetch and store data in the returning object.

```javascript
import Scrapeteer from 'scrapeteer';

const defaultTagsConfig = [
  {
    saveAs: 'title',
    query: 'title',
    attrsToFetch: ['innerHTML'],
  },
];

const onRequestTagsConfig = [
  {
    saveAs: 'title',
    query: 'title',
    attrsToFetch: ['innerHTML'],
  },
  {
    saveAs: 'image',
    query: 'meta[name="og:image"], meta[property="og:image"]',
    attrsToFetch: ['content'],
  },
];

const scrapeteer = new Scrapeteer(defaultTagsConfig);

(async () => {
  await scrapeteer.launch();

  const withDefault = await scrapeteer.extractFromUrl('http://amazon.com');
  const withOnRequest = await scrapeteer.extractFromUrl(
    'http://amazon.com',
    onRequestTagsConfig
  );

  console.log('With default tags config: ', withDefault);
  console.log('With OnRequest tags config: ', withOnRequest);
})();
```

To prepare Scrapeteer for use is necessary to follow a sequence of steps:

1. Load Scrapeteer Class.
2. Instantiate it. Is possible to set predefined tag rules here, but is optional.
3. Call the `async` method `.start()` which will load puppeteer with optimized settings.

### Tags Config:

```javascript
const tagsConfig = [
  {
    saveAs: 'title',
    query: 'meta[name="og:title"], meta[property="og:title"]',
    attrsToFetch: ['content'],
  },
  {
    saveAs: 'image',
    query: 'meta[name="og:image"], meta[property="og:image"]',
    attrsToFetch: ['content'],
  },
];
```

An object with tag rules are defined by three parameters:

1. **saveAs:** A string that represents the key in the returning object.

2. **query:** A string that will be used in a [document.querySelectorAll](https://developer.mozilla.org/en-US/docs/Web/API/Document/querySelectorAll) during the scraping.

3. **attrsToFetch:** A array containing all the attributes to fetch from the html tags resulting from the query.

Scrapeteer will use what's defined in the query to fetch html tags, going one by one of the resulting list and get the value from each attribute requested in attrsToFetch. After that it will store everything that was found into an array of strings and save in the resulting object by a key named as the saveAs value.

It goes by the sequence: `Fetch a tag by query -> Get values from Attributes -> Add to resulting object as defined`

You execute the `async` method `.extractFromUrl` to extract the data from a url. A set of tag rules can be used as a second parameter which will be used stead of the default one.

```Javascript
const withDefault = await scrapeteer.extractFromUrl('http://amazon.com') // Goes with default
const withOnRequest = await scrapeteer.extractFromUrl('http://amazon.com', onRequestTagsConfig)
```

The return will be an object with your scraped data. In case of no set of tag rules defined, the return will be an empty object.

```javascript
{
 "title": [
   "Here it comes an result",
   "here it comes another result",
 ],
 "image": [
   "http://imagine_a_image_here.png",
   "http://imagine_another_image_here.png",
   "http://imagine_one_more_image_here.png",
 ],
 "author": [] // empty array means no data was extracted from this set of tag rules
}
```

## Contributing

Pull requests are welcome. For major changes, please open an issue first to discuss what you would like to change.

Please make sure to update tests as appropriate.

## License

[MIT](https://choosealicense.com/licenses/mit/)

---
_Source: https://npm.io/package/scrapeteer · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
