# node-ckan-crawler

> NodeJS based crawler for CKAN sites

Latest version **0.0.3** (published 2014-07-30) · MIT license · 0 weekly downloads

## Install

```sh
npm install node-ckan-crawler
pnpm add node-ckan-crawler
yarn add node-ckan-crawler
bun add node-ckan-crawler
```

## Health

**Score 15/100 (F)** — status: abandoned.

Positive: no vulnerabilities.

Warnings: low downloads; no types; no esm support; pre 1.0.

Negative: abandoned; low maintenance score.

## Facts

| | |
|---|---|
| Version | 0.0.3 |
| Published | 2014-07-30 |
| First published | 2014-07-25 |
| Weekly downloads | 0 |
| License | MIT |
| TypeScript types | none |
| Module format | CommonJS |
| Dependencies | 2 |
| Known vulnerabilities | 0 (+5 in 1 direct dependencies) |
| Install scripts | no |
| Author | Hafiz Ismail |
| Maintainers | sogko |
| Keywords | crawler, ckan |

## Links

- npm: https://www.npmjs.com/package/node-ckan-crawler
- Repository: https://github.com/sogko/node-ckan-crawler
- Issues: https://github.com/sogko/node-ckan-crawler/issues
- npm.io page: https://npm.io/package/node-ckan-crawler

## Dependencies (2)

- [lodash](https://npm.io/package/lodash.md) ^2.4.1
- [crawler](https://npm.io/package/crawler.md) ^0.2.6

## Recent versions

- 0.0.3 (latest) — 2014-07-30
- 0.0.2 — 2014-07-25
- 0.0.1 — 2014-07-25

## README

node-ckan-crawler
=================

A simple and fast NodeJS based crawler for sites powered by CKAN [http://ckan.org](http://ckan.org)

* Uses the CKAN ```package_search``` Action.Get API to crawl packages / datasets


## Install 
```
npm install node-ckan-crawler
```

## Usage

```
var CKANCrawler = require('node-ckan-crawler');

var crawler = new CKANCrawler();

crawler.queueSite('http://datahub.io/');
crawler.on('content', function(response, content){
  console.log('content', response.uri, content.length);
});
```

### More examples
See more examples found in ```examples\```

## API

### Events
#### Event: 'content'
When response received from the site has been parsed and results ready for consumption

```response``` an [http.IncomingMessage](http://nodejs.org/api/http.html#http_http_incomingmessage) object returned from [mikeal's request()](https://github.com/mikeal/request)

```body``` a JSON object of the response.body

````
crawler.on('content', function(response, body) {
    ...
});

````

#### Event: 'beforeQueue'
When next link is ready to be added to the crawler queue.
Return a non-true value to skip the link

```url``` a string of the next link ready to be added to the crawler queue

```next``` a callback function

````
crawler.on('beforeQueue', function(url, next) {
    next(true); // to add the link to the queue
    // next(false) // to skip link
});

````


#### Event: 'queued'
After a link was added to the crawler queue.

```url``` a string of the next link ready to be added to the crawler queue

````
crawler.on('queued', function(url) {
    ...
});

````

#### Event: 'drain'
When crawler has drained its queue and has no more links to crawl

````
crawler.on('drain', function() {
    ...
});

````
#### Event: 'error'
When an error has occurred

````
crawler.on('error', function(err) {
    ...
});

````

### Methods

#### queueSite(url)
Queue a CKAN powered site by specifying its base API url 

Example: 
``` 
crawler.queueSite('http://datahub.io')
```

## Known Issues



## Credits

* [Hafiz Ismail](https://github.com/sogko) 

## Links
* [wehavefaces.net](http://wehavefaces.net)
* [twitter.com/sogko](https://twitter.com/sogko)
* [github.com/sogko](https://github.com/sogko)
* [medium.com/@sogko](https://medium.com/@sogko)

## License
Copyright (c) 2014 Hafiz Ismail. This software is licensed under the [MIT License](https://github.com/sogko/node-ckan-crawler/raw/master/LICENSE).

---
_Source: https://npm.io/package/node-ckan-crawler · Machine-readable twin of the npm.io package page. Health data is recomputed on every publish._
