Scrapy MongoDB Queue

MongoDB-based components for scrapy that allows distributed crawling

Available Scrapy components

Scheduler
Duplication Filter

Installation

From pypi

  $ pip install git+https://github.com/jbinfo/scrapy-mongodb-queue

From github

  $ git clone https://github.com/jbinfo/scrapy-mongodb-queue.git
  $ cd scrapy-mongodb-queue
  $ python setup.py install

Usage

Enable the components in your settings.py:

  # Enables scheduling storing requests queue in redis.
  SCHEDULER = "scrapy_mongodb_queue.scheduler.Scheduler"

  # Don't cleanup mongodb queues, allows to pause/resume crawls.
  MONGODB_QUEUE_PERSIST = True

  # Specify the host and port to use when connecting to Redis (optional).
  MONGODB_SERVER = 'localhost'
  MONGODB_PORT = 27017
  MONGODB_DB = "my_db"

  # MongoDB collection name
  MONGODB_QUEUE_NAME = "my_queue"

Author

This project is maintained by Lhassan Baazzi (GitHub | Twitter | LinkedIn)

Name		Name	Last commit message	Last commit date
Latest commit History 6 Commits
scrapy_mongodb_queue		scrapy_mongodb_queue
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
requirements.txt		requirements.txt
setup.cfg		setup.cfg
setup.py		setup.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

scrapy_mongodb_queue

scrapy_mongodb_queue

.gitignore

.gitignore

LICENSE

LICENSE

README.md

README.md

requirements.txt

requirements.txt

setup.cfg

setup.cfg

setup.py

setup.py

Repository files navigation

Scrapy MongoDB Queue

Available Scrapy components

Installation

Usage

Author

About

Releases

Packages

Contributors 2

Languages

License

jbinfo/scrapy-mongodb-queue

Folders and files

Latest commit

History

Repository files navigation

Scrapy MongoDB Queue

Available Scrapy components

Installation

Usage

Author

About

Resources

License

Stars

Watchers

Forks

Languages