Ad
 
Learn More

Open Source Web Crawlers using GitHub

A curated collection of the best open source tools for automatically crawling and extracting data from websites. using GitHub.

LLM-ready web crawler built for AI data pipelines
  • Stars


    79,404
  • Last commit


    14 hours ago
  • License


    Apache-2.0
Open-source web crawler and scraper that produces clean, structured output optimized for LLMs, RAG pipelines, and AI agents. Supports async crawling, CSS/XPath/LLM extraction, and stealth browser control.

Open Source Alternative to:

Favicon