Ad
 
Learn More

Self-hosted Scraping Platforms & SDKs

A curated collection of the best self-hosted platforms and SDKs for building web scraping solutions.

Turn any website into clean, AI-ready data via API
  • Stars


    171,968
  • Last commit


    21 hours ago
  • License


    AGPL-3.0
API for AI agents to search, scrape, crawl, and interact with the live web, returning clean Markdown, structured JSON, or screenshots from any page.

Open Source Alternative to:

LLM-ready web crawler built for AI data pipelines
  • Stars


    79,367
  • Last commit


    10 hours ago
  • License


    Apache-2.0
Open-source web crawler and scraper that produces clean, structured output optimized for LLMs, RAG pipelines, and AI agents. Supports async crawling, CSS/XPath/LLM extraction, and stealth browser control.

Open Source Alternative to:

No-code web scraping, crawling, and extraction platform
  • Stars


    17,280
  • Last commit


    22 hours ago
  • License


    AGPL-3.0
Turn any website into structured data using a visual recorder, natural language prompts, or API. Includes proxy rotation, scheduling, and webhook integrations.

Open Source Alternative to:

Self-hosted SERP API for Google, Bing, Yandex, and more
  • Stars


    1,300
  • Last commit


    1 month ago
  • License


    MIT
Query Google, Bing, Yandex, Baidu, DuckDuckGo, and Ecosia through one REST API. Self-host free under MIT, or use the managed cloud option.

Open Source Alternative to:

Favicon