Open-source web crawler and scraper that produces clean, structured output optimized for LLMs, RAG pipelines, and AI agents. Supports async crawling, CSS/XPath/LLM extraction, and stealth browser control.
No-code web scraping, crawling, and extraction platform
Stars
17,360
Last commit
13 hours ago
License
AGPL-3.0
Turn any website into structured data using a visual recorder, natural language prompts, or API. Includes proxy rotation, scheduling, and webhook integrations.
The purpose-built time series platform for high-velocity ingestion and real-time queries at scale, without sacrificing performance or cost. Download InfluxDB for free.