Self-hostable web content extraction API for RAG pipelines: clean Markdown, JSON, and crawling.
by Sparkfetch
SparkFetch is a self-hostable web fetching and extraction API. You send a URL; it returns cleaned Markdown, structured JSON, or plain text. It also crawls entire websites with depth and page limits, handling link discovery and deduplication.
The project is built for AI applications, RAG pipelines, and research tools that need reliable web data without managing Puppeteer or proxy rotation. The live site is a landing page with a brief description; the GitHub repo has a detailed README and releases.
SparkFetch is an alternative to services like Reader or Firecrawl, but open-source and self-hostable. It's a developer tool for teams that want to own their scraping infrastructure.
Built with
Last checked 2026-09-01 11:44 UTC
Open live app ↗| timestamp | result | latency |
|---|---|---|
| 2026-09-01 11:44 UTC | live | 1423ms |
| 2026-08-26 21:23 UTC | live | 619ms |
| 2026-08-26 20:15 UTC | live | 49ms |