EPISODE · Jul 26, 2026 · 6 MIN
Lessons from crawling the web at scale [2026w17]
from Lead Prompt Podcast · host John Collins
In this episode of Lead Prompt, I dive into the engineering realities of scaling Greppr.org past 41 million indexed documents and 0.5TB of data. Moving past management theory, I break down the engineering optimizations required to survive the chaos of crawling the web at scale. Discover how I bypassed geo-location redirects using proxies, defeated crawler tar pits with custom fail-fast timeouts, implemented rapid early-stage filters for NSFW content and non-text rich media, and eliminated index bloat through automated canonical URL deduplication. Show notes are here: https://leadprompt.sh/a/737-Lessons-from-crawling-the-web-at-scale-2026w17 Keywords: greppr, web crawling, systems engineering, solo founder, apache solr, web scraping, backend development, data indexing, web crawler at scale, canonical urls, crawler tar pits, proxy routing, geo ip redirection, nsfw filtering, rich media filtering, solr architecture, software engineering podcast, lead prompt, tech infrastructure
Embed this episode
NOW PLAYING
Lessons from crawling the web at scale [2026w17]
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.