Lessons from crawling the web at scale [2026w17] episode artwork

EPISODE · Jul 26, 2026 · 6 MIN

Lessons from crawling the web at scale [2026w17]

from Lead Prompt Podcast · host John Collins

In this episode of Lead Prompt, I dive into the engineering realities of scaling Greppr.org past 41 million indexed documents and 0.5TB of data. Moving past management theory, I break down the engineering optimizations required to survive the chaos of crawling the web at scale. Discover how I bypassed geo-location redirects using proxies, defeated crawler tar pits with custom fail-fast timeouts, implemented rapid early-stage filters for NSFW content and non-text rich media, and eliminated index bloat through automated canonical URL deduplication. Show notes are here: https://leadprompt.sh/a/737-Lessons-from-crawling-the-web-at-scale-2026w17 Keywords: greppr, web crawling, systems engineering, solo founder, apache solr, web scraping, backend development, data indexing, web crawler at scale, canonical urls, crawler tar pits, proxy routing, geo ip redirection, nsfw filtering, rich media filtering, solr architecture, software engineering podcast, lead prompt, tech infrastructure

Episode metadata supplied by the publisher feed · Published Jul 26, 2026

Embed this episode

NOW PLAYING

Lessons from crawling the web at scale [2026w17]

0:00 6:57

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Lead Prompt Podcast?

This episode is 6 minutes long.

When was this Lead Prompt Podcast episode published?

This episode was published on July 26, 2026.

Can I download this Lead Prompt Podcast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!