01
Train while crawling
Chose Embedding pages while the crawl still runs over crawling the whole site, then embedding it.
A trainer thread indexes new pages every three seconds, so a chatbot answers from its first pages long before a 200-page crawl finishes.
Trade-off: Two threads share crawl state, and a final pass has to catch pages that land after the last tick.
02
A namespace per chatbot
Chose Serverless Pinecone with one namespace per chatbot over a self-hosted RediSearch index on the app server.
It keeps each customer’s knowledge apart and takes vector search off the single server that already runs the web app, Redis and workers.
Trade-off: Every search became a network hop, so I later cached the client per process to remove a 1–2s round trip.
03
Proxy only when blocked
Chose Direct fetches with a stealth-proxy fallback on 403, 429 or 503 over routing every page through the scraping proxy.
Pages that respond within five seconds cost nothing to fetch; only sites that block bots use paid ScrapingBee requests, capped at ten at once.
Trade-off: Blocked pages wait out the direct attempt first, and there are two fetch paths to maintain.