01
Deterministic scoring
Chose A published rubric that does the arithmetic over letting an LLM score each page.
A score that changes between runs on identical content destroys trust. LLMs only label content, at temperature 0, cached by content hash.
Trade-off: A new classifier version shifts scores, so it is versioned and regression-gated like a rubric change.
02
Citations through official APIs
Chose Sampling the engines’ official APIs over scraping the consumer chat apps.
APIs are stable, within the terms of service and cheap enough. Three samples per prompt and engine damp the noise.
Trade-off: API answers can differ from what people see in the apps, so every result is labelled “API sampling”.
03
The agent dispatches, workers run
Chose A Celery pipeline that the agent dispatches over running the crawl inside the agent loop.
Audits take minutes, must survive a lost chat thread, and the REST button and the agent run the same pipeline.
Trade-off: Two execution worlds: progress has to be relayed from the workers to the chat over Redis.