A cache-aware web crawler that systematically crawls UWI web domains, extracts page metadata, classifies pages by campus/category/type, and presents everything in a searchable, filterable web directory with inline editing and scheduled job management.
How it knows: Polite by construction: robots.txt honoured including Crawl-Delay, ETag and If-Modified-Since conditional requests, and a semaphore capping concurrency per domain - a crawler that behaves well when nobody is watching.
Nicholas Smith - all work