About the DisclosureLens crawler
This page tells you how to identify our crawler, how often it fetches, and how to make it slow down. We fetch a fixed set of regulator pages at the per-source rate caps below, and we respond to abuse@disclosurelens.com within one business day.
Identification
Fetches identify as DisclosureLens in the HTTP User-Agent header, with a contact address:
DisclosureLens <contact email>SEC EDGAR and most regulator fetches carry exactly that form (per the SEC user-agent policy); HHS OCR fetches append this site's URL; leak-site screenshot fetches identify as DisclosureLens-Screenshot-Fetcher. A small number of state portals block all non-browser clients outright — those receive a standard browser user-agent, at the same rate caps, because their operators left no bot-accessible path to public records.
Politeness policy
- SEC EDGAR: max 8 requests/sec (within SEC's 10/sec documented cap)
- HHS OCR: 1 request/sec, sequential
- State AG portals: 1 request per 3 seconds, jittered
- GDPRhub (EU DPA decisions): 1 request/sec
- ransomware.live API: 1 request/sec
We back off exponentially on HTTP 429 and 5xx responses, and per-source rate caps are enforced in a shared HTTP layer — no fetch path can exceed them.
What we collect
Only material that is already public on the source regulator's portal. We archive raw filings to immutable storage for evidentiary purposes (audit replay, schema migration). On disclosure detail pages we render the full SEC and HHS filing inline from the originating regulator's public record under the fair report privilege (see methodology §4.5). State AG, EU DPA, and threat-actor leak-site sources are not redistributed inline; those records carry only a link to the originating regulator URL.
Block requests
If you operate a regulator portal and need our bot to slow down or stop, email abuse@disclosurelens.com. We respond within one business day and apply the requested rate limit immediately.
Why we publish this page
Most bots that hit your AG portal don't identify themselves and don't publish a contact. We'd rather you block us correctly than rate-limit-bomb your portal trying to figure out who's causing the spike. The published User-Agent, the abuse@ contact, and the per-source rate caps above are all part of the same contract — read them as "here's how to make us behave," not as a marketing artifact.