Privacy
This notice explains two separate things, because DisclosureLens involves two very different groups of people:
- People who use this site or the API — what we collect from you, which is very little.
- People who may appear inside the public records we index — breach notifications and related documents published by regulators, which can contain personal data about individuals who never interacted with us.
Who we are
DisclosureLens is operated by [PLACEHOLDER: legal entity name, legal form, and registered address]. For privacy questions and requests, contact corrections@disclosurelens.com or use the report / removal page.
[PLACEHOLDER: data protection officer, if one is appointed] [PLACEHOLDER: EU/UK Article 27 representative, if required] [PLACEHOLDER: lawful basis relied on — e.g. legitimate interests and/or the journalism / special-purposes exemption — and the balancing assessment behind it]
1. If you use the site or the API
Browsing without an account
Nearly all of DisclosureLens is browsable anonymously; only account settings and the admin tools require signing in. If you browse without signing in:
- We set no cookies of our own, and we run no analytics, advertising, or third-party tracking scripts.
- Your requests appear in our web-server access logs (see Server logs below).
- Two browser
localStoragekeys are used, both only on interaction and both staying in your browser: one remembers what you pinned with the “compare” feature, the other remembers your light/dark theme choice.
If you create an account
Accounts are handled by Clerk, our identity provider. Clerk operates the sign-in and sign-up screens. We never receive or store your password. From Clerk we receive and store:
- Your Clerk user id, your primary email address, and whether it is verified
- An organization name — which, if you have not set one, may fall back to the first name on your Clerk profile
- A role claim, read from the session token on each request to grant operator access — it is not stored in our database
We additionally store, for accounts that use those features:
- Alert subscriptions — the saved query you want to be notified about, an optional name you give it, delivery frequency, and an unsubscribe token
- Email delivery records — recipient address, which template was sent, when, the provider’s message id, and any bounce or complaint. Addresses that hard-bounce or complain are added to a suppression list so we stop mailing them.
- API keys — stored only as a hash plus a short non-secret prefix; the key itself is never stored
- Usage counters — per-day API call and processing volumes, for capacity and billing accounting
We do not collect payment data. No payment processor is integrated and no card data reaches us.
Server logs
Our web server records standard access logs. Authorization headers and credentials are redacted before writing. Logs rotate on size with a configured maximum retention of approximately 180 days [PLACEHOLDER: confirm whether logs are copied or backed up anywhere beyond the production host]. We periodically analyze these access logs with a self-hosted tool (no third party involved) to understand traffic. No IP address, user-agent, device, or location field is stored in our application database.
Cookies
We set none. If you sign in, Clerk sets the session cookies needed to keep you signed in. [PLACEHOLDER: enumerate the Clerk cookie names, purposes, and lifetimes, and decide whether a consent mechanism is required in your target markets]
2. Personal data inside the records we index
This is the part that most often matters to people who have never used the site. We index documents that regulators and other sources publish, and those documents can contain personal data — for example the text of a breach-notification letter, or a health-sector breach report. We also index threat-actor leak-site claims, which are unverified allegations published by criminal groups.
Being precise about what is and is not filtered:
- Automated redaction covers email addresses, US Social Security numbers (in standard dashed form), US-format telephone numbers, and payment-card-shaped numbers in the extracted narrative summary, and every redaction is logged.
- Natural-person names are not currently redacted. If a published source document names an individual, that name can appear in our extract.
- SEC and HHS filing bodies — the regulator's own public record — render inline for all visitors. All other raw source artifacts (state-AG letters, leak-site screenshots) are restricted to operators.
If personal data about you appears in a record, use the report / removal page. We would rather hear from you than not.
Who else processes data
We use the following providers. We do not sell personal data, and we do not use it for advertising.
- DigitalOcean — hosting. All application components and databases run on a single server.
- Clerk — authentication and account management
- Amazon Web Services (SES and SNS, US East / Ohio) — sending alert and transactional email, and receiving bounce/complaint notifications
- Cloudflare — authoritative DNS only. Our site and API domains resolve directly to our own server; Cloudflare does not proxy, terminate, or see your traffic to them. Cloudflare R2 is separately used to store archived source documents.
- Let’s Encrypt — TLS certificate issuance
- Language-model and embedding providers — used to extract structured records from source documents, not from your account data. Extraction currently runs on a self-hosted model on operator-controlled hardware; a minority of documents may be escalated to Anthropic, and Voyage AI generates embeddings from record summaries.
[PLACEHOLDER: confirm zero-retention / no-training terms and any DPA with each provider] - Search engines, via a self-hosted search proxy — entity-resolution lookups send organization names (not your personal data) to public search engines
For clarity, because they appear in our configuration but are not in use: we run no payment processor, no error-reporting service, and no third-party analytics or product-telemetry backend (our only traffic analysis is the self-hosted server-log review described above).
[PLACEHOLDER: data residency — the regions of the hosting, object storage, identity, and email tenants] [PLACEHOLDER: international transfer mechanism (SCCs / DPAs) for each provider]
Retention
[PLACEHOLDER: define a retention period for each category — account records, email delivery logs, suppression lists, usage counters, admin audit logs, archived source documents, and server access logs.]
Two things we should state plainly today, because they are current practice rather than aspiration:
- We do not currently run an automated deletion or purge process. Archived source documents are kept indefinitely by default.
- Closing your account does not currently erase your record. When an account is deleted, we disassociate the identity provider link but retain the customer row — including the email address — as part of the audit trail.
[PLACEHOLDER: confirm the legal basis and retention period for this, or change the behaviour to a hard delete on request]
Your rights
Depending on where you live, you may have rights to access, correct, delete, port, or object to the processing of your personal data, and to complain to a supervisory authority.
To exercise any of them, use the report / removal page or email corrections@disclosurelens.com. Requests are handled manually — there is no self-service export or deletion tool today. We aim to acknowledge within 48 hours and may need to verify your identity before acting on a deletion request.
[PLACEHOLDER: US state privacy disclosures — whether any activity constitutes a “sale” or “sharing”, the opt-out mechanism if so, and the status of any data-broker registrations] [PLACEHOLDER: minimum age / children’s data position]
Changes
We will post any change here and update the date below. [PLACEHOLDER: effective date]