DisclosureLens

Sources & Methodology

This page documents what DisclosureLens does to a regulator-published filing between the moment it appears on a public portal and the moment it appears in the API. Eight pipeline stages, one schema, four model-grade accuracy targets — and, for the field the compliance clocks depend on, a measurement rather than a target.

Accuracy targets

These are service-level objectives — the accuracy the pipeline is designed and tuned to hit — not independently measured or audited performance figures. Field-level accuracy varies by source and field; report discrepancies via the corrections process below. Where we have measured a field against its sources, the measurement is published and supersedes the objective: see discovery dates.

Discovery dates — measured, not assumed

Every breach-notification clock in this product runs from a discovery date, so that one field decides whether a filing is reported as on time or late. It is the field we have audited most heavily, and the audit found something a conventional accuracy percentage cannot express: a date can be transcribed perfectly and still be the wrong date.

A statutory clock needs first awareness — the day the filer knew something was wrong. Notification letters routinely date two other events in the same sentence, and both are the wrong concept:

A ±1-day transcription metric is blind to both: the date is copied correctly and is wrong by months in meaning. So the field is measured two different ways, and how the regulator publishes it decides which one applies.

Where the regulator publishes the date as a labelled field

Washington (dateaware), Oregon (discovery_dates_raw) and Texas (Breach_Discovered_Date__c) publish discovery as a structured form field, as does the majority of Maine’s intake. There is no inference to get wrong. Every such record was re-read from its archived source document and compared against what we stored: 3,781 of 3,781 exact, zero mismatches (Washington, Oregon and Texas; an earlier pass verified 470 of 470 Maine form rows).

Where the date exists only in prose

In the remaining jurisdictions the date has to be read out of the letter. Measured against a hand-labelled set of 210 filings drawn at random per jurisdiction across nine states — of which 185 are decidable, the rest genuinely ambiguous in the source — prose-inferred discovery dates are correct 87.6% of the time (162 of 185; 95% CI 82.0–91.6%). That figure, not the transcription objective above, is the one to rely on for a prose jurisdiction. Per-state precision ranges from 81.6% (New Hampshire, 38 cases) to 100% (Vermont, 11 cases); the smaller samples carry correspondingly wide intervals.

What we have done about it

Rather than restate the dates, we built lexical detectors for the two wrong-concept shapes, held them to a measured precision floor before use, and then had a human read every filing they selected against its archived letter. As of 24 August 2026:

The detectors do not find every instance. Our own estimate is that 1,100–2,100 filings carry the post-investigation defect; the 858 corrected or marked to date represent part, not all, of it. A monthly sweep now reviews only filings no human has seen before — the most recent surfaced eight. An unmarked filing means nobody has examined it, not that it has been verified.

When we decline to grade a filing

Where a discovery date is known to be the wrong concept and the letter offers no replacement, the honest answer is not to guess and not to grade. Those records carry a visible marker instead of a verdict:

Each marker is a statement about our data, never about the filer, and neither counts toward an on-time ratio. A marker means a verdict was withheld; its absence means the clock either ran or never applied. The withholding is reversible: if a filer or reader can point to a first-awareness date in the source document, the record is re-adjudicated and the clock runs again — that is a corrections request, and it is welcome.

Source list

We ingest primary regulatory sources plus two declared low-trust supplements (leak-site claims and press coverage), each labelled as such on every record. The authoritative, always-current list — with per-source freshness SLOs — is the live source-health table. Current coverage:

Source kinds — every source, one grammar

DisclosureLens carries seven source kinds, from statutory filings to attacker self-reports. Each has a fixed badge, a fixed render behavior, and — crucially — a fixed honest ceiling on what it can populate. The badge tells a reader how much to trust a record before they read a word of it; clicking any source badge on a disclosure page opens this reference in place.

Source kindSource bodyPopulatesCompliance clockTypical conf.
SEC filings
Federal · material events
Inline render — full filing embedded (fair-report §4.5)discovery + materiality dates · attack vector · threat actor; rarely affected countsSEC 4-business-day (hard)62–87%
HHS OCR
Federal · health (HIPAA)
Inline render — full filing embedded (fair-report §4.5)affected count (required) · PHI data types · breach typeHHS 60-day (submission date only — the portal omits the discovery date)67–97%
State AG
State breach notices
Link-out — body stays at the regulator; we link + extract (§4.5)per-state resident counts · data types · discovery dateState 30/60-day · indicative (web-form dates run systematically late)58–95%
EU DPA
GDPR supervisory authorities
Link-out — body stays at the regulator; we link + extract (§4.5)data categories · cross-border scope · outcome + articles (decisions)GDPR 72-hour (indicative)95%
Singapore PDPC
PDPA enforcement
Link-out — decision stays at the regulator; where it publishes a grounds-of-decision PDF we link and extract itoutcome + obligations · fine (SGD) and affected count where a grounds document states themnot assessable — enforcement decision, not a breach notice60–95%
Press
Media / market postings
Link-out — journalistic coverage links to the source URLincident type + narrative only (may be machine-translated)not assessable60%
Leak site
Threat-actor extortion blogs
Aggregated via ransomware.live — claims, not filingsactor name · victim claim · ransom/leak statusnot assessable50%

Color roles are consistent across the feed, incident, and disclosure views: teal — confirmed · on-time · high confidence; amber — claim · indicative · caveat; rust — unverified · late · discrepancy. Confidence bands describe what we observe in the current corpus, not a guarantee.

Threat-actor leak sites

We additionally ingest ransomware-group leak-site claims via ransomware.live, a public aggregator of structured listings from threat-actor extortion blogs (Source: Ransomware.live). These rows are segregated into a dedicated, publicly-browsable “Claims” feed, kept separate from confirmed regulatory filings so a threat-actor assertion is never presented as a filed disclosure. They render with a danger-tone “Leak Site” chip and a claim banner on the detail view.

Why this source exists: leak-site postings often predate the victim's regulatory filing by weeks to months, and that gap is meaningful — it's the “knew but didn't disclose” window that informs governance, insurance, and plaintiff-bar analysis. Every regulatory disclosure carries a graduated pre_disclosure_leak_gt_30d / gt_90d / gt_180d compliance flag when a same-victim leak-site listing predates it by the corresponding threshold.

Treat leak-site rows as the threat actor's assertion, not a confirmed breach. We do not link to .onion claim URLs from the UI even when ransomware.live exposes them. For groups with US sanctions exposure — where OFAC has designated the group or its operators (Evil Corp as an entity; LockBit operators, Feb 2024; Trickbot-gang members tied to Conti, 2023) — we surface a sanctions-caution callout. Public reporting on sanctioned-group activity is permitted (informational-materials exemption, 50 U.S.C. § 1702(b)(3)); payment or material support to those groups is not.

Enforcement decisions

Regulator rulings are a fourth record category, browsable at /enforcement and kept distinct from disclosures, claims, and press reports: an enforcement decision is a regulator's ruling after the fact, not anyone's notification of their own breach. The named organization is the respondent in a regulatory proceeding, not necessarily a breach victim, and a decision page is never described as a breach notification. Fines render as filed, in the regulator's own currency — we do not substitute converted or news-reported figures. Where a decision is adjudicated to concern the same incident as a filed disclosure, the incident's lifecycle state advances to enforced.

Cross-source linkage

When a ransomware group publicly claims a victim, and the same victim later files a regulatory breach notice, we link the two filings so the disclosure timeline reads as one incident. Linked filings render in a “Linked disclosures” card on each side's detail page with a confidence chip and the gap-days between the leak-site claim and the regulatory filing.

Signals used. The matcher requires the same resolved victim entity on both sides (a one-side-empty or unresolved-name pair is not a candidate). Within a 180-day window, candidate pairs are scored on: entity-id match (always 1.0 in the candidate pool by construction), narrative similarity via Voyage-3-large embeddings (cosine), threat-actor cross-reference (1.0 when both sides resolve to the same threat-actor entity-id, 0.6 for a lower-cased string match), data-type overlap, geographic overlap, malware-family overlap, and affected-count order-of-magnitude. Weights are documented in packages/python/df-core/df_core/incident_match.py; all matcher source is open-source in the repo.

Confidence tiers. Each rendered link carries a tier chip:

Operator review. Borderline pairs default to a human review queue rather than auto-merging silently. An operator reviews the side-by-side comparison + per-signal score breakdown and Confirms (link both sides into one incident), Rejects (writes a sticky negative record so the matcher won't re-propose), or Defers. The operator-led policy is intentional — the public-facing surface only carries links a human has approved, OR the matcher has scored above the high-confidence threshold.

False-positive handling. A wrongly-linked pair can be unlinked by an operator via the admin merge tool; the unlink writes a sticky record so the matcher won't re-link the same pair. Customers who spot a false-positive in their own filings can email [email protected]; we prioritize first-party correction requests over the standard review queue.

Limitations. A threat-actor leak claim is not proof a breach occurred — some leak-site listings are exaggerated, recycled from old breaches, or fabricated. The linkage we surface is “same victim entity, same time window, narrative similarity”; the chip language reflects that confidence rather than a forensic finding. The matcher does not cross sources outside the same victim entity — perpetrator-only linkage (e.g., “all the Akira victims”) lives in the threat-actor pages, not in the disclosure-level cross-source links.

Extraction pipeline

Each source document passes through:

  1. Sanitization — Unicode normalization, dangerous-HTML stripping, bidi override removal, prompt-injection guard.
  2. Triage — the self-hosted Qwen3.6-35B-A3B model (served via MLX) classifies whether the document is a breach disclosure and selects an extraction template.
  3. Structured extraction — the same self-hosted Qwen3.6-35B-A3B model, with Instructor + Pydantic schema validation, produces a BreachDisclosure v1 object with an overall confidence score. Per-field confidence is scored at extraction time to drive the hard-pass thresholds below, but is not yet persisted or served.
  4. Hard-pass review — high-stakes fields carry per-field escalation thresholds (threat-actor and malware attribution at 0.85; affected counts, industry tags, and materiality dates at 0.66; body-ungrounded discovery dates at 0.50). A below-threshold field triggers a harder re-extraction pass, and named attributions additionally face an adversarial verify pass before persistence.
  5. Entity resolution — name + LEI/CIK lookup via the GLEIF and SEC EDGAR registries.
  6. Cross-jurisdiction dedup — same incident filed across multiple jurisdictions (SEC + state AG + OCR) is grouped under a canonical incident id.
  7. PII redaction — structured PII (email addresses, SSNs, phone numbers, and payment-card numbers) appearing in incident narratives is replaced with type-tags before customer-visible fields are emitted. Natural-person names are not yet redacted.
  8. Human review queue — any record below the confidence threshold is reviewed by a DisclosureLens operator.

Extraction audit trail

Every extracted record carries:

GET /v1/disclosures/{id}/extraction returns the full envelope. Re-extraction with a newer prompt or model produces a new envelope, not a mutation; the record carries the version chain.

Compliance note: this also satisfies EU AI Act Article 50, which takes effect August 2, 2026. We built it because customers — security analysts, GRC reviewers, plaintiff-bar researchers — need to defend their use of an extracted value in front of someone who will challenge it. The regulation arrived later.

Corrections

Email [email protected] — 48-hour SLA from receipt.

Corrections are applied to the record and, where the correction changes a compliance verdict, the verdict is withdrawn or re-derived rather than annotated around. We do this to our own findings unprompted as well: see the discovery-date audit, which withdrew 129 previously published verdicts. A subject of a record may also request a published right of reply, which appears on the record itself.

Provenance

Every record carries:

Signed documents and evidence packages

Four deliverables leave the platform as signed PDFs — the entity scorecard, the compliance report, the broker benchmark letter, and the per-incident evidence package. The evidence package assembles every regulatory filing linked to one incident into a single document: timeline, notification clocks, scope revisions, cross-filing comparison, named litigation metrics, data-sensitivity profile, entity breach history, and a per-filing chain-of-custody appendix. Sections with nothing recorded say so explicitly (“none recorded”) rather than being omitted. A walkthrough of a real package, with screenshots, is at /evidence.

Signature. Packages are signed PAdES-B (ETSI EN 319 142) with byte-range tamper detection: any post-signing modification of the bytes invalidates the signature. The signer certificate is our own published signing certificate — not issued by a certificate authority — and the document says so in its own footer, which prints the signer name, certificate fingerprint, and generation timestamp. A PDF reader that only recognizes certificate-authority chains will list the signer as unknown; verification means checking that the fingerprint matches the published one and that the signature over the byte range is intact, in a reader’s signature panel or programmatically. Each generated package is also recorded server-side: template, SHA-256 of the signed bytes, signer fingerprint, and timestamp.

Chain of custody. The appendix records three layers per linked filing:

A purchased package freezes at first download — the copy is stored under a content-addressed key and served unchanged for the life of the link, so later changes to the live record do not change the document. The package reports filed facts with their provenance and draws no legal characterization; an operator affidavit describing the collection and signing methodology can be commissioned for a matter that requires one.

§4.5 Editorial claim — Fair Report Privilege

DisclosureLens publishes structured extracts of public records. Raw filings from federal public-domain regulators (SEC EDGAR, HHS OCR) are rendered inline on the disclosure detail page from the originating regulator’s public record, under the Fair Report Privilege. Raw filings from state attorney-general portals, EU DPA registers, UK ICO datasets, Australia OAIC reports, and threat-actor leak-site aggregators have their structured extracts published (and indexed) on our public surfaces, while their raw source bodies are linked to rather than redistributed in full. The editorial work — schema design, entity resolution, cross-jurisdiction linking, confidence scoring — is the original contribution that anchors the FRP claim across all sources.

SEC + HHS inline rendering. SEC 8-K / 10-K filings and HHS OCR breach- report rows are federal public-domain records with no third-party copy restriction and no defamation surface (regulator-filed, not third-party-asserted). Inline rendering surfaces the original document alongside our structured extract so a reader can verify or correct any field. The rendered document is server-side sanitized (script / iframe / form / style stripped) and embedded in a sandboxed iframe with no script execution; the dashboard never modifies or paraphrases the original body, only the structural rendering.

State AG + leak-site posture (unchanged). State AG portals carry varying per-portal terms-of-use; threat-actor claims carry defamation risk by their nature (the claim is the threat actor’s accusation, not a confirmed breach). Both surfaces retain the “structured extract + link to source” posture so the editorial transformation remains the public-facing contribution.