Kestryl · Powered by PIIScan

Prove what personal data you hold — and what you did about it.

Kestryl finds personal data across databases, documents, scans, and images, then produces the evidence a regulator, auditor, or customer can accept. PIIScan 0.6.0 extends remediation from database records to clean copies of supported documents without altering an original file.

PIISCAN ENGINE 0.6.0 · BENCHMARKED 04 SEP 2026

The problem

Everyone takes privacy seriously. Few can prove it on demand.

Personal data collects quietly in enriched leads, support transcripts, scanned forms, and images. When a deletion request, an audit, or a regulator’s inquiry arrives, sincerity is not a defense. Evidence is. Kestryl is built so “show me exactly what you searched for and what you did” is immediate and repeatable.

Three outcomes

Find it everywhere. Act only when you decide. Show your work.

i

Find it everywhere

Database tables, spreadsheets, Office files, email, archives, PDFs, scans, and images are read by the same engine to the same standard. A finding means the same thing whether it came from a warehouse column or a scanned form.

ii

Act only when you decide

Discovery changes nothing. Remediation requires three separate approvals and a confidence threshold you set, so no single person and no single setting can trigger it. Every action, and every refusal, is recorded.

iii

Show your work

Each run records which rules were in force, what was found, what was done, and what was deliberately left alone in a form privacy, legal, and audit teams can present without translation.

Payment cards, IBANs, and similar identifiers are checked against their published validation. Numbers that cannot exist never reach your findings as real PII. A deterministic rule-set fingerprint answers what was searched for and when.

New in PIIScan 0.6.0

Remediation that respects the original.

Supported document writers produce cleansed copies in a folder you designate. Word, Excel, PowerPoint, OpenDocument, email, archives, and plain-text/data formats are covered where the release supports rewriting. The original source document is never rewritten.

Discovery and read support

  • PDFs and scanned pages
  • Images and their readable metadata
  • Legacy Word .doc
  • Audio/video metadata and tags where supplied

Remediation and rewrite support

  • Word, Excel, PowerPoint
  • OpenDocument formats
  • Email and archives
  • Supported text and data formats
Reading is wider than writing

Kestryl reads more formats than it rewrites and tells you which is which on every run. PDFs, scanned pages, and images are found and reported, not rewritten. Legacy Word .doc files are read but refused for cleaning until saved in a modern format.

Measured, not promised

Fast enough to check everything. Deliberate where it changes data.

These are PDI-reported measurements on synthetic data, benchmarked 04 SEP 2026. They were measured on a single processing thread; production deployments run many threads and are sized during scoping.

12 million records, fully scannedtables only · one thread · median of three runs
6.06 seconds
Records scanned per secondsame run · end to end
1.98Mrows/sec
12 million records plus 12,000 documentswarm run · cold-storage first pass: 61.3 seconds
47.65seconds
Same 12 million records, remediation appliedeach masked record is its own recoverable step
~4minutes
Records plus documents, remediation and clean copiesmasked records plus cleansed document copies
~10.6minutes / 635 seconds

12M timing finalized at 6.06 seconds from PDI-reported measurements on synthetic data, benchmarked 04 SEP 2026. All other supplied benchmark values are retained.

Governance

Nothing is masked by accident.

  • DefaultFindings are recorded and source data is untouched. Discovery is always safe to run.
  • Three gatesThe rule must allow the action, configuration must permit it, and the operator must confirm it at run time.
  • ConfidenceA threshold you set determines when a finding is eligible for remediation. Lower-confidence matches remain visible for human review.
  • VaultIn vault mode, removed values can be restored by the key holder through a deliberate, recorded step. Retention limitations remain explicit.
  • NamesOptional name detection is a report-only heuristic where supplied; it is not deterministic remediation.

RULESET FINGERPRINT · DETERMINISTIC ACROSS ENVIRONMENTS

Why Pacific Data Integrators

Evidence that fits the systems your data lives in.

Kestryl comes from a firm with 15+ years and 100+ implementations across Informatica, Salesforce, Snowflake, and Databricks. It joins PDI’s data-quality portfolio on the conviction that trustworthy data is not a report produced for an audit, but a property you can prove at any moment.

FAQ

Frequently asked questions

What does Kestryl discover?

Kestryl reads structured records, spreadsheets, Office files, email, archives, PDFs, scans, and images through one controlled pipeline. Payment cards, IBANs, and similar identifiers are checked against their published validation before they are reported as findings.

Does remediation rewrite original documents?

No. Supported document remediation produces a cleansed copy in a folder you designate. The original source document is not rewritten. PDFs, scanned pages, images, and legacy Word .doc files are read or reported but are not rewritten.

When can Kestryl mask a value?

Discovery is read-only. Remediation requires three separate approvals: the rule must allow it, configuration must permit it, and the operator must confirm it at run time. A confidence threshold is also set before a value becomes eligible.

How are Kestryl actions evidenced?

Each run records the rules in force, findings, actions, refusals, and deliberately untouched values. A deterministic rule-set fingerprint and chained audit log make the result repeatable and tamper-evident.