Home/Agents/
Incept ClassifyIQ™
Data Privacy & Compliance
 · 

Incept ClassifyIQ™

Function: AI Data Classification

A four-layer pipeline that will not tag anything sensitive on a single signal. Discovery and sampling, then pattern and statistical analysis, then semantic classification by the model, then a confidence score that decides whether a tag is applied automatically or queued for a human. Reviewer decisions feed back into prompt tuning, so it improves on your data specifically.

40+ PII patterns · dual registration to your platform and governance catalogs · audit trail per tag

  • Databricks Unity Catalog
  • Microsoft Purview
  • Informatica
40+
PII patterns matched
4
pipeline layers
5
PII categories
4
sensitivity tiers
[Interactive demo]

Walk the pipeline yourself.

Sample data, in your browser. A faithful simulation of the real flow. Approve or reject at step four and watch the score change.

Incept DataIQ™

analytics.sales.customer_master

Sample data · not a client system

Profiling analytics.sales.customer_master

Eight dimensions across every column, five of them shown here. Nothing is written; this is a read.

Column
Type
Null %
Distinct
Detected pattern
Signal
customer_id
string
0.0%
4,812
^C[0-9]{6}$
clean
customer_name
string
0.2%
4,780
free text
clean
tax_id
string
12.2%
4,201
^[0-9]{2}-[0-9]{7}$
issue
country_code
string
0.0%
47
ISO 3166-1 alpha-2
review
email
string
8.1%
3,944
RFC 5322
issue
created_date
date
0.0%
1,204
ISO 8601
clean
status
string
0.0%
4
enum, 4 values
clean
credit_limit
decimal
31.7%
892
numeric
review

4,812 rows · 8 columns · 2 columns flagged · 2 for review

Business metadata, drafted

Table and column descriptions written from the profile and the column context. Every one is a draft until a steward accepts it.

Object
Drafted description
customer_master
Master record for every business customer the organization sells to. One row per customer, keyed on customer_id. Sourced from the order system and enriched with credit attributes.
customer_id
System-generated unique identifier for a customer. Format C followed by six digits. Primary key.
tax_id
Government-issued tax identification number, formatted NN-NNNNNNN. Required for any customer invoiced in the current fiscal year.
country_code
ISO 3166-1 alpha-2 country code for the customer’s registered address.
credit_limit
Approved credit ceiling in reporting currency. Null where no credit assessment has been completed.

1 table + 4 columns described · awaiting steward

Rules recommended, with the evidence

Seven rules proposed from the profile. Each carries the proof that produced it, so approval is a judgment call rather than a leap of faith.

customer_id, Must not be null

Completeness

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE customer_id IS NULL;
Evidence

0.0% null across 4,812 rows. Already clean, worth enforcing as a hard constraint before it drifts.

customer_id, Must be unique

Uniqueness

SELECT customer_id, COUNT(*) AS occurrences
FROM analytics.sales.customer_master
GROUP BY customer_id
HAVING COUNT(*) > 1;
Evidence

4,812 distinct values across 4,812 rows. Uniqueness holds today; the rule locks it.

tax_id, Must match the tax ID format

Validity

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE tax_id IS NOT NULL
AND tax_id NOT RLIKE '^[0-9]{2}-[0-9]{7}$';
Evidence

97.2% of non-null values match the pattern. 118 do not, mostly nine-digit strings missing the hyphen.

tax_id, Null rate must stay under 5%

Completeness

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE tax_id IS NULL;
Evidence

12.2% null against a 5% domain threshold. 587 rows. This is the largest single rule failure in the table.

email, Must be a parseable email address

Validity

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE email IS NOT NULL
AND email NOT RLIKE '^[^@\\s]+@[^@\\s]+\\.[^@\\s]+$';
Evidence

4,359 of 4,422 non-null values parse. 63 do not: trailing semicolons and two addresses in one field.

status, Must be one of the accepted values

Validity

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE status NOT IN ('ACTIVE','INACTIVE','PENDING','BLOCKED');
Evidence

Exactly 4 distinct values observed, all within the expected set. Cardinality of 4 on 4,812 rows reads as a controlled vocabulary.

country_code, Must exist in the country reference set

Consistency

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master cm
LEFT JOIN ref.country c
ON cm.country_code = c.alpha_2
WHERE c.alpha_2 IS NULL;
Evidence

47 distinct codes. Two (XK, AN) are absent from ref.country, affecting 12 rows. Neither is a current ISO 3166-1 code.

7 rules · 4 quality dimensions

Steward approval

Nothing reaches production without this step. Uncheck anything you would not stand behind. In the real product, the SQL is editable too.

customer_id, Must not be null

Completeness

Evidence

0.0% null across 4,812 rows. Already clean, worth enforcing as a hard constraint before it drifts.

customer_id, Must be unique

Uniqueness

Evidence

4,812 distinct values across 4,812 rows. Uniqueness holds today; the rule locks it.

tax_id, Must match the tax ID format

Validity

Evidence

97.2% of non-null values match the pattern. 118 do not, mostly nine-digit strings missing the hyphen.

tax_id, Null rate must stay under 5%

Completeness

Evidence

12.2% null against a 5% domain threshold. 587 rows. This is the largest single rule failure in the table.

email, Must be a parseable email address

Validity

Evidence

4,359 of 4,422 non-null values parse. 63 do not: trailing semicolons and two addresses in one field.

status, Must be one of the accepted values

Validity

Evidence

Exactly 4 distinct values observed, all within the expected set. Cardinality of 4 on 4,812 rows reads as a controlled vocabulary.

country_code, Must exist in the country reference set

Consistency

Evidence

47 distinct codes. Two (XK, AN) are absent from ref.country, affecting 12 rows. Neither is a current ISO 3166-1 code.

7 of 7 approved, nothing will run

Execute approved rules

Approve at least one rule

Executed on schedule

Approved rules run against the table. Failures produce downloadable rejected records; scores register to the catalog.

97.7%

Overall score

93.9%

Completeness

100.0%

Uniqueness

98.7%

Validity

99.8%

Consistency

Rule
Dimension
Result
Failed rows
customer_id, Must not be null
Completeness
PASS
0
customer_id, Must be unique
Uniqueness
PASS
0
tax_id, Must match the tax ID format
Validity
FAIL
118
tax_id, Null rate must stay under 5%
Completeness
FAIL
587
email, Must be a parseable email address
Validity
FAIL
63
status, Must be one of the accepted values
Validity
PASS
0
country_code, Must exist in the country reference set
Consistency
FAIL
12

No rules approved, so nothing ran.

Back to approval

7 rules executed · 780 rejected rows · scores registered to the catalog

No rules executed · nothing registered to the catalog

[HOW IT WORKS]

The pipeline, step by step.

  1. Discovery and sampling Catalog API enumerates the estate. Metadata harvested, stratified sampling per column.
  2. Pattern and statistical analysis 40+ PII patterns: Social Security number, credit card number with checksum validation, email, phone, postal code, IP address, international bank account number. High cardinality plus a low null rate means a likely identifier. Format fingerprints flag encoded values.
  3. Semantic classification Column name, 20 sample values, table context and neighboring columns go to the model. Cross-column inference catches composite risk. Your taxonomy is injected into the prompt.
  4. Confidence and review Layers 1 and 2 agreeing auto-tags. Layer 3 disagreeing flags. No match queues. Reviewer decisions tune the prompt.

Sensitivity tiers: Restricted (PII, PHI, financial, regulated) · Confidential · Internal · Public. Continuous monitoring detects new tables and alerts on unclassified production data.

[CAPABILITIES]

What it does.

  • Direct identifiers. Name, email, phone, SSN, passport, driver’s license.
  • Quasi-identifiers. Date of birth, postal code, job title, gender, the re-identification risks people miss.
  • Financial. Credit card, bank account, IBAN, tax ID, salary data.
  • Health. Patient ID, diagnosis codes, prescription data, insurance ID.
  • Technical. IP address, device ID, geolocation, session tokens, cookies.
  • Dual registration. Tags go into your data platform catalog, where they drive attribute-based access control, and into your governance catalog for one audit trail.
[RELATED]

Also in Data Privacy & Compliance.

Incept AccessIQ™

Turns classification tags into access policy, then watches the drift.

Incept LakeGuard™

Governs who can reach what across S3 and Lake Formation, tag by tag.

Incept ConsentIQ™

Answers a privacy request by knowing where every copy of a person lives.

Incept DocuGov™

Extends governance past the database, into contracts, reports and the rest of your documents.

See it on your data.

Discovery is read-only, takes five days, and costs nothing. We assess your environment and tell you honestly whether this agent is worth your time.

Book a Consultation