Walk the pipeline yourself.
Sample data, in your browser. A faithful simulation of the real flow. Approve or reject at step four and watch the score change.
Runs Microsoft Purview and Fabric governance conversationally: scans, classification, glossary, lineage and sensitivity labels, so an Azure-standardized organization can operate governance through the stack it has already bought rather than a third-party platform.
Sample data, in your browser. A faithful simulation of the real flow. Approve or reject at step four and watch the score change.
Eight dimensions across every column, five of them shown here. Nothing is written; this is a read.
4,812 rows · 8 columns · 2 columns flagged · 2 for review
Table and column descriptions written from the profile and the column context. Every one is a draft until a steward accepts it.
1 table + 4 columns described · awaiting steward
Seven rules proposed from the profile. Each carries the proof that produced it, so approval is a judgment call rather than a leap of faith.
customer_id, Must not be null
Completeness
0.0% null across 4,812 rows. Already clean, worth enforcing as a hard constraint before it drifts.
customer_id, Must be unique
Uniqueness
4,812 distinct values across 4,812 rows. Uniqueness holds today; the rule locks it.
tax_id, Must match the tax ID format
Validity
97.2% of non-null values match the pattern. 118 do not, mostly nine-digit strings missing the hyphen.
tax_id, Null rate must stay under 5%
Completeness
12.2% null against a 5% domain threshold. 587 rows. This is the largest single rule failure in the table.
email, Must be a parseable email address
Validity
4,359 of 4,422 non-null values parse. 63 do not: trailing semicolons and two addresses in one field.
status, Must be one of the accepted values
Validity
Exactly 4 distinct values observed, all within the expected set. Cardinality of 4 on 4,812 rows reads as a controlled vocabulary.
country_code, Must exist in the country reference set
Consistency
47 distinct codes. Two (XK, AN) are absent from ref.country, affecting 12 rows. Neither is a current ISO 3166-1 code.
7 rules · 4 quality dimensions
Nothing reaches production without this step. Uncheck anything you would not stand behind. In the real product, the SQL is editable too.
customer_id, Must not be null
Completeness
0.0% null across 4,812 rows. Already clean, worth enforcing as a hard constraint before it drifts.
customer_id, Must be unique
Uniqueness
4,812 distinct values across 4,812 rows. Uniqueness holds today; the rule locks it.
tax_id, Must match the tax ID format
Validity
97.2% of non-null values match the pattern. 118 do not, mostly nine-digit strings missing the hyphen.
tax_id, Null rate must stay under 5%
Completeness
12.2% null against a 5% domain threshold. 587 rows. This is the largest single rule failure in the table.
email, Must be a parseable email address
Validity
4,359 of 4,422 non-null values parse. 63 do not: trailing semicolons and two addresses in one field.
status, Must be one of the accepted values
Validity
Exactly 4 distinct values observed, all within the expected set. Cardinality of 4 on 4,812 rows reads as a controlled vocabulary.
country_code, Must exist in the country reference set
Consistency
47 distinct codes. Two (XK, AN) are absent from ref.country, affecting 12 rows. Neither is a current ISO 3166-1 code.
7 of 7 approved, nothing will run
Approve at least one rule
Approved rules run against the table. Failures produce downloadable rejected records; scores register to the catalog.
97.7%
Overall score
93.9%
Completeness
100.0%
Uniqueness
98.7%
Validity
99.8%
Consistency
No rules approved, so nothing ran.
Back to approval7 rules executed · 780 rejected rows · scores registered to the catalog
No rules executed · nothing registered to the catalog
Runs governance natively inside Databricks Unity Catalog, no external suite in the loop.
Governs the AWS Glue Data Catalog and Lake Formation without a third-party catalog.
Stands up and operates open-source DataHub, a full catalog with no license line item.
Runs the day-to-day work of your catalog whenever you ask in plain language.
One command takes a new dataset from unknown to governed.
Discovery is read-only, takes five days, and costs nothing. We assess your environment and tell you honestly whether this agent is worth your time.
Book a Consultation