Home/Agents/
Incept PipelineIQ™
Data Quality
 · 

Incept PipelineIQ™

Quality checks that run inside the pipeline rather than after it. The agent implements stage gates at each hop (source, staging, warehouse, mart and consumption), so a bad load is stopped before it reaches anybody, instead of being discovered in a dashboard a week later.

  • Azure Data Factory
  • AWS Glue
  • Apache Airflow
  • dbt
  • Databricks Workflows
[Interactive demo]

Walk the pipeline yourself.

Sample data, in your browser. A faithful simulation of the real flow. Approve or reject at step four and watch the score change.

Incept DataIQ™

analytics.sales.customer_master

Sample data · not a client system

Profiling analytics.sales.customer_master

Eight dimensions across every column, five of them shown here. Nothing is written; this is a read.

Column
Type
Null %
Distinct
Detected pattern
Signal
customer_id
string
0.0%
4,812
^C[0-9]{6}$
clean
customer_name
string
0.2%
4,780
free text
clean
tax_id
string
12.2%
4,201
^[0-9]{2}-[0-9]{7}$
issue
country_code
string
0.0%
47
ISO 3166-1 alpha-2
review
email
string
8.1%
3,944
RFC 5322
issue
created_date
date
0.0%
1,204
ISO 8601
clean
status
string
0.0%
4
enum, 4 values
clean
credit_limit
decimal
31.7%
892
numeric
review

4,812 rows · 8 columns · 2 columns flagged · 2 for review

Business metadata, drafted

Table and column descriptions written from the profile and the column context. Every one is a draft until a steward accepts it.

Object
Drafted description
customer_master
Master record for every business customer the organization sells to. One row per customer, keyed on customer_id. Sourced from the order system and enriched with credit attributes.
customer_id
System-generated unique identifier for a customer. Format C followed by six digits. Primary key.
tax_id
Government-issued tax identification number, formatted NN-NNNNNNN. Required for any customer invoiced in the current fiscal year.
country_code
ISO 3166-1 alpha-2 country code for the customer’s registered address.
credit_limit
Approved credit ceiling in reporting currency. Null where no credit assessment has been completed.

1 table + 4 columns described · awaiting steward

Rules recommended, with the evidence

Seven rules proposed from the profile. Each carries the proof that produced it, so approval is a judgment call rather than a leap of faith.

customer_id, Must not be null

Completeness

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE customer_id IS NULL;
Evidence

0.0% null across 4,812 rows. Already clean, worth enforcing as a hard constraint before it drifts.

customer_id, Must be unique

Uniqueness

SELECT customer_id, COUNT(*) AS occurrences
FROM analytics.sales.customer_master
GROUP BY customer_id
HAVING COUNT(*) > 1;
Evidence

4,812 distinct values across 4,812 rows. Uniqueness holds today; the rule locks it.

tax_id, Must match the tax ID format

Validity

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE tax_id IS NOT NULL
AND tax_id NOT RLIKE '^[0-9]{2}-[0-9]{7}$';
Evidence

97.2% of non-null values match the pattern. 118 do not, mostly nine-digit strings missing the hyphen.

tax_id, Null rate must stay under 5%

Completeness

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE tax_id IS NULL;
Evidence

12.2% null against a 5% domain threshold. 587 rows. This is the largest single rule failure in the table.

email, Must be a parseable email address

Validity

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE email IS NOT NULL
AND email NOT RLIKE '^[^@\\s]+@[^@\\s]+\\.[^@\\s]+$';
Evidence

4,359 of 4,422 non-null values parse. 63 do not: trailing semicolons and two addresses in one field.

status, Must be one of the accepted values

Validity

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master
WHERE status NOT IN ('ACTIVE','INACTIVE','PENDING','BLOCKED');
Evidence

Exactly 4 distinct values observed, all within the expected set. Cardinality of 4 on 4,812 rows reads as a controlled vocabulary.

country_code, Must exist in the country reference set

Consistency

SELECT COUNT(*) AS failed_records
FROM analytics.sales.customer_master cm
LEFT JOIN ref.country c
ON cm.country_code = c.alpha_2
WHERE c.alpha_2 IS NULL;
Evidence

47 distinct codes. Two (XK, AN) are absent from ref.country, affecting 12 rows. Neither is a current ISO 3166-1 code.

7 rules · 4 quality dimensions

Steward approval

Nothing reaches production without this step. Uncheck anything you would not stand behind. In the real product, the SQL is editable too.

customer_id, Must not be null

Completeness

Evidence

0.0% null across 4,812 rows. Already clean, worth enforcing as a hard constraint before it drifts.

customer_id, Must be unique

Uniqueness

Evidence

4,812 distinct values across 4,812 rows. Uniqueness holds today; the rule locks it.

tax_id, Must match the tax ID format

Validity

Evidence

97.2% of non-null values match the pattern. 118 do not, mostly nine-digit strings missing the hyphen.

tax_id, Null rate must stay under 5%

Completeness

Evidence

12.2% null against a 5% domain threshold. 587 rows. This is the largest single rule failure in the table.

email, Must be a parseable email address

Validity

Evidence

4,359 of 4,422 non-null values parse. 63 do not: trailing semicolons and two addresses in one field.

status, Must be one of the accepted values

Validity

Evidence

Exactly 4 distinct values observed, all within the expected set. Cardinality of 4 on 4,812 rows reads as a controlled vocabulary.

country_code, Must exist in the country reference set

Consistency

Evidence

47 distinct codes. Two (XK, AN) are absent from ref.country, affecting 12 rows. Neither is a current ISO 3166-1 code.

7 of 7 approved, nothing will run

Execute approved rules

Approve at least one rule

Executed on schedule

Approved rules run against the table. Failures produce downloadable rejected records; scores register to the catalog.

97.7%

Overall score

93.9%

Completeness

100.0%

Uniqueness

98.7%

Validity

99.8%

Consistency

Rule
Dimension
Result
Failed rows
customer_id, Must not be null
Completeness
PASS
0
customer_id, Must be unique
Uniqueness
PASS
0
tax_id, Must match the tax ID format
Validity
FAIL
118
tax_id, Null rate must stay under 5%
Completeness
FAIL
587
email, Must be a parseable email address
Validity
FAIL
63
status, Must be one of the accepted values
Validity
PASS
0
country_code, Must exist in the country reference set
Consistency
FAIL
12

No rules approved, so nothing ran.

Back to approval

7 rules executed · 780 rejected rows · scores registered to the catalog

No rules executed · nothing registered to the catalog

[HOW IT WORKS]

The pipeline, step by step.

  1. Instrument the hops Checks placed at source, staging, warehouse, mart and consumption.
  2. Gate Go / no-go decisions on counts, completeness and rule outcomes at each stage.
  3. Halt or continue A failed gate stops the pipeline rather than propagating the problem downstream.
  4. Route the failure The right owner is notified with the failing records attached.
  5. Publish Results land in the results store and the catalog scorecard.

[CAPABILITIES]

What it does.

  • Native to your orchestrator. Implemented as Data Factory activities, Glue jobs, Airflow tasks or dbt tests, not a bolt-on.
  • Parameterized, not hard-coded. Checks configured from an operational data model, so a new domain is configuration.
  • Stage gates, not surprises. Every hop can stop the pipeline. Bad data does not become somebody else’s problem.
  • Written back to governance. Scores publish to the catalog so quality is visible where people find the data.
[RELATED]

Also in Data Quality.

Incept DataIQ™

Profiles any table, drafts the metadata, writes the rules, and proves each one.

Incept PulseDQ™

Watches quality scores over time and tells you what is drifting, and what to do.

Incept FabricIQ™

Profiling and rule generation native to Microsoft Fabric and OneLake.

Incept RedshiftIQ™

The DataIQ pattern, running natively on Redshift and Athena.

Incept StewardFlow™

Reads a failed rule, works out why, and routes it to the person who owns it.

See it on your data.

Discovery is read-only, takes five days, and costs nothing. We assess your environment and tell you honestly whether this agent is worth your time.

Book a Consultation