Buying data from a source you don’t know means taking the seller’s word for it. At DataDiti nothing enters the catalogue unchecked — six layers of checks run before anything can go live, and we show the results as they came out.

The data itself may be equally good either way. What differs is whether you have any way of knowing that before you pay.
Run in order over every incoming asset, without exception — including assets from long-standing sellers, every time they submit a new version.
Data types, value ranges, formats, and cross-column consistency are tested against the rules in force in the data’s jurisdiction.
28 of 28 versions passed this layer Catches: swapped columns, impossible dates, region codes that never existedMatched against trusted reference sets per jurisdiction: master region codes, administrative boundaries, entity registries — and the references are versioned.
27 of 28 versions passed this layer Catches: data still using the old region list after a jurisdiction splitWe look for overlap with other datasets by identity, coordinates, or name, then measure the match rate.
28 of 28 versions passed this layer Catches: figures that match no other source — and a seller cannot control third-party dataBenford’s law, dead-data detection, timestamp clustering, distributions that are too tidy, and subtle duplication.
28 of 28 versions passed this layer Catches: fabricated rows, old datasets whose dates were simply shiftedA new version is compared with the previous one: missing columns, row-count jumps, distribution shifts, units changed without a changelog.
27 of 28 versions passed this layer Catches: silent changes that break existing buyers’ integrationsOrigin, chain of ownership, and resale rights are reviewed — including anachronism tests and metadata forensics.
0 of 28 versions passed this layer Catches: data with no right to be resold, or whose age does not add upSome checks are hard gates: failing one means rejection, not merely a low score. Assets that fall here never appear in your search results.
Every product states its verification level on the catalogue card, before you even open it. You pay for the depth you actually need.
Metadata and data dictionary complete; full automated checks not yet passed.
0 products in the catalogue Enough for: early exploration, feasibility checks, data you will validate yourselfPassed six automated layers — structure, anchors, triangulation, anomalies.
24 products in the catalogue Enough for: business analysis, planning, most operational needsPlus manual checks: 30–50 record ground-truth sampling and legal document review.
2 products in the catalogue Enough for: high-value decisions, regulatory compliance, evidence before third partiesThere is no blanket promise of "clean" — that is an unbounded promise no one can be held to. What you get is defined in tiers, and the level is stated on every product.
| Level | What has been done | In the catalogue |
|---|---|---|
| L0 | Exactly as supplied by the seller. Kept forever as evidence, never sold. | never sold |
| L1 | Schema, types, units, and region codes standardised. Original columns retained. | 21 products |
| L2 | Values validated against their domain, duplicates removed, outliers flagged. | 2 products |
| L3 | Enriched: geocoding, official region codes, canonical IDs, joined with other layers. | 3 products |
We publish the formula. If the only way to raise a score is to genuinely improve the data, opening the formula costs no one anything.
How fully the required columns are populated and how wide the coverage is. Placeholders count as empty.
How far the data period is from today, measured against the seller’s own stated refresh schedule.
Agreement between periods: missing columns, row-count jumps, distribution shifts, units changed without a changelog.
Match rate on the overlap with other sources. With no overlap the status reads “not cross-verified yet” — not a zero score.
| Data category | Completeness | Freshness | Consistency | Cross-accuracy |
|---|---|---|---|---|
| Demography & Population | 30% | 20% | 10% | 40% |
| Geospatial & Boundaries | 30% | 5% | 20% | 45% |
| Economy & Trade | 20% | 40% | 10% | 30% |
| Agriculture & Food | 15% | 40% | 10% | 35% |
| Infrastructure & Transport | 30% | 5% | 20% | 45% |
| Environment & Climate | 15% | 40% | 10% | 35% |
| Public Health | 30% | 20% | 10% | 40% |
| Property & Land | 20% | 40% | 10% | 30% |
Living assets need continuous supervision, not a single check at the start.
No amount of checking catches everything. What matters is not a claim of never being wrong, but a clear route when it happens.
Trap rows, precision variation, unique row ordering, and metadata watermarks. We state this openly in the terms because the deterrent effect outweighs the tracing ability — and because that is what makes good data willing to come here.