Not scraped, not modeled, not guessed. Here's the exact standard every row is held to.
Every fact is checked against a real, primary source for its domain — not a secondhand summary of one. For the company-landscape dataset, that means an official company site, an official press release, a regulatory filing, or reputable, named tech/business press (TechCrunch, Reuters, Bloomberg, SiliconANGLE, and similar). For the verification datasets, it means the actual system of record: NIST's CODATA database and real executed computation for math and physical constants, ClinicalTrials.gov and openFDA for clinical trials and drug/device safety, USASpending.gov for federal contracts, SEC EDGAR's XBRL filings for corporate financials, the published npm/PyPI artifact for code claims, actually-executed Python/Node.js for language and runtime semantics, Crossref/OpenAlex/Unpaywall for citations, Project Gutenberg's public- domain full text for literary attribution, the CIA World Factbook and the UK Companies House register for geographic and company/officer records, and Wikidata as a second independent check across most of these domains. We do not use paywalled proprietary databases (PitchBook, CB Insights) as sources, and we do not reproduce their data — that would be exactly the kind of unverifiable secondhand claim this dataset exists to avoid.
If a funding amount, founding year, or headquarters isn't publicly confirmed, the field is null. In the verification datasets, a claim that can't be confirmed against a source is marked unverifiable, not forced into true or false. Nothing is ever estimated, interpolated, or inferred from a "similar" company or claim. A confidently wrong number or verdict is worse than an honest gap — it's the difference between a dataset you can defend to your own stakeholders and one you can't.
Every row carries a source_urls field — the actual pages the facts were checked against. This is what most free, community-uploaded datasets don't give you: a way to verify the claim yourself rather than trust it blind.
The one-time export is a snapshot, versioned by month. The updates subscription re-verifies the full dataset and adds newly notable companies on a monthly cadence — funding rounds, acquisitions, and pricing changes move fast in this market, and a static export goes stale within a quarter.
We don't claim exhaustive coverage of every company in a category — new entrants and quieter regional players are inevitably missed in any snapshot. We claim that what is included is correct and sourced, which is a different and, we'd argue, more useful bar.