171 real AI/ML companies here, cited to a real source, not guessed, plus verified facts across math, science, history, language, law, medicine, government, finance, and code elsewhere in the catalog. Use it to test whether your model is hallucinating on any topic, or to ground an eval set, a RAG pipeline, or an agent in facts that are verifiably true.
The remaining 19% isn't filled in with a plausible-sounding number — it's marked null. A confidently wrong figure is worse than an honest gap.
Smaller, fixed-scope datasets for verification, eval, and agentic training — same sourcing standard, verified rather than generated.
Custom scope, a different vertical, or a specific data shape for your pipeline.