Home
Loading

aVenture is in Alpha: During this preview period, you should expect the research data to be limited and may not yet meet our exacting standards. We've made the decision to provide early access to our data to showcase the product as we build, but you should not yet rely upon it alone for your investment decisions.

aVenture is in Alpha: During this preview period, you should expect the research data to be limited and may not yet meet our exacting standards. We've made the decision to provide early access to our data to showcase the product as we build, but you should not yet rely upon it alone for your investment decisions.

Get in touch

  • Contact

  • Request a demo

  • Request data updates

  • Add a company

Research

  • Companies

  • Investors

  • People

aVenture

  • Sitemap

  • Feature requests

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Blog/Research Platform

Why Provenance Is the Foundation of Trustworthy Company Data

Data provenance — knowing where every fact came from — is what separates verifiable company data from guesswork. Here's how we built it in from day one.

William A. Callahan, CFA
William A. Callahan, CFACEO at aVenture
May 3, 2026·Updated Jul 13, 2026·8 min read

Ask a simple question about a private company — how many people work there, what it actually sells, who invested in the last round — and you can usually find an answer. The harder question, and the one that matters when real money or a real decision is on the line, is: where did that answer come from?

That question is what data provenance is about. Provenance is the recorded origin and history of a fact: what source it came from, who wrote it down, when, and what has happened to it since. It is the difference between verifiable company data and confident-sounding guesswork. And in private company research — where there are no mandatory quarterly filings, no audited statements, and no single authoritative registry — provenance is not a nice-to-have. It is the foundation everything else stands on.

We built aVenture, a research platform for private company intelligence, around that conviction. This post explains what field-level provenance means in practice, why most company data you encounter doesn't have it, and how we engineered our platform so that every fact can answer for itself.

The problem: numbers with no origin

Most company databases present facts as flat values. An employee count is a number in a cell. A funding total is a figure on a profile. There is no visible answer to the questions a careful researcher immediately asks:

  • Did this come from the company itself, a news article, or someone's estimate?
  • When was it last checked?
  • Has anyone — including the company — disputed it?
  • Was it entered by a person, imported from a feed, or generated by software?

When data has no origin, every downstream use of it inherits an invisible risk. An analyst cites the number in a memo. The memo informs a decision. Months later, the number turns out to have been a three-year-old estimate scraped from a stale page — and nobody can reconstruct where it entered the chain. That failure mode is common precisely because provenance is usually discarded at the moment of data entry, and it can never be reconstructed afterwards.

Due diligence data quality is, at its core, a provenance problem. The diligence standard isn't "we found a number"; it's "we can show where the number came from and why we believed it."

What field-level provenance looks like

On our platform, provenance is not an attribute of a company profile as a whole. It is recorded at the level of the individual fact — because a single profile is assembled from many sources of very different reliability, and treating them as one blob hides exactly the distinctions that matter.

Every fact in our research graph carries four things:

1. The kind of source it came from

A fact ingested from a company's own website is not the same as a fact from a news article, which is not the same as a fact from an independent blog, a staff review, an AI research agent's finding, or a correction submitted by a user. We record the source category with the fact itself, so a researcher can weigh it accordingly. First-party statements carry a different kind of authority — and a different kind of bias — than third-party reporting, and the data model should preserve that distinction rather than flatten it.

2. Its verification status

Facts are explicitly marked as unconfirmed, confirmed, or disputed. Disputed facts additionally record who disputes them — the company itself, a related party, or a third party. This matters more than it might sound: a headcount figure disputed by the company is a very different signal than one disputed by a competitor, and a researcher deserves to see both the claim and the disagreement rather than whichever one happened to be written last.

3. The actor that wrote it

Every write to the graph records whether it was made by a human staff member or an AI research agent — and when it was an agent, the specific model that produced it. We think this is non-negotiable in an era when much research data is machine-assisted. If software contributed a fact, the record should say so plainly, and it should say which software. Accountability that stops at "the system added this" isn't accountability.

4. When it happened

Timestamps on creation and on every subsequent change, so the age of a fact is always visible — not just the age of the profile it sits on.

Corrections that don't erase history

Provenance would be incomplete if it only covered how a fact arrived. It also has to cover what happened afterwards.

Every change to every record on our platform lands in an audit trail: what changed, what the value was before, who or what changed it, and when. When a fact gets corrected — and in private company research, facts get corrected constantly, because companies pivot, teams grow, and old reporting goes stale — the correction doesn't silently overwrite history. The previous value, its source, and the change itself remain part of the record.

This has a practical payoff beyond tidiness. If you cited a figure in March and it reads differently in June, you can see exactly when it changed and why. Research you did in the past remains explainable in the present. For anyone whose work product is a memo, a model, or an investment committee document, that's the difference between diligence you can defend and diligence you have to shrug about.

Keeping junk out: validation at the moment of writing

There is a second, quieter half of data quality that provenance alone doesn't solve: making sure a fact is well-formed before it enters the graph at all.

Every category of research fact on our platform is governed by a validation contract that is enforced at write time. Numeric facts must fall within allowed ranges or take allowed values. Categorical facts must use a defined set of options rather than free-form strings. Even narrative research text is checked for shape — length and structure — before it is accepted.

The effect is that free-form junk cannot quietly accumulate. A malformed value doesn't become someone else's cleanup project six months later; it is rejected at the door. Combined with provenance, this gives the graph a property we care a lot about: everything in it is both traceable (you know where it came from) and well-formed (you know it satisfied the rules for its type when it was written).

Staleness is a data point too

A true fact from four years ago can be more misleading than no fact at all. So beyond source and structure, our records carry currency signals: facts and links are flagged as current or historical, and companies themselves carry an operating status, so a dissolved or acquired company doesn't masquerade as a going concern in your results.

Treating staleness explicitly — rather than letting old data sit indistinguishable from fresh data — is part of the same discipline. Provenance tells you where a fact came from; currency tells you whether it still deserves your trust today.

Why we hold ourselves to this

We've written before, in our draft privacy policy, about how we see public research data: as something that serves the public interest by helping investors and researchers make informed decisions. We take that framing seriously, and it cuts both ways. If research data is going to inform real decisions, then the burden is on the platform to make that data inspectable — to show its work, fact by fact, the way we'd expect any serious researcher to.

Source-backed research is the standard we'd want as users, and it's the standard we think the private markets deserve. A claim without a source is an opinion wearing a suit.

What this means for your research

Concretely, building on provenance-first data means:

  • You can cite what you find. Facts arrive with their origin attached, so moving from "the database says" to "according to the company's site as of this date" takes no extra work.
  • You can defend your diligence. The chain from source to fact to your memo is reconstructible, even months later, even after corrections.
  • You can weigh conflicting information honestly. Disputed facts show the dispute instead of hiding it, and first-party claims are distinguishable from third-party reporting.
  • You can trust the freshness signal. Stale facts and inactive companies are flagged as such rather than blending in.

We have not commercially launched yet — we're building this in the open, deliberately, because we think getting the foundation right matters more than shipping a bigger pile of unattributed data. If the approach resonates with how you think company research should work, we'd genuinely like you to try it early and tell us where it falls short.

Join the research preview waitlist at aventure.vc/free-research.

Filed under

Data Provenance·Due Diligence·Data Quality·Private Markets

About the author

William A. Callahan, CFA
William A. Callahan, CFACEO at aVenture
View Research Profile→

Most read

1

aVenture vs. AlphaSense: Market Intelligence Search vs. Venture Research Graph

May 18, 2026·4 min read
2

aVenture Joins Techstars 2025 Cohort

Nov 28, 2025·1 min read
3

Introducing Advanced Comparables Analysis

Dec 17, 2025·2 min read
4

A Draft Privacy Policy (v0.1): Public Research Data vs. User Data

Mar 2, 2026·4 min read
5

Understanding Venture Capital Valuations in 2025

Dec 15, 2025·1 min read

Recent

1

aVenture vs. Preqin Pro: Company-Level vs. Fund-Level Private Market Data

Jul 7, 2026·4 min read
2

aVenture vs. PitchBook: Which Private Market Research Platform Fits Your Workflow?

Jul 2, 2026·5 min read
3

aVenture vs. Crunchbase Pro: Private Company Data Platforms Compared

Jun 27, 2026·4 min read
4

Product and Service Intelligence: Knowing What a Company Actually Sells

Jun 22, 2026·7 min read
5

aVenture vs. CB Insights: Tech Intelligence Platforms Compared

Jun 17, 2026·4 min read

Most read

1

aVenture vs. AlphaSense: Market Intelligence Search vs. Venture Research Graph

May 18, 2026·4 min read
2

aVenture Joins Techstars 2025 Cohort

Nov 28, 2025·1 min read
3

Introducing Advanced Comparables Analysis

Dec 17, 2025·2 min read
4

A Draft Privacy Policy (v0.1): Public Research Data vs. User Data

Mar 2, 2026·4 min read
5

Understanding Venture Capital Valuations in 2025

Dec 15, 2025·1 min read

Recent

1

aVenture vs. Preqin Pro: Company-Level vs. Fund-Level Private Market Data

Jul 7, 2026·4 min read
2

aVenture vs. PitchBook: Which Private Market Research Platform Fits Your Workflow?

Jul 2, 2026·5 min read
3

aVenture vs. Crunchbase Pro: Private Company Data Platforms Compared

Jun 27, 2026·4 min read
4

Product and Service Intelligence: Knowing What a Company Actually Sells

Jun 22, 2026·7 min read
5

aVenture vs. CB Insights: Tech Intelligence Platforms Compared

Jun 17, 2026·4 min read