Newsletters




Bolstering Data Products for AI with CData and Prophecy


As more teams push AI into real use cases, a common challenge is emerging: Access to data is not the same as having data you can use. This is driving a shift toward data products, curated, trusted, and reusable datasets built to support analytics and AI in a consistent way.

But what does it take to build data products that deliver real value?

DBTA recently held a webinar, Data Products for AI: Delivering Context, Quality, and Reusability, with Jerod Johnson, director, technology evangelism at CData and Nathan Tong, Lead FDE at Prophecy, who shared how to design, govern, and deliver data products that support modern AI initiatives.

Every data product that works well made a small number of deliberate architectural decisions. Every one that doesn't is usually missing one of them, said Johnson.

There are 4 practices that make a data product:

Choose the delivery mechanism deliberately. Live/virtualized access or CDC-based replication—chosen by what the source and consumer actually need, not habit.

Build a semantic layer above the raw schema. Metadata alone isn't context—context is what lets a person or a model actually reason about the data.

Embed governance at the product boundary. Centralized policy, locally enforced, with lineage that travels with the data.

Architect for reuse from day one. A real data product serves more than one consumer without being rebuilt for each.

Johnson presented an architecture in practice which includes the following:

Delivery Decision: Standardize on one embedded connectivity layer across sources—a live-access decision made once, reused everywhere.

Context Layer: Metadata and sample rows feed an LLM that generates synthetic context—descriptions, entity suggestions, relationships—building a shared glossary.

Reuse: Finished models push to GitHub, Azure DevOps, or dbt—the same layer now feeds an MCP server for AI agent access too.

Choose the delivery mechanism the source allows—not the other way around, Johnson explained. However, governance needs to be part of the architecture and not an afterthought.

Consistent security and lineage rules applied the same way across every domain, even as day-to-day ownership stays decentralized.

According to Johnson, there is a checklist to consider before building a data product that includes asking the following:

  • Did you choose the delivery mechanism deliberately, or default to whatever was already connected?
  • Does someone outside your team understand what this data means without asking you?
  • If ownership is decentralized, is governance still centrally consistent?
  • Could a second team or an AI agent use this today without a new integration project?

Tong asked, can your business users prep data on your cloud data platform? Or are they using desktop tools?

Copilots are the productivity layer for cloud data platforms, noted Tong. By using Alteryx Import, you can move users and assets to Propehcy on your cloud data platform.

For the full webinar, featuring a more in-depth discussion, Q&A, and more, you can view an archived version of the webinar here.


Sponsors