Dataiku Shows the Flow. MyDataWork Shows the Work Around It.

Dataiku gives you a lot. The Flow is one of the most complete analytical surfaces in the market — a visual DAG of datasets, recipes, models, and outputs that captures an end-to-end pipeline from raw input to deployed model. You get real lineage out of the box: column-level dependencies, a data catalog across projects, and visibility into how data transforms through every recipe step. Add Dataiku Govern and you get governance workflows, sign-offs, metrics, attachments, and business initiatives on top.

Within Dataiku, that’s a strong story. The Flow itself is lineage. If your entire analytical world lived inside Dataiku, the Flow and Govern would get you a long way.

But it doesn’t. The work that feeds your Flow — and the work that consumes its output — often lives in tools Dataiku doesn’t naturally govern as part of the Flow experience. The Flow tells you how the project works, and Govern can add formal oversight if your organization has implemented it. What neither automatically gives the practitioner is a cross-tool view of the surrounding analytical work: the files that feed the project, the dashboards that consume it, the parallel work happening outside it, and the portfolio-level value story the business will eventually ask for.

That’s the gap MyDataWork is built to fill.

A scenario most Dataiku teams will recognize

Nadia Haddad is the lead data scientist for a subscription business’s churn-prediction practice. Her team owns the customer_churn Dataiku project — the Flow that ingests behavioral events, account history, and support signals, scores every account weekly, and writes the results to a table the rest of the business consumes.

Nadia knows the technical surface cold. She knows every recipe in the Flow, the input datasets, the model versions, and the column-level lineage from raw events to churn score. Dataiku’s catalog and her Flow give her all of that.

What neither tells her:

  • That the partner-supplied CSV her Flow ingests every Monday is maintained by hand on a colleague’s laptop — and if it stops arriving, the score silently goes stale
  • That the regional Tableau dashboard the retention team reviews every morning is built on her churn scores, in a tool her Dataiku catalog has no awareness of
  • That a sales analyst has been building a parallel propensity model in a Python notebook and a Power BI report, unaware her churn score already exists
  • That the VP of Customer Success is the real stakeholder for this work — and that the CFO will ask “what is the churn program actually saving us” in next quarter’s planning, with no portfolio-level answer ready

These are not failures of the Flow. They are context gaps around the Flow. And for most practitioners, that context lives one layer above any single platform.

What MyDataWork adds

Dataiku shows the Flow. MyDataWork shows the work around the Flow — and the business purpose above it. In practice that’s three things.

Cross-tool inventory. The Windows Connector reads Nadia’s Dataiku DSS project export from its .zip — pulling the project name and its recipe and dataset counts, badged as “Dataiku” — and catalogs it alongside the artifacts that sit outside the Dataiku catalog’s natural field of view: the SQL and Python notebooks in her project folders, the partner CSV, the staging SQL in her team’s GitHub repo, the Excel reference file her business partner maintains, and the Snowflake tables her Flow reads from. One inventory, every tool.

dataiku catalog 1400

The asset catalog brings every tool into one inventory — the customer_churn project (badged “Dataiku”) sitting alongside the Excel, SQL, Python, Power BI, Tableau, and Snowflake work around it.

Cross-tool lineage. For the SQL, Python, and BI assets around the project, MyDataWork infers dependency edges from the structural references in each file — the upstream-and-downstream picture that extends past any single tool. The Dataiku project sits in that same workspace as a first-class asset: MyDataWork reads the dataset and recipe references inside the project export and connects the Flow into the larger map automatically, alongside the work feeding and consuming it.

dataiku lineage 1400
The lineage view places the customer_churn project at the center of its cross-tool dependencies — upstream SQL and Python feeding it, downstream Power BI and Tableau consuming its output — every edge auto-inferred from the table references inside the work itself. The cross-tool picture no single platform catalog assembles.

Business context. This is the layer technical tools often don’t capture consistently across the full practitioner workflow — and where much of the day-to-day value becomes visible.

Use case documentation. Nadia links the churn project to a documented use case: Reduce voluntary churn in the mid-market segment. She sets the baseline (5.2% monthly churn), the current state (4.1%), and the target (3.5%). She records the estimated value ($380K/year in retained revenue) and updates the realized value as the program delivers.

02 usecase detail 1400

Nadia’s churn-reduction use case documents the baseline, current state, and target, with realized value tracking the program as it delivers. The VP of Customer Success is linked as the stakeholder.

Stakeholder linking. The VP of Customer Success is added as the use case stakeholder, visible on the use case and on the lineage view, so anyone touching the underlying assets knows who depends on them. The Workspace Agent’s missing-stakeholder check flags any active use case that doesn’t have someone assigned.

On-demand workspace review. Nadia runs the Workspace Agent when she wants a structured read of her workspace. It surfaces patterns across six categories — duplicate parallel work, missing stakeholders, stale use cases, high-value work tracking below estimate, dependency breaks, and assets that have quietly become hidden infrastructure. Once the parallel propensity notebook is in her catalog, the agent flags it as likely duplicate work the next time she runs Analyze.

dataiku suggestions 1400

The Workspace Agent surfaces findings across categories. One finding recommends consolidating the parallel propensity work into Nadia’s documented churn use case — the kind of duplication a Flow can’t see, because it lives in a different tool.

Portfolio reporting. When the CFO asks what the churn program is worth, Nadia generates a portfolio export — linked assets, documented use case, baseline-to-current improvement, value realized to date. The conversation moves from “we built a model” to “here’s what it produces.”

Where MyDataWork sits relative to your existing stack

MyDataWork is not an enterprise data catalog, and it’s not a replacement for Dataiku Govern. Tools like Atlan, Alation, and Collibra — and Govern itself — serve top-down governance for teams rolling out catalog discipline at scale. The Dataiku community has asked about external-catalog integration for years; the usual answers are plugins, custom scripts, or ingestion into one of those enterprise catalogs.

MyDataWork is built for the practitioner: the lead data scientist, the analyst, the analytics engineer building a working practice from the bottom up. Your Flow keeps doing what it does. Your Govern setup, if you have one, stays in place. MyDataWork adds the work-context layer above both — the documentation of why each project matters and what it produces — without an enterprise catalog implementation.

Try it on the work that matters most

MyDataWork’s Explorer plan is 90 days, free, no credit card. For a Dataiku-centric practitioner, it is enough room to start with one important project and the assets around it: the dashboards stakeholders open, the notebooks and reference files the work depends on, and the outcome metrics leadership cares about.

Build the documentation practice around those, link the people who care, set the outcome metrics that matter, and let the agent surface what you’d miss. When you outgrow it, the paid plans lift the asset caps and your work carries forward — the workspace you built becomes the foundation you expand from.

Where this matters now

Dataiku is where your organization builds models. MyDataWork is where the practice around those models becomes legible: the surrounding assets, stakeholders, outcomes, duplication, dependencies, and value story.

Every organization is being asked to “do something with AI.” Whether those agents come from Dataiku, your BI platform, your cloud provider, or an internal build, they will need more than pipeline metadata — they’ll need to know which analytical products serve which business outcomes, who depends on what, and which work is worth protecting, migrating, or retiring. Your Flow holds the pipeline. MyDataWork holds the context any agent needs to reason well about it. Without that layer, AI initiatives produce confident output disconnected from your actual analytical priorities. With it, the agents arriving over the next year have something real to work from.

Start here. Email only, no credit card. Connect your Dataiku project export and the files around it, and you’ll be cataloging your analytical surface within minutes.

Related Posts

Scroll to Top