PYX data product agentic engineering
← Field notes·03·Discovery·6 min

Your semantics are already written. They are in a spreadsheet.

The most complete model of your business isn’t in your warehouse. It’s in a workbook your finance lead has kept for nine years. Stop calling it shadow IT.

The most complete model of your business isn’t in your warehouse. It’s in a workbook your finance lead has been maintaining for nine years.

And every modernisation programme I’ve joined has started by deciding to ignore it. We call it shadow IT. We put it on a decommission list. Then we run interviews to recover, from human memory, logic that’s sitting in column AK, fully specified, and reconciled against the general ledger every month since 2017.

Read the formulas as what they are and the workbook stops being a liability. A nested IF is an exclusion policy. A SUMIFS is a grain plus a filter. A named range is a concept with a definition and a scope. A column of hardcoded rate-class codes with a comment saying “per Dave, exclude these” is a business rule with an owner and a provenance, which is more than most of your certified tables can claim.

The real reason we ignored it

Throughput. Nobody could read four thousand formulas, resolve the cross-sheet references and cluster them into concepts inside the time a discovery phase allows. So we substituted something cheaper: ask the person what the workbook does. They tell you the eight rules they remember. The other forty walk out of the building with them.

That constraint is gone. Reading four thousand formulas and proposing what they mean is exactly the kind of work a model does well: high volume, pattern-dense, with a human sitting immediately downstream to refuse the wrong ones. The unlock isn’t that the AI understands your business. It’s that it can afford to read all of it. You never could.

Discovery stops being an interview about the artifact and becomes a reading of it, with the interview saved for what the artifact can’t say.

The practice

Every formula is a candidate, never a conclusion. Each proposal arrives with the cell it came from. If a proposed concept can’t point at its evidence, it’s an invention, and it gets refused.
Cluster before you argue. Forty variations of the same margin calculation are one concept with a versioning problem, and seeing them side by side is usually the moment somebody says “those two should never have differed”.
Score coverage instead of ticking a checklist. “This model answers 34 of the 41 questions the workbook answers” is more honest than any signed-off requirements document, and it names the seven arguments you still have to have.
Bring the reports in too. An existing Power BI page is a specification. Every card on it is a measure somebody relies on.

There’s a role change hiding in here. The requirements-gathering workshop, six people in a room reconstructing from memory what a system already states in full, doesn’t survive this method. It shouldn’t. What replaces it is shorter, sharper and much harder to fake: an analyst arguing with a page of evidence-linked proposals, ruling on them one at a time.

The workbook never becomes your system of record. It becomes what it always was and was never allowed to be: the specification. Once its rules are claims on a requirement, bound to concepts, checked by controls, you can finally decommission it for the right reason. Not because it was unsanctioned. Because everything it knew now lives somewhere that can be tested.

Where this is today. Requirements, claims and evidence-linked concept review are live. Direct workbook ingestion, parsing formulas straight into candidate concepts, is in design. Today the workbook is what you read from and cite.
Start a conversation

Have a data problem in mind?