Roughly 62% of data and analytics project failures are attributed to organisational and process issues rather than technical ones. Understand the processes before you build the data models, the analytics, the reports, the machine learning pipelines, or the AI.

The photograph above is from a process discovery workshop we ran recently with a multi-brand e-commerce business (details blurred to protect confidentiality). A wall, four colours of sticky notes, and a room full of people mapping their order-to-cash flow end to end for the first time in years, which was facilitated by our team at Lashan Digital .
What struck me, as it does in almost every engagement, was how much of what governs the data exists as tacit knowledge, understood by the people who operate the process, encoded nowhere else. That knowledge is accessible. It is in the room. The issue is that data programmes are not designed to surface it. Discovery, in a data context, means understanding the source systems. The process that created the data is treated as someone else’s concern, and so the programme is built on an incomplete picture of what the data actually represents.

Data programmes treat discovery as a technical exercise by design. Source systems are catalogued. Schemas are profiled. Data dictionaries are populated, or at least started. The assumption embedded in this approach is that the data contains everything you need to understand the data, that the structure of the tables reflects the logic of the business. It does not. The structure reflects what was built, when it was built, by people who have often since left. The logic lives elsewhere.
The reason data programmes default to technical discovery is partly structural and partly incentive-driven. Data teams are assessed on delivery: pipelines built, models shipped, dashboards released. Discovery work that cannot be directly mapped to a technical output is the first thing compressed when budgets tighten. Most programmes begin with a technical brief (migrate this platform, build this warehouse, connect these systems) which pulls engagement toward architecture and infrastructure from the outset. The business process sits outside that frame, treated as something to be resolved in time rather than first. By the time the data model is being validated, the process assumptions baked into its design are already structural.
Every field in a source system has an origin, a business decision, an encoded rule, or an operational process, most of which predate the current team and few of which are documented. The data describes the business as it has actually operated: the workarounds, the manual interventions, and the logic that lives in institutional memory rather than in any specification.

The types of logic that accumulate in enterprise systems over time are varied and consequential. Revenue recognition conventions agreed informally and never revisited. Cost allocation rules that reflect a business structure which no longer exists. Exception handling that was designed for edge cases and became standard practice because the underlying problem was never fixed. Fields that meant one thing at implementation and something different now. These are not data quality issues in the conventional sense. They are encoded history. Schema profiling surfaces the field. It cannot surface what the field means, how its meaning has shifted, or what business event it is actually recording.
That knowledge requires a different kind of engagement, one with the people who operate the process, not the people who manage the platform. This is the gap that process discovery is designed to close. It is a prerequisite for technical discovery, not a supplement to it.
The session in that photograph ran for half a day. Three things came out of it that no schema profiling exercise would have found. A margin calculation that was structurally an estimate presented as fact. A pricing discrepancy on an insurance product that existed nowhere in any report. A supplier cost recovery process that had never been formalised and was absorbing losses silently. All three were identified by asking the people in the room to write down, in past tense, every event between a customer placing an order and the finance team closing the books. A facilitated, low-tech process that, in most data programmes, never happens at all.

What process discovery surfaces is not a reflection of how well a business is run. It is a reflection of how much has accumulated in institutional memory rather than in documented systems. Every organisation has it. The structured process exists to bring it into the open so it can be designed around, not to catalogue what is missing.
The gap persists because technical milestones drive data project timelines. Process understanding produces no artefact that maps to a sprint or a milestone, so it gets compressed or skipped. The consequences surface in UAT: a disputed number, a metric with three interpretations, a report that does not reconcile. Re-engineering a data model at that stage costs multiples of what a structured discovery session would have cost at the outset.

Understand the processes before you build the data models, the analytics, the reports, the machine learning pipelines, or the AI. Go in with the people who operate the process in the room, not just the people who manage the platform. The output is a documented event map that drives entity identification, informs grain decisions, and anchors metric definitions.
The data will only ever be as good as the understanding behind it.
#DataStrategy #DataEngineering #ProcessDiscovery #DataGovernance #DigitalTransformation #EnterpriseArchitecture #LashanDigital
