Research & Methods

Data Science Has a Science Problem

Why more data, better dashboards, and artificial intelligence cannot substitute for measurement and scientific reasoning

Why data abundance, sophisticated analytics, and AI cannot replace sound measurement, scientific reasoning, and defensible inference.

More Data, Less Certainty

Organizations have never had more data or more sophisticated ways to analyze it. Employee sentiment can be tracked continuously. Customer behavior can be observed across millions of interactions. Executive dashboards update in real time. Artificial intelligence can interrogate datasets, synthesize info, and generate analyses in seconds.

However, more data has not necessarily made organizations more scientific.

That distinction matters. Being data-driven means using information. Being scientific means being disciplined about what that information allows us to conclude. Modern organizations have become exceptionally good at the former, while frequently neglecting the latter.

That is the science problem hiding inside data science.

The academic–practitioner gap has long reflected one dimension of this problem. Research and practice do not simply possess different levels of scientific knowledge; they often operate through different incentives, timelines, and channels that knowledge spreads through and becomes useful (Cascio & Aguinis, 2008; de Man, 2022).

Researchers can become increasingly specialized and removed from organizational practice, while practitioners and organizational leaders can become increasingly removed from the methodological principles underlying the evidence they use for decision making.

Those who want answers are often leading the ones with the answers in whatever direction they want to go.

The result is an uncomfortable question: in the rush to become data-driven, have organizations compromised scientific rigor?

Data is everywhere and readily available, but availability is not the same thing as evidence. A number does not become meaningful merely because it can be measured, visualized, or placed in front of an executive team.

Too often, data has become less of a mechanism for challenging assumptions than one for validating them. The presence of data itself can become a symbol of rigor; something that gives an argument the appearance of objectivity whether or not the underlying inference deserves it.

The “Science” in Data Science

Across industries from big tech to nonprofit research organizations, decision-makers often encounter data primarily through a graph, dashboard, or simple descriptive stat (likely a non-inferential percentage).

That is understandable. Executives are generally not scientists and cannot be assumed to reasonably inspect models or assumptions as if they are.

The danger comes when slidedecks of data visuals and graphs equate to understanding.

The data is believed at face value: this is where the dot or line is on the chart, so this is how it must be responded to.

The dot moved. Something must have happened.

The line went up. Something must be working.

The line went down. Someone needs an explanation.

That very sentiment is the opposite of rigor.

Data science has more emphasis today than ever before, but organizational practitioners can become heavily preoccupied with the data while treating the science like the awkward sibling in the room that nobody particularly wants to hear from.

It can feel almost absurd to remind people working in analytics and organizational leadership what science is. The problem, however, is not whether people know the textbook definition. It is whether they consistently behave according to its logic when interpreting evidence.

Why the Scientific Method Exists

Simply put, the scientific method is a disciplined way of testing what we think we know.

It asks us to evaluate ideas through systematic observation, analysis, comparison, and skepticism rather than relying exclusively on intuition or surface-level judgment.

That discipline exists for a reason. Human beings are remarkably good at constructing explanations after something happens; we naturally cannot help but reach for the low-hanging explanatory fruit. We observe an outcome, narrate a plausible story, and quickly profess the narrative as explanation.

Within organizations, that tendency can become especially consequential, because leaders are constantly being asked to react to changing numbers.

Suppose employee engagement rises this month.

The tempting conclusion is immediate: whatever leadership has been doing must be working.

Then the following month, the pulse-survey dot dips and suddenly everyone wants to know what went wrong.

But what actually caused either movement?

Maybe leadership did.

Maybe compensation changed.

Maybe the work-life balance changed.

Maybe the respondent traits changed.

Maybe seasonality matters.

Maybe a major organizational event occurred.

Maybe the instrument itself is unreliable.

Maybe the difference is essentially noise and error, which can even be accounted for with greater precision than most realize if they bother to dig deep enough.

A chart can tell you that a number moved, but it generally cannot explain why.

Causal inference requires evidence capable of ruling out plausible alternative explanations. Observing some movement over time, even when that movement is displayed beautifully on a Microsoft Power BI graph or dashboard, does not accomplish that by itself (Shadish, Cook, & Campbell, 2002).

Observation ≠ Explanation
Observed KPIEmployee engagement ↑
LeadershipCompensationWorkloadRespondent compositionSeasonalityOrganizational eventsMeasurement error / noise

A movement in the metric is an observation. It does not establish which explanation, if any, is causal.

Measurement Comes Before Analysis

The problem begins even earlier.

Before asking why engagement increased, there is a more basic question:

What exactly are we calling “engagement?”

Is it enthusiasm?

Effort?

Job satisfaction?

Commitment?

Motivation?

Energy?

Intent to stay?

That question is not semantic nitpicking. Engagement is a construct: something we infer rather than directly observe. Calling a number an “engagement score” does not by itself establish that the number measures engagement.

That interpretation requires a defensible chain from theory to measurement to validation. Construct validity concerns whether the evidence actually supports the interpretation being made from a measure as opposed to simply whether the measure reliably produces a number and calling it just that (Cronbach & Meehl, 1955).

How was engagement measured?

How reliable is that measurement instrument?

Does the measure behave as expected?

Is it distinct from adjacent concepts?

Does a change in the score represent a meaningful change in the underlying phenomenon?

These questions sound technical because some of them are technical.

But their consequences are not, because they determine whether the number on the dashboard deserves to be there in the first place.

The Executive Does Not Need Your Factor Analysis

CEOs and senior leaders are not particularly interested in learning about your factor analysis or multiple regression model.

Nor should they necessarily have to be.

Executives need interpretable information that helps them make decisions. The burden is on researchers, analysts, and data scientists to make the output simple without making the reasoning overly simplified.

Dots on a line may be all an executive ultimately needs to see.

However, somebody needs to know why those dots are trustworthy.

The rigor should exist underneath the simplicity. Scientific trust and integrity should be embedded in the data.

Organizations nevertheless remain remarkably quick to move from observation to explanation and then from explanation to cause without much thoughtfulness towards the processes and rigor in method leading to such conclusion.

No one needs to hear again that correlation does not equal causation.

Full stop.

So why do organizations with all their scientists and PhDs treat correlations as such?

Measurement Before Analysis
  1. 01Decision
  2. 02Construct
  3. 03Measurement
  4. 04Data
  1. 05Analysis
  2. 06Interpretation
  3. 07Action

Downstream Analytical Tools

DashboardsStatistical ModelsAlgorithmsAI

Analytical tools operate downstream from defining the decision and measuring the construct.

AI Makes the Problem More Important, Not Less

Artificial intelligence is not going to remedy this automatically.

AI may prove to be one of the most consequential technologies of our era, built on a long cumulative history of advances in computation, statistics, communication, and information technology.

We are living through a technological paradigm shift.

However, the sophistication and capability of the tool does not eliminate the need to assess and understand the evidence being fed into it.

If anything, it makes that obligation even more critical than before.

Feed all the data you want into a modern AI system. It can synthesize enormous quantities of information, write and execute statistical code, identify patterns, interpret spreadsheets, generate hypotheses, and perform calculations that would once have required more hours than we can fathom now at this point in time.

That is precisely why the danger is subtle.

Computational sophistication is not epistemic quality.

AI can make weak analyses look remarkably sophisticated.

It can help calculate a confidence interval.

It can help estimate factor loadings.

It can help run a regression model.

It can help summarize thousands of rows of data in seconds.

What those capabilities cannot do by themselves is make a poorly defined construct valid or turn observational data into causal evidence. It cannot repair an inappropriate research design or determine whether the question being asked was the right question in the first place. The methodological judgment still belongs to people and its influence extends only as far as organizations are willing to preserve it within boardrooms of C-suite level execs and other leadership stakeholders.

A mature approach to AI therefore requires evaluating systems and their outputs for qualities such as validity, reliability, transparency, and interpretability rather than assuming that computational sophistication itself establishes absolute and complete trust (National Institute of Standards and Technology, 2023).

The Opportunity

Luckily, there is a much more optimistic way to look at the current scientific and technological phenomena.

AI is an extraordinary analytical resource, and it can augment nearly every stage of research and decision-making in coexistence with sound human judgement and methodological rigor.

It can help researchers write and inspect code.

It can pressure-test hypotheses.

It can identify data-based inconsistencies.

It can accelerate lit reviews.

It can explore alternative explanations and narratives.

It can automate repetitive analytical work through rigorous data pipelining.

It can make sophisticated analytical techniques accessible to far more people than ever before.

That should raise the ceiling for scientific rigor, not lower the floor.

Dependency on AI should not mean abandoning careful thinking, skepticism, measurement, or thorough evaluation of what we believe to be true or accurate.

AI should not be the reason we supplant the scientific method, but should instead be the reason we make fuller use of it.

The goal should not simply be to produce answers faster, but to become better at determining which answers deserve support.

Before the Next Dot Moves

Before building the next dashboard or deploying the next model or asking an AI system what the data mean, organizations should ask three questions:

What decision are we actually trying to improve?

What construct are we truly trying to measure?

What evidence is necessary to justify any conclusion we may have?

Everything else, including the models, dashboards, algorithms, visualizations, and AI, is all secondary.

If organizations want those dotted lines to keep moving in the right direction then they should start by crossing their T’s.

Because the real danger is not that organizations will stop using data.

Quite the opposite.

Organizations will undoubtedly use more of it and even faster through increasingly complex and sophisticated systems and processes.

The danger is that they will confuse the sophistication of the system with the quality of the evidence. They will continue to digest novel and more frequent graphs, slidedecks, and dashboards. They will not be able to help themselves with the new shiny object. And they will believe it to be the truth.

Data science does not need less science as technology improves.

It needs more.

References

  1. Cascio, W. F., & Aguinis, H. (2008). Research in industrial and organizational psychology from 1963 to 2007: Changes, choices, and trends. Journal of Applied Psychology, 93(5), 1062–1081.
  2. Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
  3. de Man, A.-P. (2022). A temporal view on the academic–practitioner gap. Journal of Management Inquiry, 31(2), 200–207.
  4. National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). U.S. Department of Commerce.
  5. Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.

Disclosure

This article is original PrimeStata thought leadership. It is not a representative consulting example and is not commercially sponsored.

Discuss a Decision