Photo By: XR Expo
Artificial intelligence is increasingly being used to simulate how customers might respond to changes in products, pricing, advertising and other interventions. But while these simulations can offer valuable insights, researchers have found that basic AI simulations may sometimes overestimate the impact of those interventions.
The reason, according to Shenbo Xu, co-founder of Kapnova, is that AI customers can behave very differently from real ones.
“Basic AI simulations tend to overestimate intervention effects because AI is more decisive and deliberate than real customers,” Xu said.
Human shoppers rarely make decisions in the controlled, highly focused way that an AI simulation does. Real customers may be distracted, indifferent or inconsistent. Their decisions can also be shaped by circumstances that are difficult to capture in a customer profile: they may be in a hurry, briefly notice an advertisement, lose interest or simply fail to pay attention to a change that a simulation considers important.
An AI system, by contrast, is typically given a defined customer profile and explicitly asked to make a decision. That setup can encourage the model to focus closely on the differences between two versions of a product, promotion or intervention.
“An AI, by contrast, is given a customer profile and asked to make a decision, so it tends to focus heavily on the differences between two versions and respond more strongly than a real person would,” Xu said.
Researchers have described this phenomenon as “behavioral amplification”—a situation in which relatively small differences between interventions can produce disproportionately large predicted changes in behavior.
The missing context behind customer decisions
One of the more fundamental problems is that customer profiles rarely capture everything that influences a real-world decision.
An AI model might know that a customer has previously purchased a particular type of product, has certain preferences or belongs to a particular demographic group. But that information does not necessarily reveal what is happening at the moment the decision is made.
“The deeper issue is that AI often knows a customer’s profile—such as purchase history, preferences, or demographics—but not their immediate context, which can strongly influence real-world behavior,” Xu said.
That immediate context can be difficult to model. A shopper might respond differently to the same promotion depending on whether they are browsing casually, rushing to complete a purchase or distracted by something else. Even a customer with a strong historical preference for a product may ignore an intervention when circumstances change.
This creates a fundamental distinction between what a customer is likely to say or explain and what that customer is actually likely to do.
When language and behavior diverge
Large language models introduce another complication. These systems are trained extensively on language, and people tend to express their preferences more clearly in words than they act on them in everyday life.
That means an AI can sometimes construct a convincing explanation for why a particular promotion, product or message should appeal to someone without accurately predicting whether that person would actually click, purchase or respond.
“There is also a mismatch between how AI is trained and what it is being asked to predict,” Xu said. “Large language models learn primarily from text, where people express opinions and preferences much more clearly than they act on them.”
The distinction matters for businesses using AI to test potential interventions before deploying them to customers. A simulation may indicate that changing a message or offer will have a substantial effect, while the eventual real-world response could be considerably smaller.
Historical data can help—but only to a point
One way to address the problem is through calibration against historical behavioral data. By comparing AI-generated predictions with what customers have actually done in the past, researchers and businesses can adjust simulations to better reflect observed behavior.
Xu said this approach can substantially reduce the gap between simulated and real-world outcomes, but cautioned that calibration should not be viewed as a complete solution.
“Historical-data calibration can substantially reduce this error, but it is ultimately a correction rather than a complete solution,” he said.
The broader challenge, he argues, is to make AI simulations more closely connected to observed human behavior rather than relying primarily on the way people describe their preferences.
From simulated customers to observed behavior
As businesses increasingly turn to AI for market research and experimentation, the distinction between simulated behavior and actual behavior is likely to become more important.
AI simulations can provide a fast and relatively inexpensive way to explore potential customer responses. But their predictions need to be interpreted with an understanding of how artificial decision-making differs from the messy, contextual nature of human behavior.
“Closing the gap will likely require grounding AI more directly in observed human behavior, not just the language people use to describe their preferences,” Xu said.
For companies using AI to forecast the effects of an intervention, the lesson is not necessarily to abandon simulations. Rather, it is to recognize what the simulations are measuring—and what they may be missing.
An AI can reason about why a customer should respond to a change. Predicting whether that customer will actually respond, however, may require something more difficult: understanding the context in which a real human makes the decision.