Key points are not available for this paper at this time.
In online retail, pricing policies are often implemented when the demand for a product is unknown. This stream of work is known as dynamic pricing with demand learning. Such dynamic pricing policies are usually based on resolving exploitation-exploration trade-offs prevalent in sequential decision-making under uncertainty. Should a price be revised based on the realised revenues in prior history, or should it be revised based on the promise of a better performance in the future? We examine whether prices should be selected solely based on apparent revenue consideration or should they also account for the behavioural underpinnings of the purchase intention. This paper proposes a data-driven dynamic pricing policy based on reinforcement learning that maximises revenues while incorporating the cumulative (un)fairness concerns based on customers’ prior experience with previous price revisions. Cumulative (un)fairness accumulates (un)fairness for all prior price revisions and is termed as ‘sticky’ fairness concerns. We find that the proposed policy favours implementing smaller prices compared to higher prices in realising higher revenues. We test this counterintuitive finding by conducting numerical and lab experiments in various demand settings. The proposed dynamic pricing policy demonstrates logarithmic regret in time.
Rathore et al. (Sat,) studied this question.