Skip to main navigation Skip to search Skip to main content

A data-driven approach to demand forecasting and fulfillment in fast fashion retailing

  • Chongshou LI

    Student thesis: Doctoral Thesis

    Abstract

    Fast fashion retailing has been playing an important role in the apparel industry in recent years and makes a significant contribution to the market size growth for this industry. The main characteristics of this new business are quick response and changes in assortment and affordable prices with high fashion styles. Due to them, there are new problems appearing in the operations of this retailing. This doctoral dissertation studies three of them and provides data-driven approaches with a real data set from a multinational fast fashion retailer in Singapore, where it operates over twenty stores and a warehouse. This real data set includes daily sales transactions records, monthly inventory information over two years and attributes information associated with each item such as color, size, the heel height range (for some shoes only) and the name of the designer etc. Many fast fashion products are just sold in one season (also known as “one shot”) and the initial demand forecasting is significant for the revenue. The first problem that this thesis solves is how to forecast demand for products to be launched. These products are not carried at any stores and no historical sales data exist. The present thesis proposes a new demand model to it based on the factorization machine which was recently proposed as a generic predictor and applied in recommendation systems. This model takes the effects of both single and pairwise attributes of items into account. Three loss functions with different purposes and two algorithms of alternative least squares and stochastic gradient descent are proposed to estimate the parameters in the model; one loss function implements an important idea in demand forecasting that the overestimation should be given less penalty than that of the underestimation under the same absolute gap between the forecast and the ground-truth; the stochastic gradient descent is employed for minimizing this loss function and to refine forecasts such that the overestimation ratio is increased without largely decreasing prediction accuracy; how to choose useful attributes is also addressed. The forecasting error in terms of mean absolute percentage error (MAPE) over the real test set including 1,047 stock keeping units (SKUs) is 13.97% at the aggregate chain level and comparable with the state-of-art. Besides the new products, this dissertation addresses the issue of demand forecasting for existing products. In order to respond the market quickly, the present thesis provides a method to demand forecasting for each item on each store on each day (at the aggregation level of SKU-store-day). The challenge of this problem comes from that demand series at this aggregation level is highly intermittent. Namely, each demand series is composed by non-zero demand sizes and a certain number of zero demand sizes. The prediction involves not only the demand size but also the demand interval between two adjacent non-zero demand sizes. In order to tackle this challenge, this thesis proposes a decomposition-aggregation method; the cross-sectional aggregation over all SKUs and stores on each day is first performed and the series of SKU-store-day demand is consolidated to a new series noted as daily demand index; the exponential smoothing with seasonality of weekly variation and that by public holidays is utilized to daily demand index forecasting; widely used Syntetos-Boylan approximation (SBA) forecasts including demand sizes and intervals are computed; then a greedy heuristic for the aggregation-decomposition is proposed to produce the targeted forecasts at the aggregation level of SKU-store-day, which combines both daily demand index forecasts and SBA forecasts. The computational results show that the all forecasting errors in terms three measures of two widely used and a new proposed are less than that of the off-the-shelf methods of SBA and Croston in many software packages. Moreover, this study proposes the new measure based on the MAPE for intermittent demand forecasting which is completely free of the problem of “division by zero”. Remarkably, in the process of demand index forecasting, influences of public holidays in both Singapore and China on fast fashion sales in Singapore are quantified and reported. After demand forecasts are generated, the problem of how to effectively realize them appears. The realization here involves delivering and picking up items among the warehouse and stores in this retail network in Singapore. This problem is modeled as the vehicle routing problem with simultaneous pickup and delivery and transshipment (VRPSPDT), a new and interesting variant of vehicle routing problem. The distinguished characteristics of this problem is that collected items from stores (vertices) can be used by other stores (vertices) demanding them; it is multi-item; the number of items shipped can be as large as 11,745 which increases the difficulties of this problem significantly. This model first implements the idea of inventory transhipment policy into the routing literature and should be widely adopted in the reverse logistics. An algorithm based on the adaptive memory programming is proposed to it. Computational results show that the proposed algorithm can effectively solve the VRPSPDT. And 66 benchmark instances of the VRPSPDT from the real world are constructed for the community. The last but not least, this thesis supplies well-organized data sets for the community and proposes two techniques for, respectively, data preprocessing and anonymization for sales data analytics for retailing. Data preprocessing is a necessary step in any data-driven work. In this study, a distinguished task is to compute the daily stock level for each item at each store during the time period examined, which is critical for demand estimation. This retailer does not update the stock level every day for every item at every store. And original data set includes only monthly inventory information; stockouts on each day should be identified and be taken into account when training demand models. This thesis presents a method to it and calculates the daily inventory level for each item at each store. Regarding the data anonymization, it is trivial and easy to disguise any categorical information such as department names, product number, size, color etc. Besides it, an important concern from the retailer for sharing the data set to the community is that the revenue is very sensitive information and must not be accessed by the public or others. At the same time, quantities of sales are preferred to remain for keeping the nature of the data. So, this study proposes to linearly transform prices such that the revenue is not able to be known; the masked price is also reasonable and reserves important properties of the original information such as being positive and maintaining vertical differentiation among items in the same department (category). After that, the data could be contributed to the community for motivating new research problems and teaching innovations without violating the business confidentiality and losing key characteristics.
    Date of Award2 Oct 2015
    Original languageEnglish
    Awarding Institution
    • City University of Hong Kong
    SupervisorChi Hang Stephen LEUNG (Supervisor) & Leong Chye Andrew LIM (Supervisor)

    Keywords

    • Retail trade
    • Fashion merchandising
    • Forecasting

    Cite this

    '