E-commerce warehousing: learning a storage policy

Adrien Rimélé,Philippe Grangier,Michel Gamache,Michel Gendreau,Louis-Martin Rousseau
DOI: https://doi.org/10.48550/arXiv.2101.08828
2021-01-22
Abstract:E-commerce with major online retailers is changing the way people consume. The goal of increasing delivery speed while remaining cost-effective poses significant new challenges for supply chains as they race to satisfy the growing and fast-changing demand. In this paper, we consider a warehouse with a Robotic Mobile Fulfillment System (RMFS), in which a fleet of robots stores and retrieves shelves of items and brings them to human pickers. To adapt to changing demand, uncertainty, and differentiated service (e.g., prime vs. regular), one can dynamically modify the storage allocation of a shelf. The objective is to define a dynamic storage policy to minimise the average cycle time used by the robots to fulfil requests. We propose formulating this system as a Partially Observable Markov Decision Process, and using a Deep Q-learning agent from Reinforcement Learning, to learn an efficient real-time storage policy that leverages repeated experiences and insightful forecasts using simulations. Additionally, we develop a rollout strategy to enhance our method by leveraging more information available at a given time step. Using simulations to compare our method to traditional storage rules used in the industry showed preliminary results up to 14\% better in terms of travelling times.
Optimization and Control,Robotics
What problem does this paper attempt to address?
The problem that this paper attempts to solve is how to minimize the average cycle time for robots to complete requests through dynamic storage strategies in e - commerce warehousing. Specifically, the paper focuses on in warehouses using the Robotic Mobile Fulfillment System (RMFS), when robots are available at the operator station, it is necessary to decide in real - time which location to store the shelves in order to reduce the average travel time of robots and thus improve the overall throughput of the warehouse. The paper proposes an approach based on the Partially Observable Markov Decision Process (POMDP) and Reinforcement Learning (RL), especially using the Deep Q - learning agent to learn effective real - time storage strategies, taking advantage of repeated experiences and insights from simulations. In addition, an extended strategy has been developed to enhance the performance of the method by leveraging more available information in a given time step. Through comparison with traditional storage rules, the preliminary results show that this method can improve the travel time by up to 14%.