MMDP: A Mobile-IoT Based Multi-Modal Reinforcement Learning Service Framework

Puming Wang,Laurence T. Yang,Jintao Li,Xue Li,Xiaokang Zhou
DOI: https://doi.org/10.1109/tsc.2020.2964663
IF: 11.019
2020-07-01
IEEE Transactions on Services Computing
Abstract:With the development of GPS technology, a new Mobile Internet of Things (M-IoT) is emerging, which perceives the city's rhythm and pulse day and night to collect a large scale of city data. It is urgent to innovate M-IoT service system for these large-scale and heterogeneous data. To cope with the problem, this article proposes a Mobile-IoT based multi-modal reinforcement learning service framework from data perspective, which has three highlights, i) Developing Action-aware High-order Transition Tensor ($AHTT$<math>AHTT</math>) to fuse the heterogeneous data from M-IoTs in a unified form. ii) Developing Multi-modal Markov Decision Process ($MMDP$<math>MMDP</math>) to model the multi-modal reinforcement learning for M-IoT service framework. iii) Developing Tensor Policy Iteration algorithm ($TPIA$<math>TPIA</math>) to solve the optimal tensor policy. Due to using tensor keeps the multi-modal relations of the context information in the process of solving the optimal policy. The proposed M-IoT service system provides more personalized service for taxi drivers. The experiment results shows that most taxi drivers earn more revenue according to the tensor policy.
computer science, information systems, software engineering
What problem does this paper attempt to address?