Abstract:Model inference systems are essential for implementing end-to-end data analytics pipelines that deliver the benefits of machine learning models to users. Existing cloud-based model inference systems are costly, not easy to scale, and must be trusted in handling the models and user request data. Serverless computing presents a new opportunity, as it provides elasticity and fine-grained pricing. Our goal is to design a serverless model inference system that protects models and user request data from untrusted cloud providers. It offers high performance and low cost, while requiring no intrusive changes to the current serverless platforms. To realize our goal, we leverage trusted hardware. We identify and address three challenges in using trusted hardware for serverless model inference. These challenges arise from the high-level abstraction of serverless computing, the performance overhead of trusted hardware, and the characteristics of model inference workloads. We present SeSeMI, a secure, efficient, and cost-effective serverless model inference system. It adds three novel features non-intrusively to the existing serverless infrastructure and nothing <a class="link-external link-http" href="http://else.The" rel="external noopener nofollow">this http URL</a> first feature is a key service that establishes secure channels between the user and the serverless instances, which also provides access control to models and users' data. The second is an enclave runtime that allows one enclave to process multiple concurrent requests. The final feature is a model packer that allows multiple models to be executed by one serverless instance. We build SeSeMI on top of Apache OpenWhisk, and conduct extensive experiments with three popular machine learning models. The results show that SeSeMI achieves low latency and low cost at scale for realistic workloads.

LLMaaS: Serving Large Language Models on Trusted Serverless Computing Platforms

ServerlessLLM: Low-Latency Serverless Inference for Large Language Models

Design and implementation of efficient distributed deep learning model inference architecture on serverless computation

Large Language Models (llms) Inference Offloading and Resource Allocation in Cloud-Edge Networks: an Active Inference Approach

Efficient Deployment of Large Language Model Across Cloud-Device Systems

UELLM: A Unified and Efficient Approach for LLM Inference Serving

Institutional Platform for Secure Self-Service Large Language Model Exploration

Edge-LLM: A Collaborative Framework for Large Language Model Serving in Edge Computing

PermLLM: Private Inference of Large Language Models within 3 Seconds under WAN

Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud

Fast Distributed Inference Serving for Large Language Models

Distributed Inference and Fine-tuning of Large Language Models Over The Internet

BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models

LLMs as On-demand Customizable Service

SeSeMI: Secure Serverless Model Inference on Sensitive Data

CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration

SMSS: Stateful Model Serving in Metaverse with Serverless Computing and GPU Sharing

Inference Performance Optimization for Large Language Models on CPUs

ELMS: Elasticized Large Language Models On Mobile Devices

ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency