AI inference is the process of using a trained machine learning model to make predictions on new, unseen data by applying learned patterns. This course is designed for developers, data scientists, and ML engineers interested in quickly deploying AI inference services on Cloud Run. It is useful for those familiar with cloud-based serverless application deployment solutions, but who may not have experience with running AI inference using Google Cloud serverless products. The course includes examples that deploys a model for AI inference with GPUs and integrates gen AI apps with data storage services.
What you'll learn
deploy AI inference services on Cloud Run
integrate AI applications with data storage services
utilize GPUs for model deployment
Course objectives
understand the process of AI inference
learn about serverless application deployment on Google Cloud