Multimodal Generative AI: Vision, Speech, and Assistants

edX MOOC / Non-credit USD 149
Enroll now →
Multimodal Generative AI: Vision, Speech, and Assistants

About this course

This four-week course provides a hands-on deep dive into the full spectrum of modern AI capabilities. You will master Image-to-Text (Vision), Text-to-Speech (TTS), and Speech-to-Text (Whisper), before culminating in the development of sophisticated AI Assistants. By the end of the course, you’ll be able to build intelligent, multi-modal applications that can see, hear, speak, and solve complex problems.

What you'll learn

  • Image-to-Text conversion
  • Text-to-Speech synthesis
  • Speech-to-Text transcription
  • Development of AI assistants

Skills you'll gain

Related courses

Course details are provided by the platform and may change — always confirm on the provider's site. Links may be affiliate links.