Multimodal Generative AI: Vision, Speech, and Assistants

Coursera MOOC / Non-credit USD 49
Enroll now →
Multimodal Generative AI: Vision, Speech, and Assistants

About this course

We are introducing a new course to replace the "Coding with ChatGPT" course in the Generative AI specialization. This updated course will cover materials, models, and content released in 2024. Some of the new additions include material on using AI for image-to-text (vision), text-to-speech, speech-to-text, and the Assistant API. All these topics come with new labs, lessons, and exercises.

What you'll learn

  • understand the principles of multimodal generative AI
  • apply techniques for image-to-text and text-to-speech
  • utilize speech-to-text capabilities
  • work with the Assistant API effectively

Skills you'll gain

Related courses

Course details are provided by the platform and may change — always confirm on the provider's site. Links may be affiliate links.