Unify Modalities: Cross-Modal Retrieval

Coursera MOOC / Non-credit USD 49
Enroll now →
Unify Modalities: Cross-Modal Retrieval

About this course

Transform how AI systems understand and connect different data modalities. This course empowers machine learning professionals to build cutting-edge cross-modal retrieval systems that bridge the gap between text and images. You'll master the technical implementation of approximate nearest-neighbor search algorithms and design sophisticated attention mechanisms that fuse visual and textual information. Through hands-on work with production-scale tools like FAISS and real datasets like Flickr30K, you'll develop the expertise to create intelligent systems that understand content across modalities—enabling breakthrough applications in search, recommendation, and content understanding that mirror how humans naturally process diverse information types.

What you'll learn

  • understand cross-modal retrieval systems
  • implement approximate nearest-neighbor search algorithms
  • design attention mechanisms for unifying text and images
  • apply tools like FAISS in practical scenarios
  • work with production-scale datasets like Flickr30K

Course objectives

  • master the technical aspects of cross-modal retrieval
  • bridge gaps between different data modalities
  • enable applications in search and recommendation systems

Skills you'll gain

Related courses

Course details are provided by the platform and may change — always confirm on the provider's site. Links may be affiliate links.