Definition

Multimodal AI refers to models that can work with multiple types of data, or modalities.

Text is only one form of information. Multimodal systems can also work with images, video, audio, and other forms of structured or unstructured information. For example, a model may analyze an image, identify objects, or modify an image based on an instruction.

Multimodal AI expands the applications of generative AI beyond text and creates opportunities across areas such as marketing, content creation, customer support, education, and media.