Multimodal AI refers to AI systems that can process and generate more than one type of data — such as text, images, audio, or video — within a single model.

Early AI models were typically unimodal — a text model processed text, an image model processed images. Multimodal models can accept and produce multiple types of input and output, enabling tasks like describing an image in text, generating an image from a text description, or transcribing and summarizing audio.

Related Terms

Learn More

← Back to Glossary AI Glossary → Français →