Multimodal Transformers¶
The same ideas of multimodality as used in text to speech, can essentially be extended to train models based on tokens from text, audio and image together. Our output set becomes a concatenation of tokens from all modalities. All types of inputs are treated as part of the same sequence.