Models that combine vision with other modalities — such as text, audio, or video — to reason across inputs.