Designed and built with care, filled with creative elements

Top

Guy Gilboa – How to Encode World Knowledge?

Salle W

Foundation models are a key platform which implicitly encodes world knowledge. In this talk we first focus on vision-language models, such as CLIP, and investigate their geometric behavior and logic behind the high-dimensional feature encoding. For instance, we find that as image or text become more rare and distinct they are encoded further from the center of the embedding. We explain why InfoNCE loss leads to that behavior. We also find out empirically that each modality can be well modeled statistically as admitting a multivariate Gaussian distribution. This finding is […]