C4DM Seminar: Magical Embeddings: King-Queen to Guitar-Distortion
QMUL, School of Electronic Engineering and Computer Science
Centre for Digital Music Seminar Series
Seminar by: Paula Lauren
Date/time: Wednesday, 7th October 2026, 14:00 - 15:00
Location: GC601, Graduate Center, Mile End Campus, Queen Mary University of London
Teams link: https://teams.microsoft.com/meet/358341781556280?p=vOiCbw5tH4IeNwTMaq
Title: Magical Embeddings: King-Queen to Guitar-Distortion
Abstract: Word embeddings revealed that semantic relationships live as linear directions in a vector space: the arrow from "man" to "king" is the same as from "woman" to "queen." Contrastive learning carried this structure into vision with Contrastive Language-Image Pretraining (CLIP) and audio with Contrastive Language-Audio Pretraining (CLAP), raising the question: do the directions survive? This talk traces that arc from Word2Vec, Global Vectors (GloVe), and Extreme Learning Machine (ELM) based embeddings through CLIP to CLAP, asking whether the direction that distortion pushes a guitar can retrieve a distorted keyboard. Current work by the presenter shows it can, far above chance, on two independent encoders. Different training, different modalities, yet the same arithmetic emerges. The talk also provides a provenance of the presenter's journey from text-based natural language processing into digital music research.
Bio: Paula Lauren is an associate professor in the Department of Mathematics and Computer Science at Lawrence Technological University. She teaches undergraduate and graduate courses in artificial intelligence, machine learning and pattern recognition, natural language processing, algorithm design and analysis, database systems, and foundational computer science courses. She earned her Ph.D. in Computer Science from Oakland University in 2018. Prior to her doctoral studies, Paula spent over 10 years working full-time in industry, with roles ranging from computer programmer to software engineering manager. Her current research approaches the interpretability of learned audio embeddings through the geometric structure that audio effects reveal.
