Social-JEPA · Animation study
Learning to imagine without text
How Social-JEPA learns interaction dynamics in latent space, fully self-supervised. Follow the sequence, or select a step to inspect it.
A frozen RoBERTa-base encoder (125M parameters) reads the dialogue history Ht, keeping its first and most recent tokens so that stable traits and the current state both fit in context. A trainable projection πφ maps it to a 64-dimensional state st.
The predictor Pθ, a 3-layer MLP, rolls the state forward one turn at a time, conditioned on embeddings of the action the agent actually took, A, and the time gap Δt before the next reply. No text is generated at any step.
Training targets come from the conversation’s real future. The same frozen encoder reads each future history Ht+k, and πψ, an exponential moving average of πφ, maps it to a target state. Targets receive no gradient.
Each prediction is pulled toward its target, discounted by γk over K = 5 steps, while VICReg variance and covariance terms keep the embedding from collapsing. Gradients update only πφ and Pθ (about 500K parameters); πψ then tracks πφ by moving average.
Training uses only observed text, actions, and time gaps; LID-Bench’s oracle latent states are never seen.