Social-JEPA · Animation study
Narrate or imagine?
Two ways a world model can simulate how a conversation will unfold over five turns. Follow the sequence, or select a step to inspect it.
Both world models start from the same observable history: the interlocutor’s replies and the agent’s own turns so far. Social-JEPA encodes it once, with a frozen encoder and a learned projection, into a 64-dimensional latent state.
To predict one turn ahead, the generative model must write the interlocutor’s next reply token by token, given the agent’s next action and the time gap. Social-JEPA applies its predictor once to the latent state, conditioned on the same action and time gap.
A five-turn rollout repeats that loop. The generative model writes and re-reads a full reply at every step; Social-JEPA chains five predictor calls in latent space. Rollouts are 1,059–2,314× faster.
To read out the interlocutor’s state, the generated text must be re-encoded and then probed. Going through text loses 89–100% of the signal the generative models learn internally by k=5. Social-JEPA probes its predicted latent directly.
Probe R² measures how well the predicted state recovers the oracle latent state on LID-Bench ID-test; the generative value is the best of GPT-2, OPT-125M and OPT-350M.