Social-JEPA · Animation study

Narrate or imagine?

Two ways a world model can simulate how a conversation will unfold over five turns. Follow the sequence, or select a step to inspect it.

Generative world model versus Social-JEPA over a five-turn rollout Top lane: a generative world model, given the agent's action and the time gap, writes the interlocutor's reply as text, appends it to the history and repeats for five turns, then re-encodes the text and probes it for the latent state, reaching R-squared of at most 0.006 at turn five. Bottom lane: Social-JEPA encodes the history once into a latent state, applies its predictor five times conditioned on candidate actions, and probes the final latent directly, reaching R-squared 0.053. NARRATE Generative world model (GenWM) IMAGINE Social-JEPA Ht dialogue so far GenWM GPT-2 · OPT action, time gap append reply to history, generate again generated reply, word by word turns narrated 461–1,016 ms per five-turn rollout Re-encode SimCSE-RoBERTa-L g probe Ẑt+5 R² ≤ +0.006 89–100% of learned signal lost in text Ht dialogue so far Encoder RoBERTa + πφ st candidate actions and time gaps 0.43 ms per five-turn rollout g probe Ẑt+5 R² = +0.053 probed directly, no text in the loop

Both world models start from the same observable history: the interlocutor’s replies and the agent’s own turns so far. Social-JEPA encodes it once, with a frozen encoder and a learned projection, into a 64-dimensional latent state.

To predict one turn ahead, the generative model must write the interlocutor’s next reply token by token, given the agent’s next action and the time gap. Social-JEPA applies its predictor once to the latent state, conditioned on the same action and time gap.

A five-turn rollout repeats that loop. The generative model writes and re-reads a full reply at every step; Social-JEPA chains five predictor calls in latent space. Rollouts are 1,059–2,314× faster.

To read out the interlocutor’s state, the generated text must be re-encoded and then probed. Going through text loses 89–100% of the signal the generative models learn internally by k=5. Social-JEPA probes its predicted latent directly.

Step 1 of 4
Text (observation space) Latent state Replies are illustrative. Timings (RTX 4090, batch size 1) and R² at k=5 are from the paper.

Probe R² measures how well the predicted state recovers the oracle latent state on LID-Bench ID-test; the generative value is the best of GPT-2, OPT-125M and OPT-350M.