Imagine, Don’t Narrate:
The Generative Bottleneck in
World Models of Interaction Dynamics

1Flybits Labs, Creative AI Hub2Toronto Metropolitan University3MIT Media Lab

NeurIPS 2026 Main Track

On this page
Pen-and-ink cartoon. A person with glasses and a bun tells a robot, “I’m still not convinced.” The robot’s thought bubble holds no dialogue, only a simple sketch of the two of them with the person smiling, and the hand-lettered states “Trust”, with a line rising from a dot, and “Resistance” and “Stress”, each with a line falling from a dot.
Interaction world models. The person speaks in words. The agent plans its reply with its internal world model, imagining how possible future actions would shift the person’s hidden state, such as their trust, resistance, and stress. Social-JEPA does this in a learned latent space, without generating the future as text.

Summary

Modern AI agents can plan, reflect, reason, and act over long-horizon digital tasks. Even so, they cannot plan to steer a social interaction in real time. Doing so requires anticipating how latent factors like trust and resistance will evolve under the agent’s actions, fast enough to plan within a conversational turn. Generative world models approach this by narrating possible futures as text, but autoregressive generation is both too slow for real-time planning and fundamentally lossy.

Social-JEPA imagines instead. It is a joint-embedding predictive architecture for dialogue: a frozen 125M-parameter encoder reads the conversation, and just 500K trainable parameters roll the interaction forward in a learned latent space, conditioned on the agent’s candidate actions, without ever decoding text. To compare the two paradigms against known dynamics, we build LID-Bench, the first controlled testbed for interaction dynamics in which the true hidden states are known.

Three generative world models spanning 117M–350M parameters all learn interaction dynamics internally, yet by five turns ahead (k=5) they lose 89–100% of that signal when they render it as text. Social-JEPA bypasses text entirely, forecasting latent dynamics 2.8–3.9× more accurately while performing rollouts 1,059–2,314× faster. At ~11 ms per turn, real-time planning within conversational turn-taking latencies becomes feasible.

Left: multi-step probe R-squared over rollout horizons 1 to 5. Social-JEPA starts at 0.128 and stays positive at 0.053; three generative world models start near 0.03 to 0.045 and decay to about zero, and a dashed line for the Finetuned-Rep ablation falls from 0.121 to just below zero. Right: for each generative model, hatched bars show the signal learned internally and solid bars the much smaller signal that survives text rendering.
Fig. 1: Latent rollouts stay predictive; text rollouts do not. (a) How well each model’s forecast k turns ahead matches the true hidden state, measured as probe R² (higher is better; LID-Bench ID-test, 5-seed mean ± std). Social-JEPA stays positive through k=5; all three generative world models (GenWM) decay to near zero, with GPT-2 turning negative. Finetuned-Rep, an ablation of Social-JEPA without its EMA target and multi-step training, also falls below zero by k=5. (b) Hatched bars (int.) show what each generative model learns internally; solid bars (text) show what survives text rendering, with the percentage of signal lost. Social-JEPA exceeds even the generative models’ internal representations.

Narrate or imagine?

When AI agents interact over multiple turns—tutoring a struggling student, negotiating a contract, de-escalating a crisis—they must reason about the partially observable internal states of the people they talk to. Trust, resistance, and cognitive load do not evolve monotonically: an intervention that reduces confusion under low resistance may trigger a defensive “backfire effect” when resistance is high. Planning under such state-dependent dynamics requires a world model of the interlocutor.

How should that model represent the future? One approach is to narrate: generate the interlocutor’s next response in text, then re-encode it to estimate how their state has changed. The alternative is to imagine: predict future interaction states directly in a learned representation space, bypassing text entirely. In continuous visual control, latent world models already dominate observation-space approaches in both accuracy and compute [1][2]. Whether that advantage transfers to dialogue, where observations are natural language full of surface cues that tempt shortcut solutions, was an open question.

Social-JEPA

Social-JEPA learns latent interaction dynamics strictly in representation space, following the joint-embedding predictive architecture (JEPA) recipe [3]. It trains fully self-supervised on observed conversations, the agent’s actions, and the time gaps between turns. It never sees the true hidden states.

A frozen RoBERTa-base encoder reads the first and most recent tokens of the dialogue history, preserving both stable traits and the current state, and a learned projection maps it to a 64-dimensional state. A 3-layer MLP predictor then rolls that state forward one turn at a time, conditioned on the agent’s abstract strategy (for example VALIDATE or REFRAME) and on the time gap before the next reply. During training, each of the K=5 predicted states is matched to a target: the conversation’s real future, read by the same frozen encoder and an exponential-moving-average (EMA) copy of the projection. VICReg variance and covariance terms [4] keep the embedding from collapsing to a constant.

Every component matters. Training with single-step instead of multi-step rollout drops R² at k=5 from +0.053 to +0.009. Removing the EMA target as well nearly doubles the spread across seeds, turns the k=5 mean negative (−0.011), and halves the effective rank of the embedding (23.1 → 11.4). Without VICReg, training diverges within the first epoch. And a 2.8× larger generic encoder without dynamics-aware training (SimCSE-RoBERTa-Large, 355M) reaches only +0.106 at k=1, versus Social-JEPA’s +0.128.

LID-Bench

Just as MuJoCo [5] and DMControl [6] were prerequisites for progress in model-based RL, evaluating interaction world models requires a testbed with known dynamics. LID-Bench (Latent Interaction Dynamics) simulates an interlocutor with nine latent coordinates: three stable traits (openness, patience, expertise) and six dynamic ones (trust, resistance, arousal, cognitive load, confusion, stress). These are generator-defined control variables, not psychological claims. They evolve under the agent’s actions through a nonlinear, state-dependent transition with a backfire gate: in regions of high arousal and resistance, interventions reverse sign, so actions that normally build trust erode it.

GPT-4o writes each reply from descriptions of how the person responds (the scope, stance and focus of the reply), and each agent turn from a paraphrase of the agent’s intended move. It never sees the state values or the action names. Four leakage controls stop simple word-matching from reading the state off the text: a TF-IDF baseline reaches only R² = 0.085 out of distribution.

The benchmark contains 4,750 episodes of 25 turns (~119k transitions) across five domains: tutoring, medical triage, customer service, career mentoring, and persuasion. It includes five distribution shifts, each changing a single mechanism, and three tasks: multi-step state forecasting, counterfactual action effects, and planning in latent space. Models observe only the text, the agent’s actions and the time gaps; the true hidden states are used only for evaluation. LID-Bench will be released publicly.

Fig. 2: What the interlocutor says versus what is happening inside. A complete 25-turn episode from the LID-Bench ID-test split (MedTriage). Left: the conversation so far, with the agent’s action and the time gap at each turn. The model sees nothing else, and once the conversation outgrows its 512-token window it reads only the start and the latest turns. Right: the true hidden state after the reply (the oracle), used only for evaluation. From t = 1 the parent is inside the backfire region, and by t = 2 trust has fallen to zero, although their replies stay polite. At t = 12, after a long pause, the state leaves the region, and trust then slowly recovers. The shaded region marks where the backfire gate exceeds 0.5 under the default gate parameters.

Experimental results

We compare three world-modelling paradigms on LID-Bench. Generative world models (GenWM: GPT-2 117M, OPT-125M, OPT-350M) are fine-tuned to predict the interlocutor’s next reply given the history, action, and time gap. Reconstruction (ReconWM) uses the same frozen encoder and predictor as Social-JEPA but is trained to reconstruct the text. Finetuned-Rep, an ablation, keeps Social-JEPA’s architecture but drops the EMA target and multi-step rollout. Every method is scored the same way: a linear probe, trained on the training split, maps its predicted state to the true hidden state. For generative models, each generated reply is first re-encoded by a frozen SimCSE-RoBERTa-Large encoder.

Forecasting interaction dynamics

The benchmark behaves as intended: a reference model given the true hidden states and the full dynamics reaches R² > 0.94 at every horizon, while shortcut baselines fail at k=1 (TF-IDF +0.037; zero-shot Claude 3.5 Sonnet −0.132). Social-JEPA outperforms GenWM GPT-2 by 3.9× at k=1 (+0.128 vs +0.033, p < 10−7) and stays positive through k=5, while the generative models decay to near zero (Fig. 1a). OPT-125M and OPT-350M both beat GPT-2 at every horizon, yet Social-JEPA still leads by 2.8×: better pretraining and ~3× more parameters do not close the gap.

Where the signal goes

Probing each generative model’s hidden states directly, without decoding text, separates what it learns from what it can express (Fig. 1b). GPT-2 loses 56% of its learned dynamics signal through text rendering at k=1. OPT-125M loses 20% at k=1, growing to 89% at k=5. OPT-350M learns more internally at k=1, but its loss grows from 32% to 91% by k=5: scaling improves single-step encoding, not what survives text. Even read out internally, every generative model falls well below Social-JEPA.

Reconstruction fails differently. ReconWM achieves the best encoding R² of any method (0.196) yet catastrophic dynamics (−0.607 at k=1). Because it shares Social-JEPA’s encoder and predictor, this isolates the cause: the reconstruction objective forces the latent space to encode surface features needed for text, not dynamics. Generative models lose the signal at inference; reconstruction loses it at training. Observation-space coupling is the common failure mode.

Counterfactual reasoning shows the same pattern. Asked how each possible action would change the state, Social-JEPA and the generative models predict the direction of the effect about equally well (sign accuracy 0.607–0.611, against 0.524 for a baseline that ignores the state). But the generative models cannot rank actions by the size of their effect (Spearman ρ at most +0.027, versus +0.330 for Social-JEPA). Coarse directions survive text rendering; fine-grained magnitudes do not.

Robustness under distribution shift

Each of LID-Bench’s five shifts changes exactly one mechanism: delayed action effects, a lower backfire threshold, rare starting states with high arousal and resistance, a different family of transition dynamics, or noisy text (dropped words and inserted filler). At k=5, Social-JEPA remains positive on the in-distribution test set and four of five shifts, while the generative models and Finetuned-Rep are often near zero or negative, and scaling from 125M to 350M does not improve robustness. Noisy text is the exception for Social-JEPA (−0.001 at k=5); diagnostics point to the frozen encoder rather than the dynamics model.

Heatmaps of multi-step R-squared for Social-JEPA, Finetuned-Rep, and three generative models on the in-distribution test set and five distribution shifts, at horizon 1 and horizon 5. At horizon 5, Social-JEPA's column is positive on every row except Noise, while other methods are mostly near zero or negative.
Fig. 3: Robustness under five controlled distribution shifts. Multi-step probe R² at k=1 and k=5 for the trained models (5-seed mean, shared colour scale). At k=5, Social-JEPA remains positive on ID-test and 4/5 shifts; generative models and Finetuned-Rep (FT-Rep) are often near zero or negative. ReconWM is omitted (R² < −0.5).

Planning in real time

The practical payoff of latent dynamics is planning. We run the cross-entropy method (CEM: 256 candidate action sequences, 5 iterations, horizon 5) over each latent-space model and report normalised improvement (NI), where random actions score 0 and planning with the true hidden state scores 1. Social-JEPA plans well under a cooperative reward (NI = 0.640) and an adversarial one (0.553) with the same frozen dynamics model; only the reward changes. Under the adversarial reward, the policy that generated the training data scores −0.104, worse than random, so the planner is not copying behaviour it saw in training: it exploits the learned dynamics.

Table 1: Planning quality. Normalised improvement under CEM (5-seed mean ± std). The two latent-space planners are statistically indistinguishable under the cooperative reward (p = 0.92). A generative world model takes ~628–972 s per CEM step, so planning with it is infeasible.
MethodCooperativeAdversarial
Social-JEPA (ours)+0.640 ± 0.17+0.553 ± 0.25
Finetuned-Rep+0.650 ± 0.17+0.632 ± 0.22
ReconWM−0.195 ± 0.53+0.044 ± 0.16
Behaviour policy+0.309 ± 0.06−0.104 ± 0.04
GenWM (all variants)infeasible

Speed is what makes this usable. A Social-JEPA rollout takes 0.43 ms and a full CEM step 5.7 ms; the same step with a generative world model takes ~628–972 s. Social-JEPA’s complete per-turn cost, encoder plus CEM step, is ~11 ms, under 1.1% of a 1 s conversational turn-taking window [7]. It is the only paradigm where sampling-based planning is compatible with real-time interaction.

Left: normalised planning improvement for Social-JEPA stays between about 0.70 and 0.83 across planning horizons 5 to 20, above Finetuned-Rep at about 0.41 to 0.49. Right: latency on a log scale; Social-JEPA's rollout takes 0.4 ms and its CEM step 5.7 ms, while generative models take 461 to 1,016 ms per rollout and about 628 to 972 seconds per CEM step, far above a 1 second line.
Fig. 4: Fast enough to plan within a turn. (a) Planning quality across planning horizons H=5–20 for a single seed (seed 42, cooperative reward). Averaged over five seeds, the two methods are statistically indistinguishable (Table 1). (b) Computational cost on a log scale (RTX 4090, batch size 1). Social-JEPA’s rollout and CEM step fall well below the ~1 s turn-taking threshold; every generative variant exceeds it by ~3 orders of magnitude on CEM planning.

Illustrative scenario: a crisis-line call (adapted from Appendix R.3)

Suppose a caller is getting more agitated by turn 8. In about 11 ms, the agent reads the conversation so far, estimates that the caller’s arousal and resistance are high, close to the backfire region, and uses Social-JEPA to try out plans for the next five turns: 256 at a time, refined over five rounds.

Close to the backfire region, the model predicts that validating the caller, usually a trust-building move, would push arousal higher, while waiting and then summarising lets it settle first. Once the caller calms down, the agent can pursue a different goal, such as building rapport, by changing only the reward. The model itself is not retrained.

Limitations and real dialogue

LID-Bench’s latent states are generator-defined control variables, not real-world ground truth. On KODIS [8], a corpus of real multicultural dispute-resolution dialogues, the gap between what generative models learn internally and what survives in their text replicates in 12/12 runs across all three model families. It holds for three of the four traits that map onto LID-Bench’s coordinates; Honor (mapped to resistance) is the exception. Effect sizes are small, and multi-step rollout and planning are evaluated on LID-Bench only. Absolute R² values are modest (0.128 at k=1, 0.053 at k=5), reflecting how hard it is to infer state from text alone. Planning, however, only needs predicted states to rank action sequences correctly, and R² = 0.053 at k=5 is enough for NI = 0.64.

The same reward-agnostic dynamics that let one model serve cooperative or adversarial objectives could also serve manipulation under a different reward. Because the probe translates latent predictions into the quantities a reward operates on, it is the natural point of governance: deployments should pair this work with consent from those whose states are inferred, audit logs of probe outputs, and reward functions vetted against externally validated objectives.

References

[1] Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to Control: Learning Behaviors by Latent Imagination. ICLR, 2020.

[2] Julian Schrittwieser et al. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model. Nature 588, 2020.

[3] Yann LeCun. A Path Towards Autonomous Machine Intelligence. Online essay, 2022.

[4] Adrien Bardes, Jean Ponce, and Yann LeCun. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning. ICLR, 2022.

[5] Emanuel Todorov, Tom Erez, and Yuval Tassa. MuJoCo: A Physics Engine for Model-Based Control. IROS, 2012.

[6] Yuval Tassa et al. DeepMind Control Suite. arXiv:1801.00690, 2018.

[7] Tanya Stivers et al. Universals and Cultural Variation in Turn-Taking in Conversation. PNAS 106(26), 2009.

[8] James Anthony Hale, Sushrita Rakshit, Kushal Chawla, Jeanne M. Brett, and Jonathan Gratch. KODIS: A Multicultural Dispute Resolution Dialogue Corpus. NAACL, 2025.

BibTeX

@misc{platnick2026imagine,
  author    = {Platnick, Daniel and Alirezaie, Marjan
               and Rahnama, Hossein},
  title     = {Imagine, Don't Narrate: The Generative Bottleneck
               in World Models of Interaction Dynamics},
  month     = aug,
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.21733425},
  url       = {https://doi.org/10.5281/zenodo.21733425},
  note      = {Accepted to NeurIPS 2026 Main Track}
}