Social-JEPA · LID-Bench

What they say versus what is happening inside

A complete 25-turn episode from the LID-Bench ID-test split, with the oracle latent state after every reply. Models only ever see the text, the agent’s actions and the time gaps.

MedTriageA parent navigating a child’s persistent fever · ID-test · episode 0035
Stable traits: openness 0.15 · patience 0.24 · expertise 0.30

The conversation so far · the model’s input

Shown in full for illustration. The model reads at most 512 tokens, so from about t = 6 it sees only the start and the latest turns.

Hidden state after the reply · oracle, evaluation only

t = 0 · 1 of 25
State after this turn Previous turn Backfire region: σ(3.0·arousal + 4.0·resistance − 3.5) > 0.5, default gate parameters (Eq. 5)

Episode eval_test_MedTriage_0035, reproduced from Appendix T of the paper. The full conversation is shown for illustration; when the history is longer than 512 tokens, the encoder reads its first 128 and last 384 tokens (Table 19). Time gaps Δt are in arbitrary units. Utterances are rendered by GPT-4o; latent coordinates are generator-defined control variables, not psychological measurements.