SMMU: Benchmarking Social Intelligence of Multimodal Large Language Models

EMNLP (Findings) 2026
1Nanyang Technological University 2Anuttacon
✉Corresponding authors
Nanyang Technological University Anuttacon

TL;DR

Our paper evaluates

  1. Does your model understand complex social interactions over long horizons as the video unfolds?

  2. Can your model reason correctly about visual clues in long-horizon videos?

  3. Can your model behave appropriately in various social contexts?

05Taxonomy

We evaluate 3 Cognitive Dimensions and 5 Social Dimensions.

SMMU benchmark teaser showing long-horizon video context, cognitive tasks, and five social dimensions
03

Cognitive Dimensions

Comprehension

Recognize what is happening at a specific social moment.

Reasoning

Explain the social evidence behind an interpretation.

Prediction

Anticipate what is most likely to happen next.

05

Social Dimensions

06Evaluation Metric

Can your model maintain and update its understanding of videos as it unfolds?

Human understanding of character emotions, intent, perspectives, knowledge states, and relationships evolves as video context unfolds. However existing papers only ask independent MCQs.

Long-horizon video

Social representation over time

C0
Initial social state

Comprehend the situation before the pivotal event.

Establish
R
Pivotal event

Reason about the visual evidence that changes the interaction.

Interpret
C1
Outcome social state

Comprehend the situation after the event and update the representation.

Update
P
Social response

Predict the appropriate next behavior from the updated state.

Predict
Joint scoring

SMMU's Joint Evaluation Framework.

SUS

Social Understanding Score

Both comprehension checkpoints and reasoning must be correct.

SUS=C0 ∧ C1 ∧ R
SAS

Social Agency Score

A correct prediction must also show the model can act on its updated social representation. (e.g. Correct Comprehension, Wrong Reasoning)

SAS=SUS ∧ P

No partial score is awarded for partial understanding.

07Benchmark Results

Social understanding remains a hard frontier.

Even the strongest multimodal model trails the human baseline on joint social understanding and agency. Explore performance across every cognitive task and social dimension.

SMMU performance comparison across joint scores, cognitive tasks, and social dimensions.

Scroll horizontally to inspect every metric. Values are proportions; higher is better.

Selected model

08Examples

View our Dataset.

Test your Social Intelligence? Click on any video. The video will pause automatically to ask you questions at particular moments. Answer the questions to continue with the rest of the video.

Three SMMU benchmark examples showing video context, characters, social dimensions, and timestamped comprehension, reasoning, and prediction questions
09REFERENCE

Cite our work.

If SMMU supports your research, please cite the project. Final publication details will replace this placeholder.

@misc{smmu2026,
  title  = {SMMU: Benchmarking Social Intelligence of Multimodal Large Language Models},
  author = {SMMU Authors},
  year   = {2026},
  note   = {Placeholder citation - final bibliographic details forthcoming}
}