PhysAlign

A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning

Correct recognition does not guarantee correct physical-role grounding.

FIG. 1Same evidence. Different premise.
The symbol T is read correctly. Its physical owner is not. PhysAlign scores the distinction at the source evidence.
01986

source problems

023,341

localized probes

03553

paired +GT inputs

046

evaluated MLLMs

Paper-snapshot inventory · Paired variants are not additional unique probes.

THE MISSING DIAGNOSTIC

A model can read T—and still misunderstand the physics.

Answer accuracy collapses recognition, role assignment, and downstream reasoning into one outcome. PhysAlign tests the earlier question directly: did the model attach the observed content to the physical entity or relation supported by the problem?

Each query preserves the original diagram and conditions, points to one evidence location, and asks for the local physical correspondence.

01 / READ
Local contentT

What is visible at the anchored region?

02 / GROUND
Physical ownerE4

Which entity or role does that content belong to?

03 / ALIGN
Joint correctnessC ∩ G

Both must be correct on the same localized probe.

EVIDENCE IN · LOCAL TESTS OUT

Built around the source evidence.

Human-reviewed observational annotations become deterministic probes—without requiring a full diagram graph or exposing the target answer.

  1. 01
    Anchor the evidence

    Fixed image regions or exact text spans identify what is being queried.

  2. 02
    Separate reading from role

    Independent targets expose correctly read content assigned to the wrong entity.

  3. 03
    Control the local reading

    Matched +GT variants provide verified local content while keeping the target role hidden.

FIGURE 2 Answer-blind annotation, human review, deterministic compilation, paired inputs, and release packaging.
T02-single

Resolve an anchored mention to a candidate entity.

T03-image

Read a local label and identify its physical owner.

T03-text

Identify the owner of an anchored textual quantity.

THE GROUNDING BOTTLENECK

Strong perception is not reliable physical interpretation.

Across six MLLMs, PhysAlign finds a persistent gap between reading local content and assigning its physical role.

13.8%
lowest conditional role error

Even GPT-6 Astra misgrounds some correctly recognized evidence.

50.6%
conditional role error

For InternVL3.5-8B, about half of correctly read cases are assigned to the wrong role.

+1.3–8.1
paired point-estimate change (pp)

Supplying a verified local reading can help, but repairs and harms both occur.

Descriptive paper results. Most paired 95% confidence intervals include zero; a positive point estimate is not a significance claim.

FIG. 4Where recognition and grounding diverge.
C/G denote content and grounding correctness on the same joint set. +GT provides the reading, not the grounding answer.

EXPLORE THE PAPER SNAPSHOT

One benchmark. Five complementary views.

Compare role grounding, same-probe joint correctness, content recognition, and independent problem solving. Change the metric to see why a single “overall” score would hide important behavior.

PAPER SNAPSHOT

Development and test are evaluated jointly. These are release-derived diagnostics, not held-out test estimates.

PhysAlign results

Loading paper results…

Result JSON ↗

Ranks describe reported point estimates under the selected metric, not statistical significance. No composite “overall score” is constructed.

Evaluation populations & score interpretation
G · Overall grounding986 / 3,341Parents / probes · GAcc
L · Joint evaluation311 / 412Parents / probes · CAcc, JAcc, GAccL
P · Paired eligible385 / 553Release eligibility, not observed coverage
Pm · Observed pairs311 / 412 or 304 / 397Open weights / API models, respectively

These counts describe the paper snapshot. Local scores use parent- and task-balanced aggregation. SolveAcc is mean normalized credit on 986 original problems, including partial credit, not the fraction of fully correct solutions. Paired comparisons keep Base/+GT support fixed within each model, but observed support may differ across models.

All scores are percentages; changes are percentage points. Reported deltas and diagnostic ratios are transcribed from the paper, not recomputed from rounded columns. The +GT condition supplies local content, not perfect perception or the grounding answer. See manuscript §4.2, Table 2, and Appendix B.3–B.5.

SIX SOURCES · SEVEN DOMAINS

Broad evidence, precisely localized.

FIGURE 3 Descriptive domain estimates and representative source problems, from middle-school to university-level physics.

PhysAlign keeps each probe attached to the full original problem while localizing the evidence that defines the query.

Source datasets
SeePhys, LiveK12Bench, PhysElite, OlympiadBench, Gaokao-MM-Physics, and PhyX-OE.
Evaluation unit
A parent physics problem plus one localized evidence–role query.
Quality control
Answer-blind observation, human review, structural checks, and separated public/private records.

Coverage is intentionally reported as descriptive, not domain-balanced.

VERSIONING MATTERS

A frozen paper snapshot, with room to grow.

The leaderboard reproduces the paper's retained 3,341-probe inventory. Audit-corrected releases, changed protocols, and additional candidates must use a separate evaluation track; old scores must not be relabeled as results on new data.

Explore the dataset Coming soon ↗

RESPONSES, NOT JUST SCORES

Two ways the local evidence changes the diagnosis.

These manuscript examples show observable outputs. They are not estimates of error frequency or claims about a model's hidden reasoning process.

CASE APERSISTS

Read correctly. Grounded to the wrong object.

Three Qwen3.5 models read D correctly but assign it to the point charge rather than the conducting pipe, under both Base and +GT.

Manuscript Figure 9 · A supplied reading does not repair this role error.
CASE BREPAIRED

A supplied reading repairs one local relation.

InternVL3.5-8B changes the graph-axis owner to the correct candidate after receiving the verified reading. Its independent solution remains incorrect.

Manuscript Figure 10 · A repaired local relation is not a complete solution.

REPRODUCE · EXTEND · COMPARE

Evaluate another model.

Record the dataset revision, scorer commit, inference settings, and exact evaluation support. A result enters the ranked leaderboard only after maintainer review.

This static site displays results; it does not run models or receive private evaluation targets.

BUILD ON THE WORK

Citation

@misc{liang2026physalignbenchmarkevidencegroundedrole,
  title={PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning},
  author={Kecheng Liang and Haoyang Liu and Zexin Chen and Zirong Liu and Weixing Chen and Qiufeng Wang and Yang Liu and Liang Lin},
  year={2026},
  eprint={2609.33319},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2609.33319},
}

Download BibTeX ↗ · arXiv:2609.33319 ↗

Model configuration