What is visible at the anchored region?
PhysAlign
A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning
Correct recognition does not guarantee correct physical-role grounding.
source problems
localized probes
paired +GT inputs
evaluated MLLMs
Paper-snapshot inventory · Paired variants are not additional unique probes.
THE MISSING DIAGNOSTIC
A model can read T—and still misunderstand the physics.
Answer accuracy collapses recognition, role assignment, and downstream reasoning into one outcome. PhysAlign tests the earlier question directly: did the model attach the observed content to the physical entity or relation supported by the problem?
Each query preserves the original diagram and conditions, points to one evidence location, and asks for the local physical correspondence.
Which entity or role does that content belong to?
Both must be correct on the same localized probe.
EVIDENCE IN · LOCAL TESTS OUT
Built around the source evidence.
Human-reviewed observational annotations become deterministic probes—without requiring a full diagram graph or exposing the target answer.
- 01Anchor the evidence
Fixed image regions or exact text spans identify what is being queried.
- 02Separate reading from role
Independent targets expose correctly read content assigned to the wrong entity.
- 03Control the local reading
Matched +GT variants provide verified local content while keeping the target role hidden.
T02-singleResolve an anchored mention to a candidate entity.
T03-imageRead a local label and identify its physical owner.
T03-textIdentify the owner of an anchored textual quantity.
THE GROUNDING BOTTLENECK
Strong perception is not reliable physical interpretation.
Across six MLLMs, PhysAlign finds a persistent gap between reading local content and assigning its physical role.
Even GPT-6 Astra misgrounds some correctly recognized evidence.
For InternVL3.5-8B, about half of correctly read cases are assigned to the wrong role.
Supplying a verified local reading can help, but repairs and harms both occur.
Descriptive paper results. Most paired 95% confidence intervals include zero; a positive point estimate is not a significance claim.
EXPLORE THE PAPER SNAPSHOT
One benchmark. Five complementary views.
Compare role grounding, same-probe joint correctness, content recognition, and independent problem solving. Change the metric to see why a single “overall” score would hide important behavior.
Development and test are evaluated jointly. These are release-derived diagnostics, not held-out test estimates.
Loading paper results…
Ranks describe reported point estimates under the selected metric, not statistical significance. No composite “overall score” is constructed.
Evaluation populations & score interpretation
These counts describe the paper snapshot. Local scores use parent- and task-balanced aggregation. SolveAcc is mean normalized credit on 986 original problems, including partial credit, not the fraction of fully correct solutions. Paired comparisons keep Base/+GT support fixed within each model, but observed support may differ across models.
All scores are percentages; changes are percentage points. Reported deltas and diagnostic ratios are transcribed from the paper, not recomputed from rounded columns. The +GT condition supplies local content, not perfect perception or the grounding answer. See manuscript §4.2, Table 2, and Appendix B.3–B.5.
SIX SOURCES · SEVEN DOMAINS
Broad evidence, precisely localized.
PhysAlign keeps each probe attached to the full original problem while localizing the evidence that defines the query.
- Source datasets
- SeePhys, LiveK12Bench, PhysElite, OlympiadBench, Gaokao-MM-Physics, and PhyX-OE.
- Evaluation unit
- A parent physics problem plus one localized evidence–role query.
- Quality control
- Answer-blind observation, human review, structural checks, and separated public/private records.
Coverage is intentionally reported as descriptive, not domain-balanced.
A frozen paper snapshot, with room to grow.
The leaderboard reproduces the paper's retained 3,341-probe inventory. Audit-corrected releases, changed protocols, and additional candidates must use a separate evaluation track; old scores must not be relabeled as results on new data.
Explore the dataset Coming soon ↗RESPONSES, NOT JUST SCORES
Two ways the local evidence changes the diagnosis.
These manuscript examples show observable outputs. They are not estimates of error frequency or claims about a model's hidden reasoning process.
Read correctly. Grounded to the wrong object.
Three Qwen3.5 models read D correctly but assign it to the point charge rather than the conducting pipe, under both Base and +GT.
Manuscript Figure 9 · A supplied reading does not repair this role error.A supplied reading repairs one local relation.
InternVL3.5-8B changes the graph-axis owner to the correct candidate after receiving the verified reading. Its independent solution remains incorrect.
Manuscript Figure 10 · A repaired local relation is not a complete solution.REPRODUCE · EXTEND · COMPARE
Evaluate another model.
Record the dataset revision, scorer commit, inference settings, and exact evaluation support. A result enters the ranked leaderboard only after maintainer review.
This static site displays results; it does not run models or receive private evaluation targets.
BUILD ON THE WORK
Citation
@misc{liang2026physalignbenchmarkevidencegroundedrole,
title={PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning},
author={Kecheng Liang and Haoyang Liu and Zexin Chen and Zirong Liu and Weixing Chen and Qiufeng Wang and Yang Liu and Liang Lin},
year={2026},
eprint={2609.33319},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.33319},
}