Abstract

Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can be myopic to physical dynamics, or fine-tuned evaluators that can overfit to dataset-specific cues. We introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatiotemporal encoder and maps them to a scalar physical-consistency violation score through a lightweight scoring head. Higher scores indicate more severe physical violations, while lower scores indicate greater physical consistency. PhyProbe is trained with a unified objective combining pairwise ranking, regression on noisy scalar annotations, and score anchoring over heterogeneous supervision sources. Across real–generated and generated–generated benchmarks, PhyProbe reliably orders videos and anchors its scores to a stable reference scale. Although it receives no explicit general-preference labels, it also transfers competitively to broader human preference benchmarks.

Measuring a Spectrum, Not a Label

A video can look sharp and temporally coherent while still violating the dynamics of the physical world. PhyProbe treats physical consistency as a measurement problem: instead of only deciding whether something is wrong, it estimates how severe the violation is on a continuous scale anchored by real videos and clear failures.

Generated videos ranging from physically consistent to subtly and clearly inconsistent, motivating a continuous violation score.
Generated videos can remain visually realistic while spanning a wide range of physical violation severity.

How PhyProbe Works

PhyProbe uses a frozen 2B-parameter spatiotemporal visual encoder and trains only a lightweight 1M-parameter scoring head. Three complementary signals shape the score: pairwise rankings provide stable relative order, noisy human ratings provide approximate magnitude, and reference anchors map real videos near 0 and clear violations near 1. At inference time, the model needs only a single video—no prompt or reference pair.

01

Rank

Learn which of two videos contains the more severe physical violation.

02

Measure

Use scalar human ratings to ground relative order in approximate magnitude.

03

Anchor

Use targets of 0 for real videos and 1 for clear synthetic failures to anchor the score scale.

PhyProbe training and inference architecture using a frozen spatiotemporal encoder and lightweight scoring head.
One unified objective turns heterogeneous supervision into a continuous physical-consistency violation score.

Physical Video Benchmark Results

PhyProbe achieves the best weighted average and leads on three of four physical-consistency benchmarks. Its largest gains appear in real–generated and no-correspondence settings, while Gemini-3.1-Pro leads on VideoPhy2.

Pair types: R–G compares a real video with a generated video; G–G compares two generated videos.

Correspondence: image+text pairs share an input image and text prompt; text pairs share a prompt; no-correspondence pairs require neither.

Method ImplBR–G VP2G–G PhyDetExR–G VF2G–G Avg.
VideoPhy2-AutoEval28.028.530.536.332.1
V-JEPA2 Surprise33.346.349.045.546.6
VideoScore268.749.760.665.059.9
Gemini-3.1-Flash-Lite91.361.889.074.577.3
Gemini-3.1-Pro95.270.586.273.778.1
PhyProbe98.066.198.975.282.4

Pairwise accuracy (%). Higher is better. Avg. is weighted by benchmark pair count; the best result in each column is highlighted.

Evaluation settings: ImplausiBench[1] (image+text, R–G) · VideoPhy2[2] (text, G–G) · PhyDetEx[3] (no correspondence, R–G) · VideoFeedback2[4] (no correspondence, G–G).

Alignment with Human Judgments

We compare PhyProbe with human ratings after aligning the score direction so that higher means more physically plausible. Correlation ranges from −1 to 1: values closer to 1 indicate stronger agreement, while values near 0 indicate little linear agreement (Pearson) or rank agreement (Spearman).

Spearman ρ
Do PhyProbe and humans rank videos in the same order?
Pearson r
Do PhyProbe's score values rise and fall with human rating values?

Human-rating correlation

Higher is better
MethodVideoPhy2VideoFeedback2Average
ρrρrρr
VideoPhy2-AutoEval0.3590.3620.2350.2400.2980.302
V-JEPA2 Surprise0.0650.0650.1210.1330.0930.099
VideoScore20.1970.1600.4620.4300.3360.301
Gemini-3.1-Flash-Lite0.3020.3030.4190.4090.3620.357
Gemini-3.1-Pro0.2770.2400.4020.3830.3410.313
PhyProbe0.2930.3020.5960.6040.4580.466

Takeaway: VideoPhy2-AutoEval is strongest on its in-domain VideoPhy2 benchmark, but its correlation falls substantially on VideoFeedback2. PhyProbe is moderate on VideoPhy2, strongest on VideoFeedback2, and achieves the best cross-dataset average. Average correlations use an equal-weight Fisher-z aggregation.

  1. TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
  2. VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
  3. PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
  4. VideoScore2: Think before You Score in Generative Video Evaluation

Evaluator Analysis

Same Backbone, Better Evaluator

To isolate the contribution of PhyProbe's learning objective, we compare it with V-JEPA2 Surprise using the identical frozen V-JEPA2 ViT-H/16 backbone. PhyProbe improves both pairwise accuracy and human-rating correlation, showing that the gains do not come from the backbone alone.

Pairwise accuracy

%, higher is better
MethodImplBVideoPhy2PhyDetExVideoFeedback2
V-JEPA2 Surprise33.346.349.045.5
PhyProbe81.460.270.776.3
Improvement+48.1+13.9+21.7+30.8

Alignment with human judgments

Correlation, higher is better
MethodVideoPhy2VideoFeedback2
ρrρr
V-JEPA2 Surprise0.0650.0650.1210.133
PhyProbe0.3010.3030.6230.630
Improvement+0.236+0.238+0.502+0.497

Controlled comparison: both methods use V-JEPA2 ViT-H/16. The main results above use PhyProbe with PE-Core-G14-448. Accuracy improvements are in percentage points.

Cleaner Score Separation

PhyProbe assigns more distinct scores to plausible and implausible videos than VideoPhy2-AutoEval. AUC measures how reliably the score separates the two groups (0.5 is chance; 1.0 is perfect), while a larger Cohen's d indicates a wider gap between their score distributions. VideoFeedback2 results here use only the clearly implausible and plausible endpoints (human ratings 1 and 5), rather than the full rating range used in the correlation table.

Distribution separation

Higher is better
MethodImplausiBenchVideoFeedback2
AUCCohen's dAUCCohen's d
VideoPhy2-AutoEval0.5870.340.6670.74
PhyProbe0.8681.600.9803.26

Transfer to Human Preference

Although trained without explicit general-preference labels, PhyProbe also transfers to broader video preference benchmarks: MonetBench[5], Rapidata-I2V[6][7], and VideoGen-RewardBench[8]. This suggests that physical failures are an important component of perceived video quality rather than an isolated artifact category.

MonetBench

Pairwise accuracy (%)
MethodWith tiesWithout ties
Random33.249.9
V-JEPA2 Surprise52.453.8
VideoScore242.341.1
VideoPhy2-AutoEval22.516.5
VisionReward68.272.2
VideoReward56.158.2
PhyProbe72.771.0

Rapidata-I2V

Without ties
MethodAccuracy (%)
Random49.7
V-JEPA2 Surprise45.1
VideoScore245.8
VideoPhy2-AutoEval9.5
VisionReward47.1
VideoReward51.5
PhyProbe62.7

VideoGen-RewardBench

Pairwise accuracy (%)
Method Overall Motion Visual
tiesno ties tiesno ties tiesno ties
Random33.249.733.249.533.349.8
V-JEPA2 Surprise45.244.248.847.148.747.6
VideoScore257.458.954.360.655.360.1
VideoPhy2-AutoEval24.619.537.920.534.420.4
VisionReward64.667.654.661.354.859.1
VideoReward69.773.660.375.263.675.9
PhyProbe64.165.362.365.961.463.2

Pairwise accuracy (%). Higher is better; the strongest result in each column is highlighted. For tie-aware evaluation, PhyProbe uses the 20th percentile of absolute score differences on each benchmark as its tie threshold. Without-tie results compare raw scores.

  1. VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
  2. Rapidata: Image-to-Video Human Preference — Hailuo-02 vs. Marey
  3. Rapidata: Image-to-Video Human Preference — Seedance-1-Pro
  4. Improving Video Generation with Human Feedback
  5. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Interpretable Violation Scores

In these examples, coherent real and generated motion receives scores near zero, while morphing, object pop-in, interpenetration, erratic motion, and impossible rigidity receive higher violation scores.

Qualitative PhyProbe scores across videos with no violation, object morphing, pop in and out, interpenetration, erratic motion, and impossible rigidity.
Representative predictions from low-violation videos to clear physical failures. Lower PhyProbe scores are better.

Citation

@inproceedings{ku2026phyprobe,
  title   = {PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos},
  author  = {Ku, Max and Fan, Jiaojiao and Hao, Zekun and Ferroni, Francesco and Wang, Heng and Chen, Wenhu and Liu, Ming-Yu and Chattopadhyay, Prithvijit},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year    = {2026},
  eprint  = {2609.38377},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2609.38377}
}