LATENT VISUAL REASONINGResearch project · 2026

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

“Reasoning like humans. Grounded in what you see.”

1 University of Alabama at Birmingham2 Amazon AGI

* Work done during an internship at Amazon AGI.

THE IDEA, IN MOTION

From what we see to how we reason.

The original ReaLVR method diagram, showing latent rollout, answer contrast, and visual supervision.
0:00 / 0:54

Watch the inputs become tokens, follow the latent rollout, and see how visual evidence guides learning.

Extra supervision is training only · Original inference unchanged · Token highlights are illustrative.

Reasoning in latent space.
Grounded in the image.

Continuous latent tokens offer a compact way for multimodal models to reason. But a correct final answer provides little guidance on which visual details those tokens should preserve.

We identify a latent evidence-credit gap: latent tokens respond only weakly to image perturbations that change the correct answer. ReaLVR brings explicit visual supervision to the model’s own free-running latent trajectory, with no change to its architecture or inference procedure.

63.7%Five-task mean · Qwen2.5-VL-7B
+3.3 ppCompared with LVR-RL · same 7B backbone
6 backbonesThree model families · up to 235B total parameters
Paper teaser: real-world visual question, training supervision, and fixed-context latent token replacement diagnostic.
FROM THE PAPER What visual details do latent tokens preserve—and does the answer actually depend on them?

One trajectory.
Two complementary contrasts.

A training objective that makes visual evidence matter inside latent reasoning.

01 / TRAJECTORY

Learn on the model’s
own latent rollout.

Regenerate one autoregressive latent span with the current model. Keep gradients through the recurrence, before any answer tokens are supplied.

02 / WHAT

Preserve relevant
visual evidence.

A cosine margin aligns the trajectory with a relevant visual prototype and separates it from the hardest mismatched prototype. Use a region annotation when available; otherwise, pool the image.

03 / WHERE

Assign credit through
answer contrast.

Read the same latent span with correct and model-generated wrong answers. Their positive attention difference, plus a uniform baseline, allocates detached supervision weights.

ReaLVR method: relevant and mismatched visual evidence define a margin loss, while correct and wrong answer readout supplies detached weights to train the free-running latent rollout.
TRAINING ONLY The visual loss trains latent generation alongside unchanged GRPO rewards and advantages. Both supervision branches are absent at inference.

Better visual reasoning.
Across models and scales.

ReaLVR improves the average performance across the evaluated latent reasoning baselines.

FIVE-TASK AVERAGE · QWEN2.5-VL-7B

Accuracy (%)

LVR-RL
60.4
Monet-RL
62.2
ILVR-Stage2
62.9
ReaLVR
63.7
0100%

Mean ± standard deviation for ReaLVR: 63.7 ± 0.2%, across three seeds. Bars show the unweighted five-task mean.

SCALING TO 235B

Visual supervision
at frontier scale.

On Qwen3-VL-235B-A22B, ReaLVR improves all three evaluated benchmarks over LVR-SFT.

MMVP
80.9 → 81.9
BLINK
74.2 → 75.4
MME-RealWorld
70.4 → 71.0

235B total / 22B active parameters. This evaluation covers three tasks; no five-task mean is reported.

A closer look at the five benchmarks

Qwen2.5-VL-7B · accuracy (%) · three-seed mean

MethodMMVPBLINKHRBench-4KHRBench-8KMME-RealWorldMean
LVR-RL64.253.669.664.450.160.4
Monet-RL69.952.471.366.051.562.2
ILVR-Stage269.456.871.066.950.362.9
ReaLVR72.055.871.866.652.263.7

Selected latent reasoning baselines from Table 1. Bold indicates the best value in each column among the displayed methods.

Do the latent tokens
carry visual evidence?

Accuracy is only part of the story. We also probe whether latent representations respond to answer-changing visual evidence and whether answers depend on the selected tokens.

In a fixed-context diagnostic, replacing the eight most attended latent tokens reduces correct-answer probability by 4 percentage points for LVR and 11 points for ReaLVR. This measures local dependence with the remaining states held fixed.

LVR · correct-answer probability67% → 63%−4 pp
ReaLVR · correct-answer probability70% → 59%−11 pp
Answer-changing image perturbations test whether latent token trajectories respond to visual evidence.
EVIDENCE SENSITIVITY Controlled image changes reveal how latent trajectories react when the correct answer changes.

Scope & open questions

Fixed-context interventions measure local dependence, rather than a fully rerolled reasoning process or proof of faithful reasoning. Visual targets are less spatially specific without region annotations, and inference still uses a prescribed latent-token budget.

Cite ReaLVR

@misc{xiao2026realvr,
  title={Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence},
  author={Xi Xiao and Tianchen Zhao and Youngeun Kim and Zhuowei Li and Linghan Xu and Jiaye Wu and Zheng Zhang and Xiang Xu and Xuanbai Chen and Farhan Tejani and Jakub Zablocki and Julia Xu and Yifan Xing},
  year={2026},
  note={Preprint}
}