Seek-and-View Reasoning for
Multi-View Spatial Understanding

Qixiang Chen1, Cheng Zhang1, Fucai Ke1, Chi-Wing Fu2, Jianfei Cai1, Jingwen Ye1

1Monash University    2The Chinese University of Hong Kong

teaser image

(a) Existing approaches reason over fixed views, suffering from fragile cross-view alignment and the geometry-to-language bottleneck. (b) Our novel Seek-and-View reasoning paradigm seeks a question-relevant view (see bottom right) to make the spatial evidence directly observable.

Abstract

Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, we formulate a novel Seek-and-View reasoning approach to find implicit cross-view spatial evidence by locating a question-relevant view to support the spatial reasoning. To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model, a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning. Comprehensive experiments on six VLMs demonstrate consistent improvements on five benchmarks without fine-tuning. Overall, by revealing spatial evidence through view-grounded reasoning, Vantage can largely reduce reliance on language-based cross-view alignment and improve multi-view spatial understanding.

Motivation

What still limits multi-view spatial reasoning?

Fragile cross-view alignment

Strong semantic and contextual reasoning ability does not inherently enforce the structured, globally consistent geometric constraints required across views, leaving the alignment implicit and error-prone.

Geometry-to-language bottleneck

Upon compressing the spatial geometry into discrete semantic descriptions, the reasoning would potentially overlook fine-grained spatial information.

The usual paradigm

View-and-Reason

Reasoning stays passive over the fixed, sparse observations.

  • Verbalized geometry, such as cognitive maps, scene graphs, or bounding boxes
  • Injected geometric priors, such as 3D positional encoding, feature fusion, or reconstructive supervision
Ours

Seek-and-View

Reasoning seeks a question-relevant view that makes the spatial evidence directly observable.

  • Question-relevant, evidence-seeking, and reasoning-oriented visual evidence acquisition
  • Spatial cues grounded in scene geometry and presented directly as visual evidence

Humans can conceptualize the 3D structure of a scene and seek a vantage viewpoint that resolves the spatial relation. Vantage gives a VLM the same option.

Method

Overview of the Vantage pipeline

We propose Vantage, a two-stage training-free and model-agnostic framework with a viewpoint-grounded reasoning stage and a geometry-grounded evidence augmentation stage, explicitly coupling VLM reasoning with a 3D foundation model.

Vantage pipeline

(a) The VLM first analyses the question and identifies the observation needs. (b) It then plans a question-relevant view and generates reasoning guidance specifying which spatial cues should be examined. (c) A 3D foundation model reconstructs the scene and realizes the planned view in a gravity-aligned coordinate frame. (d) The VLM finally answers the question with the original observations augmented by the synthesized view and supporting reasoning context.

Results

Consistent gains, no additional training

Across five spatial reasoning benchmarks and six VLMs spanning a wide range of model scales, Vantage consistently improves the average accuracy of all six models without any additional training.

Base model + Vantage Average accuracy (%) over the five benchmarks

Per-benchmark accuracy

Accuracy per benchmark (%); see the paper for detailed subset-level results.

ModelMindCube-tinyMMSI-BenchBLINK
Multi-view
OmniSpatial
Perspective
SPINBenchAvg.% Gain
Training-scaled spatial model
SenseNova-SI-1.5-InternVL3-8B93.245.263.952.043.059.5–
Model-scaled spatial agent
GCA (Qwen3-VL-235B-A22B-Thinking)64.251.2–58.6–––
Vantage on base models
Qwen3-VL-4B26.027.941.442.153.238.1+13.1
+ Vantage38.933.742.147.253.443.1
Qwen3-VL-8B31.330.144.443.953.240.6+0.7
+ Vantage31.630.145.145.152.540.9
InternVL3-8B36.828.454.144.048.942.4+4.2
+ Vantage39.930.255.643.951.344.2
Qwen3.6-27B63.042.328.652.689.255.1+7.3
+ Vantage68.844.829.364.788.059.1
Gemma-4-31B53.536.836.151.771.149.8+9.4
+ Vantage61.937.144.455.173.954.5
GPT-5.444.234.448.147.459.346.7+14.3
+ Vantage54.341.154.153.763.953.4

The two reference rows achieve SOTA performance through either large-scale spatial-task training or tool-integrated agents built on substantially larger VLMs. In contrast, Vantage is training-free, leaves the base models unchanged, and generalizes across a wide range of model scales.

Qualitative examples

Qualitative example of Vantage with GPT-5.4. For clarity, we simplify the original model outputs and present only the key intermediate information relevant to the reasoning process.

Case study 1
Case study 2
Case study 3
Case study 4
Discussion

Where current VLMs + Seek-and-View still fail

Macro-averaged failure analysis across four models
Representative failure case for each of the four failure types

Left Macro-averaged failure analysis. Right Representative failure cases by type.

We perform a failure audit across multiple models and multi-view benchmarks to identify the dominant failure modes of Seek-and-View. Most residual errors arise from incorrect analysis and incorrect view planning, indicating that deciding what should be seen and how to reach it remains the main bottleneck. Unreliable synthesis is comparatively rare and mostly occurs under challenging visual conditions, suggesting that Seek-and-View can tolerate imperfect rendering as long as the synthesized view preserves sufficient geometric and semantic evidence. Reasoning failures arise when the final VLM under-utilizes useful evidence in the sought view or produces an inconsistent reasoning trace.

BibTeX

If you find this work interesting or relevant to your research, please consider citing our paper 😊

@misc{chen2026seekandviewreasoningmultiviewspatial,
      title={Seek-and-View Reasoning for Multi-View Spatial Understanding}, 
      author={Qixiang Chen and Cheng Zhang and Fucai Ke and Chi-Wing Fu and Jianfei Cai and Jingwen Ye},
      year={2026},
      eprint={2610.11810},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.11810}, 
}