G-ray icon

G-ray

Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity

Shuo Zhang1 Xin Su1 Wei Wang1 Jun Liu1 Xinrui Zeng1 Yongsen Chen1 Chenjie Wang3 Guibo Zhu2 Jinqiao Wang2 Bin Luo1† Liangpei Zhang1
1 Wuhan University
2 Institute of Automation, Chinese Academy of Sciences and Wuhan AI Research
3 Rongyun Robot (Guizhou) Co., Ltd.
Corresponding author.
Abstract

We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encodings commonly use image-plane positional coordinates, producing projection-dependent relative phases and inconsistent geometric cues for cross-projection attention. We introduce G-ray, a ray-level relative position encoding whose rotary phases are parameterized by camera-local ray angles. The same camera-local ray pair induces the same relative phase across projections, providing projection-invariant positional consistency. G-ray can be used directly or integrated with existing encodings, retaining complementary geometric cues without additional learned parameters.

We validate G-ray in three host encodings, RoPE, GTA, and RayRoPE, across 3D reconstruction and novel-view synthesis (NVS). Across three heterogeneous 3D reconstruction benchmarks at 50 views, G-ray leads all six averaged metrics and reduces mean pointmap relative error by 45.8% over MapAnything, with calibration supplied to both. Trained exclusively on homogeneous pinhole images, the 3D reconstruction model handles mixed pinhole and non-pinhole inputs without retraining and remains competitive on homogeneous pinhole 3D reconstruction protocols. For NVS, GTA and RayRoPE improve with G-ray under joint viewpoint and FoV variation.

0:00 / 0:00

G-ray for 3D Reconstruction under Camera Heterogeneity

We compare feed-forward 3D reconstruction when the input cameras differ in viewpoint, FoV, and projection model. G-ray maps their image-plane positions to camera-local ray angles, providing relative attention with a common positional interface across heterogeneous camera inputs.

Ground Truth

G-ray (w/C)

MapAnything (w/C)

VGGT-Ω (w/o C)

Select Dataset:

Select Scene (Input Views):

G-ray for Novel-View Synthesis under FoV Variation

We train the LVSM model with different positional encodings under FoV variation and compare their novel-view synthesis results below.

Left Video:

Right Video:

RayRoPE
RayRoPE + G-ray

Select Scene:

RE10K RE10K 0 Ref 0 RE10K 0 Ref 1
RE10K RE10K 1 Ref 0 RE10K 1 Ref 1
RE10K RE10K 2 Ref 0 RE10K 2 Ref 1
RE10K RE10K 3 Ref 0 RE10K 3 Ref 1
RE10K RE10K 4 Ref 0 RE10K 4 Ref 1
RE10K RE10K 5 Ref 0 RE10K 5 Ref 1
Objv Objaverse 0 Ref 0 Objaverse 0 Ref 1 Objaverse 0 Ref 2 Objaverse 0 Ref 3
Objv Objaverse 1 Ref 0 Objaverse 1 Ref 1 Objaverse 1 Ref 2 Objaverse 1 Ref 3
Objv Objaverse 2 Ref 0 Objaverse 2 Ref 1 Objaverse 2 Ref 2 Objaverse 2 Ref 3
Objv Objaverse 3 Ref 0 Objaverse 3 Ref 1 Objaverse 3 Ref 2 Objaverse 3 Ref 3
Objv Objaverse 4 Ref 0 Objaverse 4 Ref 1 Objaverse 4 Ref 2 Objaverse 4 Ref 3

G-ray Aligns Cross-View Attention under FoV Variation

We visualize cross-view attention on ETH3D image pairs under controlled FoV variation. The top two rows show wide-FoV images and their resized narrow-FoV crops. Red rectangles mark the shared FoV, and dashed lines link each region to its crop. Attention maps average wide-to-narrow responses over global-attention blocks and heads. The right column shows position-only priors (PE-only), computed for the rotary encodings using constant query and key features before rotation. In the weakly textured examples, G-ray concentrates attention more clearly within the shared FoV than grid-index RoPE. Its PE-only prior reflects angular alignment, whereas the grid-index prior is nearly uniform.

Cross-view attention under controlled FoV variation on ETH3D
Cross-view attention under controlled FoV variation on ETH3D. The top two rows show wide-FoV images and their resized narrow-FoV crops. Red rectangles mark the shared FoV, and dashed lines link each region to its crop. Attention maps average wide-to-narrow responses over global-attention blocks and heads. The right column shows position-only priors (PE-only).

Acknowledgements

This work was supported by the National Key R&D Program of China under Grants 2022ZD0160601 and 2022YFB3903404.

Citation

@article{zhang2026gray,
  title={G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity},
  author={Zhang, Shuo and Su, Xin and Wang, Wei and Liu, Jun and Zeng, Xinrui and Chen, Yongsen and Wang, Chenjie and Zhu, Guibo and Wang, Jinqiao and Luo, Bin and Zhang, Liangpei},
  journal={arXiv preprint arXiv:2609.15018},
  year={2026}
}