We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models.
Existing rotary relative position encodings commonly use image-plane positional coordinates, producing projection-dependent relative phases and inconsistent geometric cues for cross-projection attention.
We introduce G-ray, a ray-level relative position encoding whose rotary phases are parameterized by camera-local ray angles.
The same camera-local ray pair induces the same relative phase across projections, providing projection-invariant positional consistency.
G-ray can be used directly or integrated with existing encodings, retaining complementary geometric cues without additional learned parameters.
We validate G-ray in three host encodings, RoPE, GTA, and RayRoPE, across 3D reconstruction and novel-view synthesis (NVS).
Across three heterogeneous 3D reconstruction benchmarks at 50 views, G-ray leads all six averaged metrics and reduces mean pointmap relative error by 45.8% over MapAnything, with calibration supplied to both.
Trained exclusively on homogeneous pinhole images, the 3D reconstruction model handles mixed pinhole and non-pinhole inputs without retraining and remains competitive on homogeneous pinhole 3D reconstruction protocols.
For NVS, GTA and RayRoPE improve with G-ray under joint viewpoint and FoV variation.