See/Saw

Embodied Probing of Vision Transformer Attention in Virtual Reality

YEAR

2026

KEYWORDS

Virtual Reality, Computer Vision, Visual Computing

TOOLS

Unity (VR), C#, Python, Vision Transformer (ViT)

ADDITIONAL INFO

Published in SIGGRAPH Posters’26. Ziyan Xie, Sihui Lin (2026). “See/Saw: Embodied Probing of Vision Transformer Attention in Virtual Reality”. (ACM Student Research Competition, 2nd Place) [PAPER LINK]

About the project

See/Saw is a VR system that paints a frozen vision transformer’s patch grouping and attention onto the 3D surfaces the model is looking at. A controller aim selects a query patch, a frozen DINOv2 tokenizes the aimed view, k-means partitions the patches, and a frozen FeatUp read-out upsamples the partition to pixel resolution. A URP overlay pass then reprojects the cluster image onto scene meshes inside the capture frustum, with hue encoding the group and intensity encoding each group’s mean CLS-to-patch attention. The result is an interactive probe of how a ViT groups a scene from where the user is standing.

Model Pipeline
Smooth Scroll
This will hide itself!

See/Saw

Embodied Probing of Vision Transformer Attention in Virtual Reality

YEAR

2026

KEYWORDS

Virtual Reality, Computer Vision, Visual Computing

TOOLS

Unity (VR), C#, Python, Vision Transformer (ViT)

ADDITIONAL INFO

Published in SIGGRAPH Posters’26. Ziyan Xie, Sihui Lin (2026). “See/Saw: Embodied Probing of Vision Transformer Attention in Virtual Reality”. (ACM Student Research Competition, 2nd Place) [PAPER LINK]

About the project

See/Saw is a VR system that paints a frozen vision transformer’s patch grouping and attention onto the 3D surfaces the model is looking at. A controller aim selects a query patch, a frozen DINOv2 tokenizes the aimed view, k-means partitions the patches, and a frozen FeatUp read-out upsamples the partition to pixel resolution. A URP overlay pass then reprojects the cluster image onto scene meshes inside the capture frustum, with hue encoding the group and intensity encoding each group’s mean CLS-to-patch attention. The result is an interactive probe of how a ViT groups a scene from where the user is standing.

Model Pipeline
Smooth Scroll
This will hide itself!

See/Saw

Embodied Probing of Vision Transformer Attention in Virtual Reality

YEAR

2026

KEYWORDS

Virtual Reality, Computer Vision, Visual Computing

TOOLS

Unity (VR), C#, Python, Vision Transformer (ViT)

ADDITIONAL INFO

Published in SIGGRAPH Posters’26. Ziyan Xie, Sihui Lin (2026). “See/Saw: Embodied Probing of Vision Transformer Attention in Virtual Reality”. (ACM Student Research Competition, 2nd Place) [PAPER LINK]

About the project

See/Saw is a VR system that paints a frozen vision transformer’s patch grouping and attention onto the 3D surfaces the model is looking at. A controller aim selects a query patch, a frozen DINOv2 tokenizes the aimed view, k-means partitions the patches, and a frozen FeatUp read-out upsamples the partition to pixel resolution. A URP overlay pass then reprojects the cluster image onto scene meshes inside the capture frustum, with hue encoding the group and intensity encoding each group’s mean CLS-to-patch attention. The result is an interactive probe of how a ViT groups a scene from where the user is standing.

Model Pipeline
Smooth Scroll
This will hide itself!