See/Saw
Embodied Probing of Vision Transformer Attention in Virtual Reality
Embodied Probing of Vision Transformer Attention in Virtual Reality


2026
Virtual Reality, Computer Vision, Visual Computing
Unity (VR), C#, Python, Vision Transformer (ViT)
Published in SIGGRAPH Posters’26. Ziyan Xie, Sihui Lin (2026). “See/Saw: Embodied Probing of Vision Transformer Attention in Virtual Reality”. (ACM Student Research Competition, 2nd Place) [PAPER LINK]
See/Saw is a VR system that paints a frozen vision transformer’s patch grouping and attention onto the 3D surfaces the model is looking at. A controller aim selects a query patch, a frozen DINOv2 tokenizes the aimed view, k-means partitions the patches, and a frozen FeatUp read-out upsamples the partition to pixel resolution. A URP overlay pass then reprojects the cluster image onto scene meshes inside the capture frustum, with hue encoding the group and intensity encoding each group’s mean CLS-to-patch attention. The result is an interactive probe of how a ViT groups a scene from where the user is standing.


Embodied Probing of Vision Transformer Attention in Virtual Reality


2026
Virtual Reality, Computer Vision, Visual Computing
Unity (VR), C#, Python, Vision Transformer (ViT)
Published in SIGGRAPH Posters’26. Ziyan Xie, Sihui Lin (2026). “See/Saw: Embodied Probing of Vision Transformer Attention in Virtual Reality”. (ACM Student Research Competition, 2nd Place) [PAPER LINK]
See/Saw is a VR system that paints a frozen vision transformer’s patch grouping and attention onto the 3D surfaces the model is looking at. A controller aim selects a query patch, a frozen DINOv2 tokenizes the aimed view, k-means partitions the patches, and a frozen FeatUp read-out upsamples the partition to pixel resolution. A URP overlay pass then reprojects the cluster image onto scene meshes inside the capture frustum, with hue encoding the group and intensity encoding each group’s mean CLS-to-patch attention. The result is an interactive probe of how a ViT groups a scene from where the user is standing.


Embodied Probing of Vision Transformer Attention in Virtual Reality


2026
Virtual Reality, Computer Vision, Visual Computing
Unity (VR), C#, Python, Vision Transformer (ViT)
Published in SIGGRAPH Posters’26. Ziyan Xie, Sihui Lin (2026). “See/Saw: Embodied Probing of Vision Transformer Attention in Virtual Reality”. (ACM Student Research Competition, 2nd Place) [PAPER LINK]
See/Saw is a VR system that paints a frozen vision transformer’s patch grouping and attention onto the 3D surfaces the model is looking at. A controller aim selects a query patch, a frozen DINOv2 tokenizes the aimed view, k-means partitions the patches, and a frozen FeatUp read-out upsamples the partition to pixel resolution. A URP overlay pass then reprojects the cluster image onto scene meshes inside the capture frustum, with hue encoding the group and intensity encoding each group’s mean CLS-to-patch attention. The result is an interactive probe of how a ViT groups a scene from where the user is standing.

